Agent evaluation: Why 100% accuracy still fails safety
Standard benchmarks miss critical failures. This framework uses an internal LLM evaluator to audit multiturn conversations against safety policies and accuracy.
Standard benchmarks miss critical failures. This framework uses an internal LLM evaluator to audit multiturn conversations against safety policies and accuracy.
With 276 commits, this framework uses an LLM evaluator to test agent reasoning via multiturn dialogue instead of static string matching.