Agent testing needs an LLM evaluator, not static scripts With 276 commits, this framework uses an LLM evaluator to test agent reasoning via multiturn dialogue instead of static string matching. Jul 4, 2026 Diego Alvarez 12 min read