Evaluation types for AI agent reliability LangSmith supports four distinct evaluator types to validate agent performance. Learn how offline testing and LLM-as-judge scoring catch regressions early. Jul 14, 2026 Marcus Chen 15 min read