Coding agent scores: why 83.4% hides real costs
The Artificial Analysis Coding Agent Index v1.1 reveals how an 83.4% TerminalBench score masks variance in token efficiency and tool use.
The Artificial Analysis Coding Agent Index v1.1 reveals how an 83.4% TerminalBench score masks variance in token efficiency and tool use.
TerminalBench v2.1 uses 89 curated tasks to test if AI agents can execute complex system commands rather than just generating static code snippets.