Agent-as-a-Judge: Evaluate Agents with Agents
Mingchen ZhugeChangsheng Zhao 0002Dylan R. AshleyWenyi WangDmitrii KhizbullinYunyang XiongZechun LiuErnie ChangRaghuraman KrishnamoorthiYuandong Tian
Proposes the Agent-as-a-Judge framework alongside the DevAI benchmark to evaluate autonomous agents throughout their intermediate task trajectories, matching human-level evaluation accuracy while cutting assessment time and cost by over 97 percent.
Autonomous agentic systems are increasingly deployed to solve complex, multi-step engineering tasks such as end-to-end software development. However, existing evaluation methods fail to adequately assess these systems: traditional benchmarks focus solely on final binary outcomes rather than intermediate steps, while manual human evaluation across full workflows is prohibitively expensive and prone to inconsistency. At the same time, standard language model evaluators lack the interactive tools needed to inspect multi-file environments, complex dependencies, and execution logs.
The article introduces Agent-as-a-Judge, a framework that employs autonomous agentic systems equipped with specialized tools to evaluate other agentic systems across entire task-solving trajectories. To validate this framework, the article also presents DevAI, a new benchmark comprising 55 realistic AI development tasks featuring 365 hierarchical requirements structured as directed acyclic graphs and 125 soft preferences.
The researchers evaluated three leading open-source coding agents—MetaGPT, GPT-Pilot, and OpenHands—on DevAI using human expert panels, standard language model judges, and the proposed Agent-as-a-Judge framework. A modular architecture was developed for the judge agent, integrating components to parse workspace graphs, inspect multimodal files across 33 formats, locate specific target files, and retrieve execution trajectory feedback. Performance was measured by alignment with human consensus and absolute deviation from expert ground truth.
The investigation produced four central findings. First, current autonomous coding agents struggle with complete real-world workflows: top-performing frameworks satisfied only about 29% of dependent requirements and completed only 1.81% of full tasks. Second, Agent-as-a-Judge dramatically outperformed traditional language model evaluators, achieving an alignment rate of approximately 90% with human consensus compared to roughly 65% to 70% for standard language models. Third, Agent-as-a-Judge proved more reliable than individual human evaluators, whose agreement with the consensus ranged from 76% to 92%. Fourth, the automated agentic evaluation reduced human labor time by 97.72% (from 86.5 hours to under two hours) and financial cost by 97.64% (from roughly 30.58).
These findings demonstrate that agentic evaluation provides an accurate, scalable alternative to expensive human grading panels, removing a major bottleneck in artificial intelligence development. Beyond lowering testing expenses, the framework enables granular, step-by-step diagnostic feedback on intermediate milestones rather than coarse pass-fail scores. This rich feedback can be fed directly back into developer agents to support autonomous debugging, error correction, and iterative self-improvement.
Organizations developing or deploying autonomous agents should transition from static outcome benchmarks to intermediate, trajectory-aware agentic judges. Engineering teams should adopt the modular combination of workspace parsing, file reading, target locating, and trajectory retrieval while tailoring validation modules to domain-specific criteria. Further work should focus on integrating dynamic feedback loops between judge agents and developer agents to enable automated iterative refinement during execution.
The primary limitations include edge cases where the automated judge was misled by synthetic placeholder datasets or subtle semantic requirements, particularly during complex data preprocessing. Although the benchmark focused on 55 curated machine learning development tasks using a single primary language model backend, ablation tests across multiple underlying models showed consistently high alignment, providing strong confidence in the stability and generalizability of the Agent-as-a-Judge paradigm.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). Provides a comprehensive foundational survey of the LLM-as-a-Judge paradigm, whose failure modes and intermediate-feedback limitations Agent-as-a-Judge directly addresses and extends.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). Establishes the standard LLM-as-a-Judge methodology for evaluating generative outputs, serving as the core baseline that the source paper seeks to generalize into multi-step agentic evaluation.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). Examines the vulnerabilities and superficiality of automated LLM evaluators on instruction following, highlighting the exact evaluation bottlenecks that necessitate trajectory-level agentic judges.
- Paper: SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, John Yang et al. (2024). Introduces agent-computer interfaces for autonomous code generation and problem-solving, providing direct context for the coding agent workflows evaluated by Agent-as-a-Judge.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Presents a structured architectural taxonomy of LLM-based autonomous agents, detailing the planning and action cycles that agent evaluators must inspect.
- Paper: The Verification Horizon: No Silver Bullet for Coding Agent Rewards, Binghai Wang et al. (2026). Investigates verification and reward design for coding agents—including autonomous agent evaluators for long-horizon repositories—directly extending the verification paradigms introduced in the source.
- Paper: Automated Design of Agentic Systems, Shengran Hu et al. (2025). Applies meta-agent evaluation and iterative generation loops to search and refine agent architectures automatically, building upon the principles of agents assessing agents.
- Paper: Self Improvement via Fast Tree-search, Xinghong Fu et al. (2026). Uses model judges within a fast tree-search optimization framework to evaluate and recursively improve autonomous coding agents efficiently.
- Paper: AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security, Dongrui Liu et al. (2026). Extends trajectory-level agentic evaluation to multi-step diagnostic guardrails and safety verification across agent action execution paths.
- Paper: Towards a Science of Scaling Agent Systems, Yubin Kim et al. (2025). Analyzes the scaling behaviors, overheads, and trade-offs of single versus multi-agent systems across diverse complex workflows.
