Agent-as-a-Judge: Evaluate Agents with Agents

Mingchen ZhugeChangsheng Zhao 0002Dylan R. AshleyWenyi WangDmitrii KhizbullinYunyang XiongZechun LiuErnie ChangRaghuraman KrishnamoorthiYuandong Tian

article2025ICML138 citations

Proposes the Agent-as-a-Judge framework alongside the DevAI benchmark to evaluate autonomous agents throughout their intermediate task trajectories, matching human-level evaluation accuracy while cutting assessment time and cost by over 97 percent.

Listen

Autonomous agentic systems are increasingly deployed to solve complex, multi-step engineering tasks such as end-to-end software development. However, existing evaluation methods fail to adequately assess these systems: traditional benchmarks focus solely on final binary outcomes rather than intermediate steps, while manual human evaluation across full workflows is prohibitively expensive and prone to inconsistency. At the same time, standard language model evaluators lack the interactive tools needed to inspect multi-file environments, complex dependencies, and execution logs.

The article introduces Agent-as-a-Judge, a framework that employs autonomous agentic systems equipped with specialized tools to evaluate other agentic systems across entire task-solving trajectories. To validate this framework, the article also presents DevAI, a new benchmark comprising 55 realistic AI development tasks featuring 365 hierarchical requirements structured as directed acyclic graphs and 125 soft preferences.

The researchers evaluated three leading open-source coding agents—MetaGPT, GPT-Pilot, and OpenHands—on DevAI using human expert panels, standard language model judges, and the proposed Agent-as-a-Judge framework. A modular architecture was developed for the judge agent, integrating components to parse workspace graphs, inspect multimodal files across 33 formats, locate specific target files, and retrieve execution trajectory feedback. Performance was measured by alignment with human consensus and absolute deviation from expert ground truth.

The investigation produced four central findings. First, current autonomous coding agents struggle with complete real-world workflows: top-performing frameworks satisfied only about 29% of dependent requirements and completed only 1.81% of full tasks. Second, Agent-as-a-Judge dramatically outperformed traditional language model evaluators, achieving an alignment rate of approximately 90% with human consensus compared to roughly 65% to 70% for standard language models. Third, Agent-as-a-Judge proved more reliable than individual human evaluators, whose agreement with the consensus ranged from 76% to 92%. Fourth, the automated agentic evaluation reduced human labor time by 97.72% (from 86.5 hours to under two hours) and financial cost by 97.64% (from roughly 1,298to1,298 to 30.58).

These findings demonstrate that agentic evaluation provides an accurate, scalable alternative to expensive human grading panels, removing a major bottleneck in artificial intelligence development. Beyond lowering testing expenses, the framework enables granular, step-by-step diagnostic feedback on intermediate milestones rather than coarse pass-fail scores. This rich feedback can be fed directly back into developer agents to support autonomous debugging, error correction, and iterative self-improvement.

Organizations developing or deploying autonomous agents should transition from static outcome benchmarks to intermediate, trajectory-aware agentic judges. Engineering teams should adopt the modular combination of workspace parsing, file reading, target locating, and trajectory retrieval while tailoring validation modules to domain-specific criteria. Further work should focus on integrating dynamic feedback loops between judge agents and developer agents to enable automated iterative refinement during execution.

The primary limitations include edge cases where the automated judge was misled by synthetic placeholder datasets or subtle semantic requirements, particularly during complex data preprocessing. Although the benchmark focused on 55 curated machine learning development tasks using a single primary language model backend, ablation tests across multiple underlying models showed consistently high alignment, providing strong confidence in the stability and generalizability of the Agent-as-a-Judge paradigm.

Cover for Agent-as-a-Judge: Evaluate Agents with Agents

Abstract

Contemporary evaluation techniques are inadequate for agentic systems. These approaches either focus exclusively on final outcomes—ignoring the step-by-step nature of the thinking done by agentic systems—or require excessive manual labour. To address this, we introduce the Agent-as-a-Judge framework, wherein agentic systems are used to evaluate agentic systems. This is a natural extension of the LLM-as-a-Judge framework, incorporating agentic features that enable intermediate feedback for the entire task-solving processes for more precise evaluations. We apply the Agent-as-a-Judge framework to the task of code generation. To overcome issues with existing benchmarks and provide a proof-of-concept testbed for Agent-as-a-Judge, we present DevAI, a new benchmark of 55 realistic AI code generation tasks. DevAI includes rich manual annotations, like a total of 365 hierarchical solution requirements, which make it particularly suitable for an agentic evaluator. We benchmark three of the top code-generating agentic systems using Agent-as-a-Judge and find that our framework dramatically outperforms LLM-as-a-Judge and is as reliable as our human evaluation baseline. Altogether, we believe that this work represents a concrete step towards enabling vastly more sophisticated agentic systems. To help that, our dataset and the full implementation of Agent-as-a-Judge will be publically available at https://github.com/metauto-ai/agent-as-a-judge

Table of Contents

  • 1. Introduction
  • 2. (Step 1) Crafting a Benchmark for Automated AI Development
  • 2.1. Motivation
  • 2.2. The DevAI Dataset
  • 2.3. Preliminary Benchmark
  • 3. (Step 2) Manual Evaluation on DevAI (Human-as-a-Judge)
  • 3.1. Benchmark Baselines by Human-as-a-Judge
  • 3.2. Judging Human-as-a-Judge
  • 4. (Step 3&4) Evaluating Agents with Agents (Agent-as-a-Judge)
  • 4.1. Specific Agent-as-a-Judge for Code Generation
  • 4.2. Judging Agent-as-a-Judge and LLM-as-a-Judge
  • 4.3. Ablations For Agent-as-a-Judge
  • 4.4. Cost Analysis
  • 5. Related Work
  • 6. Discussion and Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Task Sample
  • B. Experiment Designs
  • B.1. Summary of Experiments
  • B.2. Judges and Subjects of Evaluation
  • C. Agent-as-a-Judge Pipeline
  • D. Extend Related Work
  • E. The Procedures of Creating DevAI Dataset
  • E.1. Manually draft user queries
  • E.2. Set Judging Criteria
  • E.3. Building Dependency Among Requirements
  • E.4. Refine the dataset
  • E.5. Analyse the dataset
  • E.6. Auxiliary Information
  • E.7. A Json Format of Our Sample
  • F. User experiences of code-generation agentic systems
  • G. More DevAI dataset samples
  • H. Human Evaluation Procedure
  • I. Suggest Constraints
  • J. Collected Trajectories
  • J.1. Schema
  • J.2. A Sample of Trajectory
  • K. Ablations of Agent-as-a-Judge
  • K.1. Component Ablations
  • K.2. Search Algorithms in Search Module
  • K.3. Search Algorithms in Retrieve Module
  • L. Prompt Demos of Agent-as-a-Judge
  • L.1. System Prompt for Agent-as-a-Judge
  • L.2. System Prompt for Locate Module
  • L.3. System Prompt for Retrieve Module
  • L.4. Prompt for Ask Module (for requirement check)
  • L.5. Prompt for Locate Module
  • M. Judge Evidences Collected from Agent-as-a-Judge
  • N. Analysis of failure cases
  • O. Sensitivity w.r.t the choice of the backend LLM
  • P. Additional Human Evaluation Details

Knowls

  1. Knowl 1 — Agent-as-a-Judge Framework for Evaluating Autonomous Systems

    model/method

    The Agent-as-a-Judge framework evaluates autonomous agentic systems by deploying an agentic evaluator capable of inspecting intermediate development steps, execution trajectories, and generated artifacts rather than inspecting only conversational outputs or binary end-to-end task completion.

    The framework is composed of modular components:

    • Graph Module: Parses the project directory structure, code modules, and file dependencies to create a structural graph and segment code into atomic snippets.
    • Locate Module: Maps requirements and queries to relevant files and subdirectories within the workspace.
    • Read Module: Ingests and interprets multimodal files across 33 distinct formats, including source code, images, video demonstrations, and logs.
    • Retrieve Module: Filters and extracts relevant interaction steps, execution feedback, and error messages from long runtime trajectories.
    • Ask Module: Formulates the final verification decision for a target requirement given the dynamically gathered evidence, outputting either <SATISFIED> or <UNSATISFIED> alongside a concise justification.

    While additional modules such as explicit action Planning and judgment Memory were explored, ablations indicate that omitting planning and memory achieves superior reliability by preventing error cascading and noisy multi-step inference.

  2. Knowl 2 — The DevAI Benchmark for Automated AI Development

    definition

    DevAI is an evaluation benchmark consisting of 55 comprehensive, real-world AI development tasks designed to assess the multi-step problem-solving capabilities of code-generating agentic systems.

    Each task specification comprises:

    1. User Query: A plain-text prompt specifying the AI development task (spanning machine learning subfields such as supervised learning, reinforcement learning, computer vision, natural language processing, time series forecasting, and generative models).
    2. Hierarchical Requirements: A set of milestone criteria (totaling 365 across the benchmark) structured as a Directed Acyclic Graph (DAG) with explicit prerequisite dependencies (DD). Requirements encompass dataset preparation, data preprocessing, machine learning architecture implementation, model persistence, evaluation metrics logging, visualization generation, and human-computer interaction (HCI) interfaces.
    3. Preferences: Optional, softer criteria (totaling 125 across the benchmark) describing qualitative properties such as code cleanliness, robust edge-case handling, or UI responsiveness.

    Unlike traditional benchmarks that score solely on final pass rates (e.g., pass@1pass@1), DevAI evaluates intermediate milestone completion both independently (disregarding prerequisites) and dependently (requiring all prerequisite milestones in the DAG to be satisfied).

  3. Knowl 3 — Alignment and Judge Shift of Agent-as-a-Judge versus LLM-as-a-Judge

    data/table

    Evaluating agentic outputs against human consensus reveals that Agent-as-a-Judge achieves higher alignment rates and substantially lower judge shifts compared to a standard LLM-as-a-Judge across black-box (workspace only) and gray-box (workspace plus execution trajectory) settings.

    Evaluation Setting Metric MetaGPT GPT-Pilot OpenHands
    Black-Box Setting
    LLM-as-a-Judge Alignment Rate ↑\uparrow 84.15% 65.30% 60.38%
    Agent-as-a-Judge Alignment Rate ↑\uparrow 88.52% 83.88% 90.44%
    Gray-Box Setting
    LLM-as-a-Judge Alignment Rate ↑\uparrow 68.86% 71.85% 70.76%
    Agent-as-a-Judge Alignment Rate ↑\uparrow 92.07% 86.61% 90.16%
    Human Baseline Alignment
    Average of Individual Evaluators 89.34% 84.88% 85.70%
    Human Majority Vote (3 Experts) 95.08% 93.98% 94.26%

    Agent-as-a-Judge consistently achieves 83.88%83.88\%–92.07%92.07\% alignment with the human consensus across all target systems, outperforming both LLM-as-a-Judge (60.38%60.38\%–84.15%84.15\%) and the average individual human evaluator (84.88%84.88\%–89.34%89.34\%).

  4. Knowl 4 — Benchmark Baseline Performance on DevAI

    data/table

    Three open-source code-generation agent frameworks—MetaGPT (Data Interpreter), GPT-Pilot (v0.2.13), and OpenHands (CodeAct v1.9)—were evaluated on DevAI using gpt-4o-2024-05-13 with a 30-minute (1800-second) task execution limit.

    Metric MetaGPT GPT-Pilot OpenHands
    Average Execution Cost (USD) $1.19 $3.92 $6.38
    Average Execution Time (s) 775.29 1622.38 362.41
    Average Saved Code Files 0.42 3.84 2.53
    Average Saved Code Lines 11.15 273.33 96.56
    Requirements Met (Independent) 22.13% 44.80% 42.89%
    Requirements Met (Dependent DAG) 6.55% 28.96% 28.68%
    Self-Termination Rate 41.81% 5.45% 54.54%
    Full Task Solve Rate 0.00% 1.81% 1.81%

    Current state-of-the-art agent frameworks solve only 1.81%1.81\% of full DevAI tasks (1 out of 55 tasks), while satisfying roughly 29%29\% of dependent development milestones, demonstrating that DevAI provides non-sparse intermediate evaluation signals without suffering from ceiling effects.

  5. Knowl 5 — Component Ablation Studies for Agent-as-a-Judge

    data/table

    The influence of incrementally integrating modules into the Agent-as-a-Judge framework was evaluated on the OpenHands baseline against human consensus judgments.

    Configuration / Modular Addition Human Alignment Rate
    `+ ask` 65.03%
    `+ graph` 75.95%
    `+ read` 82.24%
    `+ locate` 90.44%
    `+ search` (BM25) 86.06%
    `+ retrieve` (Gray-Box Trajectory) 90.16%
    `+ planning` 88.52%
    `+ memory` 87.97%

    Key takeaways from the modular ablations include:

    • Adding file structure parsing (graph), file reading (read), and requirement targeting (locate) progressively increases alignment from 65.03%65.03\% to 90.44%90.44\%.
    • Information retrieval (search) via BM25, Sentence-BERT (87.70%87.70\%), or fuzzy matching (85.52%85.52\%) degrades performance by introducing retrieval noise on small-to-moderate codebases.
    • Explicit planning and judgment memory degrade alignment due to decision instability and error cascading across dependent evaluations.
  6. Knowl 6 — Cost and Time Efficiency of Agent-as-a-Judge

    empirical result

    Automated evaluation via Agent-as-a-Judge substantially reduces evaluation overhead compared to human panel scoring:

    • Human-as-a-Judge: Evaluating the full DevAI benchmark across three expert annotators required a cumulative 86.5 hours (58 hours for initial scoring plus 28.5 hours for consensus deliberation). At an assumed expert wage of $15.00/hour, this amounts to $1,297.50 USD.
    • Agent-as-a-Judge: Evaluated the benchmark in 118.43 minutes total wall-clock time at an API cost of $30.58 USD, representing a 97.72%97.72\% reduction in evaluation time and a 97.64%97.64\% reduction in financial cost.
    • LLM-as-a-Judge: Required 10.99 minutes and $29.63 USD, but achieved significantly lower agreement with human consensus due to lacking targeted context locating and multimodal file parsing.
  7. Knowl 7 — Human Evaluator Disagreement and the Reliability Hierarchy

    empirical result

    Human expert evaluation of multi-step code generation exhibits notable inter-annotator variance:

    • Pairwise disagreement among three experienced AI experts (each having ≥5\ge 5 years of experience) ranged between 10%10\% and 30%30\%, with individual error rates relative to post-deliberation consensus reaching up to 23.77%23.77\%.
    • Combining independent evaluator judgments via majority voting reduced the disagreement error rate to 6.01%6.01\%. An extended panel of 10 independent evaluators demonstrated a 97.67%97.67\% agreement with the 3-expert majority vote and a 95.23%95.23\% agreement with the final consensus.
    • Based on precision-recall analysis and alignment rates, the overall reliability ordering of evaluation paradigms is: LLM-as-a-Judge<Single-Human-as-a-Judge<Agent-as-a-Judge<Ensemble/Consensus Human Judges\text{LLM-as-a-Judge} < \text{Single-Human-as-a-Judge} < \text{Agent-as-a-Judge} < \text{Ensemble/Consensus Human Judges} Agent-as-a-Judge surpasses the reliability of a single human annotator while approaching multi-expert consensus.
  8. Knowl 8 — Sensitivity of Agent-as-a-Judge to LLM Backbones

    data/table

    The alignment rate between Agent-as-a-Judge and human consensus was measured across multiple underlying large language models:

    Backbone Model Version Parameters Alignment Rate (%)
    LLaMA 3.2 90B 87.76%
    Qwen Coder 2.5 32B 88.73%
    ChatGPT gpt-4o-2024-05-13 Unknown 90.16%
    Claude claude-3-5-sonnet-20241022 Unknown 92.95%

    While all tested backbones achieve robust alignment (>87%>87\%), claude-3-5-sonnet-20241022 achieved the highest agreement (92.95%92.95\%), attributed to enhanced tool-calling and agentic reasoning capabilities.

  9. Knowl 9 — Trajectory Truncation Strategies for Execution Log Retrieval

    empirical result

    When evaluating large interaction trajectories in the gray-box setting (analyzed on GPT-Pilot), information density is non-uniformly distributed across steps:

    • Trajectory-Level Truncation: Truncating early steps (head truncation) preserves an alignment rate of 86.61%86.61\%, whereas truncating late steps (tail truncation) degrades alignment to 82.51%82.51\%. This occurs because the end of the trajectory contains dense information regarding the final project state.
    • Step-Level Truncation: Within an individual step log, middle truncation achieves 86.61%86.61\% alignment, outperforming tail truncation (83.88%83.88\%) and head truncation (86.34%86.34\%). This is because runtime errors typically manifest near the beginning of log outputs, while file paths and stack trace origins appear at the end.
  10. Knowl 10 — Failure Modes and Categorical Limitations of Agent-as-a-Judge

    limitation

    Analysis of judgment discrepancies between Agent-as-a-Judge and expert human consensus identified distinct failure distributions across task categories:

    1. Distribution of Errors: Discrepancies were highest in Data Preprocessing and Postprocessing (10 errors) and Dataset or Environment Setup (8 errors), but lowest in Performance Metrics (3), Visualization (3), and Human-Computer Interaction (3).
    2. Synthetic / Mock Artifact Deception: When an agentic developer synthesizes a fake or mock dataset file under the target path name rather than downloading the real dataset, the judge agent can be misled by file path presence and superficial format consistency, whereas humans readily spot the fabrication.
    3. Nuanced / Dynamic Logic Oversight: The judge agent occasionally misses semantic nuances in criteria—such as verifying whether hyperparameters were dynamically scheduled versus statically defined in train.py.

Coverage note — Omitted introductory text, external related work summaries, generic prompt templates from Appendix L, and repetitive task JSON listings from Appendix G, as they do not constitute standalone contributed findings.

References

  1. 1.Anthropic. The claude 3 model family: Opus, sonnet, haiku, 2024. URL https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf.
  2. 2.Arora, D., Sonwane, A., Wadhwa, N., Mehrotra, A., Utpala, S., Bairi, R., Kanade, A., and Natarajan, N. MASAI: Modular architecture for software-engineering ai agents. arXiv preprint arXiv:2406.11638, 2024.
  3. 3.Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.
  4. 4.Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., Hui, B., Ji, L., Li, M., Lin, J., Lin, R., Liu, D., Liu, G., Lu, C., Lu, K., Ma, J., Men, R., Ren, X., Ren, X., Tan, C., Tan, S., Tu, J., Wang, P., Wang, S., Wang, W., Wu, S., Xu, B., Xu, J., Yang, A., Yang, H., Yang, J., Yang, S., Yao, Y., Yu, B., Yuan, H., Yuan, Z., Zhang, J., Zhang, X., Zhang, Y., Zhang, Z., Zhou, C., Zhou, J., Zhou, X., and Zhu, T. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023.
  5. 5.Bavaresco, A., Bernardi, R., Bertolazzi, L., Elliott, D., Fernández, R., Gatt, A., Ghaleb, E., Giulianelli, M., Hanna, M., Koller, A., et al. LLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks. arXiv preprint arXiv:2406.18403, 2024.
  6. 6.Bogin, B., Yang, K., Gupta, S., Richardson, K., Bransom, E., Clark, P., Sabharwal, A., and Khot, T. SUPER: Evaluating agents on setting up and executing tasks from research repositories. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 12622–12645, 2024.
  7. 7.Cassano, F., Gouwar, J., Nguyen, D., Nguyen, S., Phipps-Costin, L., Pinckney, D., Yee, M.-H., Zi, Y., Anderson, C. J., Feldman, M. Q., et al. MultiPL-E: a scalable and polyglot approach to benchmarking neural code generation. IEEE Transactions on Software Engineering, 49(7): 3675–3691, 2023.
  8. 8.Cassano, F., Li, L., Sethi, A., Shinn, N., Brennan-Jones, A., Ginesin, J., Berman, E., Chakhnashvili, G., Lozhkov, A., Anderson, C. J., and Guha, A. Can it edit? evaluating the ability of large language models to follow code editing instructions. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=D06yk3DBas.
  9. 9.Chan, C.-M., Chen, W., Su, Y., Yu, J., Xue, W., Zhang, S., Fu, J., and Liu, Z. ChatEval: Towards better LLM-based evaluators through multi-agent debate. In The Twelfth International Conference on Learning Representations.
  10. 10.Chase, H. LangChain. https://github.com/hwchase17/langchain, 2022.
  11. 11.Chen, D., Chen, R., Zhang, S., Wang, Y., Liu, Y., Zhou, H., Zhang, Q., Wan, Y., Zhou, P., and Sun, L. MLLM-as-a-Judge: Assessing multimodal LLM-as-a-Judge with vision-language benchmark. In Forty-first International Conference on Machine Learning.
  12. 12.Chen, D., Huang, Y., Wu, S., Tang, J., Chen, L., Bai, Y., He, Z., Wang, C., Zhou, H., Li, Y., et al. GUI-WORLD: A dataset for GUI-oriented multimodal LLM-based agents. arXiv preprint arXiv:2406.10819, 2024a.
  13. 13.Chen, D., Lin, S., Zeng, M., Zan, D., Wang, J.-G., Cheshkov, A., Sun, J., Yu, H., Dong, G., Aliev, A., et al. CodeR: Issue resolving with multi-agent and task graphs. arXiv preprint arXiv:2406.01304, 2024b.
  14. 14.Chen, G. H., Chen, S., Liu, Z., Jiang, F., and Wang, B. Humans or LLMs as the judge? a study on judgement bias. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 8301–8327, Miami, Florida, USA, November 2024c. Association for Computational Linguistics. URL https://aclanthology.org/2024.emnlp-main.474.
  15. 15.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  16. 16.Cheng, C.-A., Nie, A., and Swaminathan, A. Trace is the new AutoDiff—unlocking efficient optimization of computational workflows. In ICML 2024 Workshop on Automated Reinforcement Learning: Exploring Meta-Learning, AutoML, and LLMs.
  17. 17.Clemen, R. T. Combining forecasts: A review and annotated bibliography. International journal of forecasting, 5(4): 559–583, 1989.
  18. 18.Cortes, C. Support-vector networks. Machine Learning, 1995.
  19. 19.Dong, Y. R., Hu, T., and Collier, N. Can LLM be a personalized judge? In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 10126–10141, Miami, Florida, USA, November 2024. Association for Computational Linguistics. URL https://aclanthology.org/2024.findings-emnlp.592.
  20. 20.Du, Z., Qian, C., Liu, W., Xie, Z., Wang, Y., Dang, Y., Chen, W., and Yang, C. Multi-agent software development through cross-team collaboration. arXiv preprint arXiv:2406.08979, 2024.
  21. 21.Fayyad, U., Piatetsky-Shapiro, G., and Smyth, P. From data mining to knowledge discovery in databases. AI magazine, 17(3):37–37, 1996.
  22. 22.Fu, J., Ng, S. K., Jiang, Z., and Liu, P. GPTScore: Evaluate as you desire. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 6556–6576, 2024.
  23. 23.Goodhart, C. Monetary relationships: a view from Threadneedle Street. University of Warwick, 1976.
  24. 24.Gravitas, S. Auto-GPT. GitHub repository, 2023.
  25. 25.Grofman, B., Owen, G., and Feld, S. L. Thirteen theorems in search of the truth. Theory and decision, 15(3):261–278, 1983.
  26. 26.Guo, S., Deng, C., Wen, Y., Chen, H., Chang, Y., and Wang, J. DS-Agent: Automated data science by empowering large language models with case-based reasoning. In Forty-first International Conference on Machine Learning.
  27. 27.Haque, M. M. A. FixEval: Execution-based evaluation of program fixes for competitive programming problems. PhD thesis, Virginia Tech, 2023.
  28. 28.Hastie, R. and Kameda, T. The robust beauty of majority rules in group decisions. Psychological review, 112(2): 494, 2005.
  29. 29.Hastie, T., Tibshirani, R., and Friedman, J. H. The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer, 2 edition, 2009. doi: 10.1007/978-0-387-84858-7.
  30. 30.He, X., Zhao, K., and Chu, X. AutoML: A survey of the state-of-the-art. Knowledge-based systems, 212:106622, 2021.
  31. 31.Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., et al. Measuring coding challenge competence with APPS. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  32. 32.Hochreiter, S. Long short-term memory. Neural Computation MIT-Press, 1997.
  33. 33.Hong, S., Lin, Y., Liu, B., Wu, B., Li, D., Chen, J., Zhang, J., Wang, J., Zhang, L., Zhuge, M., et al. Data interpreter: An LLM agent for data science. arXiv preprint arXiv:2402.18679, 2024a.
  34. 34.Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S. K. S., Lin, Z., et al. MetaGPT: Meta programming for a multi-agent collaborative framework. In The Twelfth International Conference on Learning Representations, 2024b.
  35. 35.Hu, S., Lu, C., and Clune, J. Automated design of agentic systems. arXiv preprint arXiv:2408.08435, 2024.
  36. 36.Huang, D., Bu, Q., Zhang, J. M., Luck, M., and Cui, H. AgentCoder: Multi-agent-based code generation with iterative testing and optimisation. arXiv preprint arXiv:2312.13010, 2023.
  37. 37.Huang, Q., Vora, J., Liang, P., and Leskovec, J. MLAgentBench: Evaluating language agents on machine learning experimentation. In Forty-first International Conference on Machine Learning, 2024.
  38. 38.Jain, N., Han, K., Gu, A., Li, W.-D., Yan, F., Zhang, T., Wang, S., Solar-Lezama, A., Sen, K., and Stoica, I. LiveCodeBench: Holistic and contamination free evaluation of large language models for code. arXiv preprint arXiv:2403.07974, 2024.
  39. 39.Jansen, P., Cotê, M.-A., Khot, T., Bransom, E., Mishra, B. D., Majumder, B. P., Tafjord, O., and Clark, P. DISCOVERYWORLD: A virtual environment for developing and evaluating automated scientific discovery agents. arXiv preprint arXiv:2406.06769, 2024.
  40. 40.Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K. R. SWE-bench: Can language models resolve real-world GitHub issues? In The Twelfth International Conference on Learning Representations.
  41. 41.Jin, H., Huang, L., Cai, H., Yan, J., Li, B., and Chen, H. From LLMs to LLM-based agents for software engineering: A survey of current, challenges and future. arXiv preprint arXiv:2408.02479, 2024.
  42. 42.Khattab, O., Singhvi, A., Maheshwari, P., Zhang, Z., Santhanam, K., A, S. V., Haq, S., Sharma, A., Joshi, T. T., Moazam, H., Miller, H., Zaharia, M., and Potts, C. DSPy: Compiling declarative language model calls into state-of-the-art pipelines. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=sY5N0zY5Od.
  43. 43.Lai, Y., Li, C., Wang, Y., Zhang, T., Zhong, R., Zettlemoyer, L., Yih, W.-t., Fried, D., Wang, S., and Yu, T. DS-1000: A natural and reliable benchmark for data science code generation. In International Conference on Machine Learning, pp. 18319–18345. PMLR, 2023.
  44. 44.LangChain-AI. LangGraph. https://github.com/langchain-ai/langgraph, 2024.
  45. 45.Larrick, R. P. and Soll, J. B. Intuitions about combining opinions: Misappreciation of the averaging principle. Management science, 52(1):111–127, 2006.
  46. 46.Levenshtein, V. Binary codes capable of correcting deletions, insertions, and reversals. Proceedings of the Soviet physics doklady, 1966.
  47. 47.Li, G., Hammoud, H., Itani, H., Khizbullin, D., and Ghanem, B. CAMEL: Communicative agents for ”mind” exploration of large language model society. Advances in Neural Information Processing Systems, 36:51991–52008, 2023.
  48. 48.Li, R., Patel, T., Wang, Q., and Du, X. MLR-Copilot: Autonomous machine learning research based on large language models agents. arXiv preprint arXiv:2408.14033, 2024.
  49. 49.Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al. Competition-level code generation with AlphaCode. Science, 378(6624):1092–1097, 2022.
  50. 50.Liu, J., Wang, K., Chen, Y., Peng, X., Chen, Z., Zhang, L., and Lou, Y. Large language model-based agents for software engineering: A survey. arXiv preprint arXiv:2409.02977, 2024.
  51. 51.Liu, T., Xu, C., and McAuley, J. RepoBench: Benchmarking repository-level code auto-completion systems. In The Twelfth International Conference on Learning Representations, a.
  52. 52.Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., Yang, K., et al. AgentBench: Evaluating llms as agents. In The Twelfth International Conference on Learning Representations, b.
  53. 53.Liu, Y., Tang, X., Cai, Z., Lu, J., Zhang, Y., Shao, Y., Deng, Z., Hu, H., Yang, Z., An, K., et al. ML-Bench: Large language models leverage open-source libraries for machine learning tasks. arXiv preprint arXiv:2311.09835, 2023.
  54. 54.Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., and Ha, D. The AI scientist: Towards fully automated open-ended scientific discovery. arXiv preprint arXiv:2408.06292, 2024a.
  55. 55.Lu, Y., Yang, X., Li, X., Wang, X. E., and Wang, W. Y. LLMScore: Unveiling the power of large language models in text-to-image synthesis evaluation. Advances in Neural Information Processing Systems, 36, 2024b.
  56. 56.Mündler, N., Müller, M. N., He, J., and Vechev, M. Code agents are state of the art software testers. arXiv preprint arXiv:2406.12952, 2024.
  57. 57.OpenAI. GPT-4 technical report, 2023.
  58. 58.Park, J. Constructive multiple-choice testing system. British Journal of Educational Technology, 41(6):1054–1064, 2010. doi: https://doi.org/10.1111/j.1467-8535.2010.01058.x. URL https://bera-journals.onlinelibrary.wiley.com/doi/abs/10.1111/j.1467-8535.2010.01058.x.
  59. 59.Phan, H. N., Nguyen, P. X., and Bui, N. D. Hyperagent: Generalist software engineering agents to solve coding tasks at scale. arXiv preprint arXiv:2409.16299, 2024.
  60. 60.Pythagora.io. GPT-Pilot: Your ai copilot for software development. https://github.com/Pythagora-io/gpt-pilot, 2023. URL https://github.com/Pythagora-io/gpt-pilot. GitHub repository.
  61. 61.Qian, C., Liu, W., Liu, H., Chen, N., Dang, Y., Li, J., Yang, C., Chen, W., Su, Y., Cong, X., Xu, J., Li, D., Liu, Z., and Sun, M. ChatDev: Communicative agents for software development. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15174–15186, Bangkok, Thailand, August 2024a. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.810. URL https://aclanthology.org/2024.acl-long.810.
  62. 62.Qian, C., Xie, Z., Wang, Y., Liu, W., Dang, Y., Du, Z., Chen, W., Yang, C., Liu, Z., and Sun, M. Scaling large-language-model-based multi-agent collaboration. arXiv preprint arXiv:2406.07155, 2024b.
  63. 63.Qiao, B., Li, L., Zhang, X., He, S., Kang, Y., Zhang, C., Yang, F., Dong, H., Zhang, J., Wang, L., et al. Taskweaver: A code-first agent framework. arXiv preprint arXiv:2311.17541, 2023.
  64. 64.Raina, V., Liusie, A., and Gales, M. Is LLM-as-a-judge robust? investigating universal adversarial attacks on zero-shot LLM assessment. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 7499–7517, Miami, Florida, USA, November 2024. Association for Computational Linguistics. URL https://aclanthology.org/2024.emnlp-main.427.
  65. 65.Reimers, N. and Gurevych, I. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Inui, K., Jiang, J., Ng, V., and Wan, X. (eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 3982–3992, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1410. URL https://aclanthology.org/D19-1410.
  66. 66.Robertson, S., Zaragoza, H., et al. The probabilistic relevance framework: BM25 and beyond. Foundations and Trends in Information Retrieval, 3(4):333–389, 2009.
  67. 67.Shi, L., Ma, W., and Vosoughi, S. Judging the judges: A systematic investigation of position bias in pairwise comparative assessments by LLMs. arXiv preprint arXiv:2406.07791, 2024.
  68. 68.Song, L., Liu, J., Zhang, J., Zhang, S., Luo, A., Wang, S., Wu, Q., and Wang, C. Adaptive in-conversation team building for language model agents. arXiv preprint arXiv:2405.19425, 2024.
  69. 69.Tan, W., Ding, Z., Zhang, W., Li, B., Zhou, B., Yue, J., Xia, H., Jiang, J., Zheng, L., Xu, X., et al. Towards general computer control: A multimodal agent for Red Dead Redemption II as a case study. In ICLR 2024 Workshop on Large Language Model (LLM) Agents.
  70. 70.Tao, W., Zhou, Y., Zhang, W., and Cheng, Y. MAGIS: LLM-Based multi-agent framework for GitHub issue resolution. arXiv preprint arXiv:2403.17927, 2024.
  71. 71.Thakur, A. S., Choudhary, K., Ramayapally, V. S., Vaidyanathan, S., and Hupkes, D. Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges. arXiv preprint arXiv:2406.12624, 2024.
  72. 72.Tian, R., Ye, Y., Qin, Y., Cong, X., Lin, Y., Pan, Y., Wu, Y., Haotian, H., Weichuan, L., Liu, Z., and Sun, M. DebugBench: Evaluating debugging capability of large language models. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Findings of the Association for Computational Linguistics: ACL 2024, pp. 4173–4198, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-acl.247. URL https://aclanthology.org/2024.findings-acl.247.
  73. 73.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  74. 74.Tufano, M., Agarwal, A., Jang, J., Moghaddam, R. Z., and Sundaresan, N. AutoDev: Automated AI-driven development. arXiv preprint arXiv:2403.08299, 2024.
  75. 75.Wang, J., Xu, H., Jia, H., Zhang, X., Yan, M., Shen, W., Zhang, J., Huang, F., and Sang, J. Mobile-agent-v2: Mobile device operation assistant with effective navigation via multi-agent collaboration. arXiv preprint arXiv:2406.01014, 2024a.
  76. 76.Wang, X., Chen, Y., Yuan, L., Zhang, Y., Li, Y., Peng, H., and Ji, H. Executable code actions elicit better LLM agents. In ICLR 2024 Workshop on Large Language Model (LLM) Agents.
  77. 77.Wang, X., Li, B., Song, Y., Xu, F. F., Tang, X., Zhuge, M., Pan, J., Song, Y., Li, B., Singh, J., et al. OpenDevin: An open platform for ai software developers as generalist agents. arXiv preprint arXiv:2407.16741, 2024b.
  78. 78.Wirth, R. and Hipp, J. CRISP-DM: Towards a standard process model for data mining. In Proceedings of the 4th international conference on the practical applications of knowledge discovery and data mining, volume 1, pp. 29–39. Manchester, 2000.
  79. 79.Wooldridge, M. Intelligent agents. Multiagent systems: A modern approach to distributed artificial intelligence, 1: 27–73, 1999.
  80. 80.Wu, Q., Bansal, G., Zhang, J., Wu, Y., Zhang, S., Zhu, E., Li, B., Jiang, L., Zhang, X., and Wang, C. AutoGen: Enabling next-gen llm applications via multi-agent conversation framework. arXiv preprint arXiv:2308.08155, 2023.
  81. 81.Wu, Y., Yue, T., Zhang, S., Wang, C., and Wu, Q. StateFlow: Enhancing LLM task-solving through state-driven workflows. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=3nTbuygoop.
  82. 82.Xia, C. S., Deng, Y., Dunn, S., and Zhang, L. Agentless: Demystifying LLM-based software engineering agents. arXiv preprint arXiv:2407.01489, 2024.
  83. 83.Xu, T., Chen, L., Wu, D.-J., Chen, Y., Zhang, Z., Yao, X., Xie, Z., Chen, Y., Liu, S., Qian, B., et al. CRAB: Cross-environment agent benchmark for multimodal language model agents. arXiv preprint arXiv:2407.01511, 2024.
  84. 84.Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K. R., and Press, O. SWE-agent: Agent-computer interfaces enable automated software engineering. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024a.
  85. 85.Yang, Z., Liu, J., Han, Y., Chen, X., Huang, Z., Fu, B., and Yu, G. AppAgent: Multimodal agents as smartphone users. arXiv preprint arXiv:2312.13771, 2023.
  86. 86.Yang, Z., Zhou, Z., Wang, S., Cong, X., Han, X., Yan, Y., Liu, Z., Tan, Z., Liu, P., Yu, D., Liu, Z., Shi, X., and Sun, M. MatPlotAgent: Method and evaluation for LLM-based agentic scientific data visualization. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Findings of the Association for Computational Linguistics: ACL 2024, pp. 11789–11804, Bangkok, Thailand, August 2024b. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-acl.701. URL https://aclanthology.org/2024.findings-acl.701.
  87. 87.Zhang, F., Chen, B., Zhang, Y., Keung, J., Liu, J., Zan, D., Mao, Y., Lou, J.-G., and Chen, W. RepoCoder: Repository-level code completion through iterative retrieval and generation. In Bouamor, H., Pino, J., and Bali, K. (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 2471–2484, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.151. URL https://aclanthology.org/2023.emnlp-main.151.
  88. 88.Zhao, R., Zhang, W., Chia, Y. K., Zhao, D., and Bing, L. Auto arena of LLMs: Automating LLM evaluations with agent peer-battles and committee discussions. arXiv preprint arXiv:2405.20267, 2024.
  89. 89.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. Judging LLM-as-a-Judge with MT-bench and chatbot arena. Advances in Neural Information Processing Systems, 36, 2024.
  90. 90.Zhou, W., Jiang, Y. E., Li, L., Wu, J., Wang, T., Qiu, S., Zhang, J., Chen, J., Wu, R., Wang, S., et al. Agents: An open-source framework for autonomous language agents. arXiv preprint arXiv:2309.07870, 2023.
  91. 91.Zhou, W., Ou, Y., Ding, S., Li, L., Wu, J., Wang, T., Chen, J., Wang, S., Xu, X., Zhang, N., et al. Symbolic learning enables self-evolving agents. arXiv preprint arXiv:2406.18532, 2024.
  92. 92.Zhuge, M., Wang, W., Kirsch, L., Faccio, F., Khizbullin, D., and Schmidhuber, J. GPTSwarm: Language agents as optimizable graphs. In Forty-first International Conference on Machine Learning.
  93. 93.Zhuge, M., Liu, H., Faccio, F., Ashley, D. R., Csordás, R., Gopalakrishnan, A., Hamdi, A., Hammoud, H. A. A. K., Herrmann, V., Irie, K., et al. Mindstorms in natural language-based societies of mind. arXiv preprint arXiv:2305.17066, 2023.

Citation

MLA
Zhuge, M., et al. “Agent-as-a-Judge: Evaluate Agents with Agents”. arXiv, 2024, http://arxiv.org/abs/2410.10934v2.
APA
Zhuge, M., Zhao, C., Ashley, D., Wang, W., Khizbullin, D., Xiong, Y., Liu, Z., Chang, E., Krishnamoorthi, R., Tian, Y., Shi, Y., Chandra, V., & Schmidhuber, J. (2024). Agent-as-a-Judge: Evaluate Agents with Agents. arXiv. http://arxiv.org/abs/2410.10934v2
Chicago
Zhuge, M., C. Zhao, D. Ashley, et al. 2024. “Agent-as-a-Judge: Evaluate Agents with Agents”. arXiv. http://arxiv.org/abs/2410.10934v2.
Harvard
Zhuge, M. et al. (2024) “Agent-as-a-Judge: Evaluate Agents with Agents”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.10934v2.
Vancouver
1. Zhuge M, Zhao C, Ashley D, et al (2024) Agent-as-a-Judge: Evaluate Agents with Agents. arXiv

BibTeX

@article{zhuge2024agent,
  title = {Agent-as-a-Judge: Evaluate Agents with Agents},
  author = {Zhuge, Mingchen and Zhao, Changsheng and Ashley, Dylan and Wang, Wenyi and Khizbullin, Dmitrii and Xiong, Yunyang and Liu, Zechun and Chang, Ernie and Krishnamoorthi, Raghuraman and Tian, Yuandong and Shi, Yangyang and Chandra, Vikas and Schmidhuber, Jürgen},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.10934v2},
  eprint = {2410.10934}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/