MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Qian HuangJian VoraPercy LiangJure Leskovec
Presents MLAgentBench, a novel evaluation suite across 13 diverse research tasks to assess how effectively language model agents can autonomously read code, run experiments, and iterate on machine learning pipelines.
Machine learning experimentation is an iterative and resource-intensive process that requires deep technical expertise, extensive trial-and-error, and careful result interpretation. While automating this workflow has long been a goal to broaden access and accelerate discovery, prior methods have struggled with end-to-end execution. The article introduces MLAgentBench, a standardized benchmark framework designed to evaluate whether autonomous language model agents can effectively design, execute, and iterate on machine learning experiments without human intervention.
The benchmark consists of 13 diverse tasks across images, text, tabular data, time series, and graphs, encompassing canonical machine learning benchmarks, competitive challenges, and code optimization problems. Agents interact with a realistic workspace environment by reading files, modifying code, executing Python scripts, and producing test set predictions. The evaluation framework tests multiple leading language models—including GPT-4, GPT-4-turbo, Claude v1.0, Claude v2.1, Claude v3 Opus, Gemini Pro, and Mixtral—measuring task success (defined as at least a 10% improvement over starter baselines) and computational efficiency across multiple runs.
The findings show that Claude v3 Opus achieved the highest overall performance with a 37.5% average success rate, outperforming existing agent frameworks such as AutoGPT and LangChain. Performance varied widely based on task maturity: agents reached up to a 100% success rate on well-established classic datasets but fell to 0% to 25% on newer or complex research problems. GPT-4-turbo proved the most efficient, consuming approximately 51% fewer tokens than the benchmark average, though its lower success rate increased the expected cost per completed task to $231.
Qualitative analysis indicates that while structured reflection, planning, and fact-checking mechanisms improve progress tracking and reduce errors, persistent failure modes remain. These include model hallucinations, poor long-term planning, debugging loops, and submission formatting errors. Additionally, executing more iterative steps frequently degraded rather than improved agent performance over longer interaction horizons.
These results demonstrate that fully autonomous machine learning experimentation is feasible for straightforward workflows but remains unreliable for complex and novel domains. Stakeholders should treat current autonomous agents as assistive tools under close human supervision rather than fully independent systems. Future efforts should focus on conducting user studies to improve human-AI collaboration, enhancing long-term planning capabilities, mitigating hallucinations, and expanding evaluation suites to cover broader scientific and creative engineering tasks.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). This comprehensive survey outlines the foundational architectures, planning methods, and action spaces of LLM-based autonomous agents that underpin the agent design in MLAgentBench.
- Paper: The Rise and Potential of Large Language Model Based Agents: A Survey, Zhiheng Xi et al. (2023). It provides essential background on how large language models function as interactive agents perceiving environments, executing tools, and planning multi-step actions.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). It establishes key evaluation paradigms and execution-based interactive benchmark environments for autonomous language agents that preceded ML-specific agent benchmarks.
- Paper: ScienceWorld: Is your Agent Smarter than a 5th Grader?, Ruoyao Wang et al. (2022). It introduces the foundational concept of testing whether language agents can perform goal-oriented scientific experimentation in interactive, simulated environments.
- Paper: AutoML: A Survey of the State-of-the-Art, Xin He et al. (2019). It reviews classic AutoML stages—such as data preparation, model search, and hyperparameter tuning—that MLAgentBench tasks language model agents to perform autonomously.
- Paper: AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoML, Patara Trirat et al. (2025). It builds beyond single-agent ML experimentation by developing a specialized multi-agent LLM framework to automate full-pipeline machine learning across diverse data modalities.
- Paper: ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models, Jinheon Baek et al. (2025). It extends autonomous machine learning and scientific experimentation upstream by focusing on iterative scientific research idea generation and experimental design.
- Paper: Accelerating Scientific Research with Gemini in the Real-World, Samuel Schmidgall et al. (2026). It advances agentic ML and scientific discovery from benchmark simulations to real-world, cross-disciplinary computational and physical laboratory execution.
- Paper: Automated Design of Agentic Systems, Shengran Hu et al. (2025). It generalizes autonomous agent experimentation to meta-level automated design, enabling agents to iteratively invent, code, and optimize entire agentic architectures.
- Paper: SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, John Yang et al. (2024). It complements file- and code-level ML agent workflows by introducing tailored Agent-Computer Interfaces that improve how language agents navigate repositories and edit code.
- Paper: Agent-as-a-Judge: Evaluate Agents with Agents, Mingchen Zhuge et al. (2025). It addresses the evaluation of complex engineering and development agents across entire interactive trajectories by employing tool-augmented judge agents.
- Paper: Code as Agent Harness, Xuying Ning et al. (2026). It formalizes and systematizes the paradigm of using executable code and harness environments as the primary interface for autonomous agent reasoning and experimentation.
