MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Qian HuangJian VoraPercy LiangJure Leskovec

article2024ICML198 citations

Presents MLAgentBench, a novel evaluation suite across 13 diverse research tasks to assess how effectively language model agents can autonomously read code, run experiments, and iterate on machine learning pipelines.

Listen

Machine learning experimentation is an iterative and resource-intensive process that requires deep technical expertise, extensive trial-and-error, and careful result interpretation. While automating this workflow has long been a goal to broaden access and accelerate discovery, prior methods have struggled with end-to-end execution. The article introduces MLAgentBench, a standardized benchmark framework designed to evaluate whether autonomous language model agents can effectively design, execute, and iterate on machine learning experiments without human intervention.

The benchmark consists of 13 diverse tasks across images, text, tabular data, time series, and graphs, encompassing canonical machine learning benchmarks, competitive challenges, and code optimization problems. Agents interact with a realistic workspace environment by reading files, modifying code, executing Python scripts, and producing test set predictions. The evaluation framework tests multiple leading language models—including GPT-4, GPT-4-turbo, Claude v1.0, Claude v2.1, Claude v3 Opus, Gemini Pro, and Mixtral—measuring task success (defined as at least a 10% improvement over starter baselines) and computational efficiency across multiple runs.

The findings show that Claude v3 Opus achieved the highest overall performance with a 37.5% average success rate, outperforming existing agent frameworks such as AutoGPT and LangChain. Performance varied widely based on task maturity: agents reached up to a 100% success rate on well-established classic datasets but fell to 0% to 25% on newer or complex research problems. GPT-4-turbo proved the most efficient, consuming approximately 51% fewer tokens than the benchmark average, though its lower success rate increased the expected cost per completed task to $231.

Qualitative analysis indicates that while structured reflection, planning, and fact-checking mechanisms improve progress tracking and reduce errors, persistent failure modes remain. These include model hallucinations, poor long-term planning, debugging loops, and submission formatting errors. Additionally, executing more iterative steps frequently degraded rather than improved agent performance over longer interaction horizons.

These results demonstrate that fully autonomous machine learning experimentation is feasible for straightforward workflows but remains unreliable for complex and novel domains. Stakeholders should treat current autonomous agents as assistive tools under close human supervision rather than fully independent systems. Future efforts should focus on conducting user studies to improve human-AI collaboration, enhancing long-term planning capabilities, mitigating hallucinations, and expanding evaluation suites to cover broader scientific and creative engineering tasks.

Cover for MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Abstract

A central aspect of machine learning research is experimentation, the process of designing and running experiments, analyzing the results, and iterating towards some positive outcome (e.g., improving accuracy). Could agents driven by powerful language models perform machine learning experimentation effectively? To answer this question, we introduce MLAgentBench, a suite of 13 tasks ranging from improving model performance on CIFAR-10 to recent research problems like BabyLM. For each task, an agent can perform actions like reading/writing files, executing code, and inspecting outputs. We then construct an agent that can perform ML experimentation based on ReAct framework. We benchmark agents based on Claude v1.0, Claude v2.1, Claude v3 Opus, GPT-4, GPT-4-turbo, Gemini-Pro, and Mixtral and find that a Claude v3 Opus agent is the best in terms of success rate. It can build compelling ML models over many tasks in MLAgentBench with 37.5% average success rate. Our agents also display highly interpretable plans and actions. However, the success rates vary considerably; they span from 100% on well-established older datasets to as low as 0% on recent Kaggle challenges created potentially after the underlying LM was trained. Finally, we identify several key challenges for LM-based agents such as long-term planning and reducing hallucination.

Table of Contents

  • 1. Introduction
  • 2. MLAgentBench: Benchmarking ML experimentation
  • 2.1. Task Specification
  • 2.2. General Environment
  • 2.2.1. ACTIONS
  • 2.3. Evaluation
  • 2.4. Tasks
  • 3. Our LM-based Agent
  • 3.1. Thinking before Acting
  • 4. Experiments
  • 4.1. Competence in Accomplishing The Objectives
  • 4.2. Research Process
  • 4.3. Efficiency
  • 5. Related Work
  • 5.1. Language Agents
  • 5.2. Language Models for AutoML
  • 5.3. AI for Automating Scientific Discovery
  • 6. Conclusion
  • Impact Statement
  • References
  • A. Agent Framework Comparison
  • B. Benchmark Details
  • C. Quantitative Analysis
  • D. Qualitative Examples
  • D.1. Research Plan and Status
  • D.2. Fact Check
  • D.3. Problem Misspecification
  • E. Efficiency
  • F. Full Example

Knowls

  1. Knowl 1 — MLAgentBench Autonomous Experimentation Framework

    model/method

    MLAgentBench is a general benchmark and execution framework designed to evaluate language model (LM)-based agents on end-to-end autonomous machine learning experimentation. Each experimentation task is specified by a tuple comprising a textual task description, a set of starter files (including starter code and datasets), and an automated evaluator.

    The interaction between an agent and the task environment proceeds in discrete time steps t=1,…,Tt = 1, \dots, T:

    1. Act: The agent inspects its memory mt−1=(o<t,r<t)m_{t-1} = (o_{<t}, r_{<t}) and current workspace state st−1s_{t-1} (the collection of files and directories in the working workspace) to generate a structured rationale rtr_t and an action ata_t: rt,at=Agent(st−1,mt−1)r_t, a_t = \text{Agent}(s_{t-1}, m_{t-1})

    2. Execution: The environment executes ata_t on the workspace st−1s_{t-1}, producing an updated workspace state sts_t and returning an observation string oto_t: st,ot=Env(st−1,at)s_t, o_t = \text{Env}(s_{t-1}, a_t)

    3. Update: The agent updates its memory history with the new rationale, action, and observation: mt=Update(mt−1,at,rt,ot)m_t = \text{Update}(m_{t-1}, a_t, r_t, o_t)

    The agent interacts iteratively until it invokes the submission action or hits execution budget limits (e.g., maximum action steps or timeout). Upon termination, the evaluator calculates a primary performance metric score on the final workspace artifact (such as a submission.csv file or model checkpoint). An experiment is defined as a success if the final metric improves by at least 10%10\% over the baseline score implemented in the starter files.

  2. Knowl 2 — Action Space and Compound Tool Calls in MLAgentBench

    definition

    The MLAgentBench environment provides agents with a standard interface of primitive file-system operations and execution actions, augmented with specialized compound actions combining file operations and modular LM sub-calls:

    • Primitive File and Execution Actions:

      • List Files: Lists files and folders within a specified directory path.
      • Read File: Returns the complete content of a specified file.
      • Write File: Overwrites a specified file with provided text content.
      • Append File: Appends provided text to a specified file.
      • Copy File: Copies a file from a source path to a destination path.
      • Inspect Script Lines: Returns file contents within a specified start and end line range (up to 100 lines).
      • Undo Edit Script: Restores a script to its state prior to the most recent modification.
      • Execute Script: Executes a specified Python script and captures stdout, stderr, and runtime exit codes.
      • Final Answer: Terminates the environment session and submits the current workspace for final evaluation.
    • Modular Compound Actions:

      • Understand File: Takes a file name and an inquiry query (e.g., asking about the model architecture), prompts a secondary LM call to read and summarize relevant sections of the file, and returns retrieved lines and summaries.
      • Edit Script: Takes a target file name, natural language edit instructions (e.g., modifying learning rate), and an output file name; it prompts an LM to apply the diff/edit and saves the resulting script.
      • Edit Script Segment: Operates identically to Edit Script but restricts edits to a specified line range, assisting in localized edits within large codebases.
  3. Knowl 3 — Prompt-Engineered Agent Architecture with Structured Rationale and Fact Checking

    model/method

    The baseline LM agent in MLAgentBench prompts an underlying language model to generate a structured rationale rtr_t prior to producing the next action ata_t in JSON format. The prompt context ptp_t provided at step tt contains the available action definitions, task description, formatting requirements, and the interaction history of the last 3 steps: (rt−3,at−3,ot−3,rt−2,at−2,ot−2,rt−1,at−1,ot−1)(r_{t-3}, a_{t-3}, o_{t-3}, r_{t-2}, a_{t-2}, o_{t-2}, r_{t-1}, a_{t-1}, o_{t-1}).

    The output string rtr_t is strictly constrained to include five sequential components before action generation:

    1. Reflection: Analysis of the previous observation ot−1o_{t-1}, diagnosing errors and code execution outputs.
    2. Research Plan and Status: An explicit, updated multi-step research plan detailing past milestones, active experiments, and subsequent goals. Newly updated status lines must be enclosed in double asterisks (**...**).
    3. Fact Check: A verification step forcing the LM to evaluate every claim added to Research Plan and Status, categorizing statements as either objectively confirmed by preceding execution output or unverified hypotheses. This prevents hallucinated progress (e.g., assuming code modifications improved metrics prior to script execution).
    4. Thought: Immediate step-level reasoning justifying the specific upcoming tool invocation.
    5. Action and Action Input: The exact tool name and its corresponding parameters formatted as a valid JSON object.
  4. Knowl 4 — MLAgentBench Benchmark Task Suite

    data/table

    MLAgentBench comprises 13 diverse ML tasks spanning 5 distinct problem categories, multiple modalities (images, text, graphs, tabular data, time series), and various ML frameworks (PyTorch, TensorFlow, JAX, Keras). The runtime for executing scripts across tasks is designed to be inexpensive (on the order of minutes).

    Category Task Type Modality Dataset Name Metric
    Canonical Tasks Classification Image CIFAR-10 Classification accuracy
    Canonical Tasks Classification Text imdb Classification accuracy
    Canonical Tasks Node Classification Graph ogbn-arxiv Classification accuracy
    Classic Kaggle Regression Tabular house-price Mean absolute error
    Classic Kaggle Classification Tabular spaceship-titanic Classification accuracy
    Kaggle Challenges Regression Time Series parkinsons-disease SMAPE score
    Kaggle Challenges Classification Image fathomnet MAP@20
    Kaggle Challenges Regression Text feedback MCRMSE
    Kaggle Challenges Segmentation Images identify-contrails Dice coefficient
    Recent Research Node Regression Graph CLRS Mean square error
    Recent Research Language Modeling Text BabyLM Perplexity
    Code Improvement Improve speed Text llama-inference Wall Clock Time
    Code Improvement Improve speed Image vectorization Wall Clock Time
    • Canonical Tasks & Classic Kaggle: Well-established benchmarks; for tasks like imdb, house-price, and spaceship-titanic, models must be built from scratch without starter baseline implementations.
    • Kaggle Challenges: Open competitions launched between August 2022 and May 2023 to test agent performance on realistic, out-of-distribution problems released around or after LM training cutoffs.
    • Recent Research: Ongoing research challenges (CLRS algorithmic reasoning and BabyLM 10M-word sample-efficient pretraining).
    • Code Improvement: Tasks focused on reducing execution runtime rather than optimizing prediction accuracy.
  5. Knowl 5 — Success Rate Across Language Models on MLAgentBench

    data/table

    Language agents built on Claude v3 Opus, Claude v2.1, Claude v1.0, GPT-4 (0613), GPT-4-turbo (0125), Gemini Pro, and Mixtral (Instruct-v0.1) were evaluated over 8 independent runs per task. A trial is deemed successful if the agent achieves ≥10%\ge 10\% metric improvement over the starter code baseline. For tasks without starter baseline models (house-price, spaceship-titanic, imdb, fathomnet), improvement is measured relative to trivial baseline predictions.

    Task GPT-4 GPT-4-turbo Claude v1.0 Claude v2.1 Claude v3 Opus Gemini Pro Mixtral Baseline
    cifar10 25.0 25.0 12.5 25.0 62.5 12.5 25.0 0.0
    imdb 25.0 12.5 0.0 0.0 25.0 0.0 0.0 0.0
    ogbn-arxiv 87.5 62.5 37.5 62.5 87.5 37.5 0.0 0.0
    house-price 12.5 87.5 75.0 87.5 100.0 100.0 12.5 0.0
    spaceship-titanic 12.5 50.0 12.5 75.0 100.0 87.5 0.0 0.0
    parkinsons-disease 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
    fathomnet 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
    feedback 12.5 37.5 0.0 37.5 87.5 0.0 0.0 0.0
    identify-contrails 25.0 62.5 12.5 25.0 0.0 0.0 0.0 40.0
    llama-inference 0.0 0.0 12.5 25.0 0.0 0.0 12.5 0.0
    vectorization 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
    CLRS 50.0 0.0 50.0 0.0 25.0 0.0 0.0 42.9
    BabyLM 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
    Average 19.2 26.0 16.3 26.0 37.5 18.3 3.8 10.4

    Claude v3 Opus achieved the highest overall average success rate (37.5%37.5\%). Across all agents, performance is high on canonical and classic tabular tasks (e.g., 100%100\% on house-price and spaceship-titanic for Claude v3 Opus) but drops to 0%0\% on complex Kaggle competitions (parkinsons-disease, fathomnet) and recent research problems (BabyLM).

  6. Knowl 6 — Comparison of Agent Frameworks: MLAgentBench Agent vs. AutoGPT and LangChain

    data/table

    The structured MLAgentBench agent design was benchmarked against adaptations of two prominent agent frameworks: AutoGPT and LangChain (using zero-shot ReAct description, lacking explicit research plan tracking and fact-checking entries). Both were evaluated using GPT-4-turbo and Claude v3 Opus over 8 runs per task.

    GPT-4-turbo Claude v3 Opus
    Task Ours AutoGPT LangChain Ours AutoGPT LangChain
    cifar10 25.0 0.0 0.0 62.5 0.0 87.5
    imdb 12.5 0.0 0.0 25.0 0.0 25.0
    ogbn-arxiv 62.5 0.0 12.5 87.5 12.5 62.5
    house-price 87.5 25.0 0.0 100.0 62.5 100.0
    spaceship-titanic 50.0 12.5 0.0 100.0 100.0 75.0
    parkinsons-disease 0.0 0.0 0.0 0.0 0.0 0.0
    fathomnet 0.0 0.0 0.0 0.0 0.0 0.0
    feedback 37.5 0.0 0.0 87.5 0.0 50.0
    identify-contrails 62.5 0.0 0.0 0.0 0.0 25.0
    llama-inference 0.0 0.0 0.0 0.0 0.0 0.0
    vectorization 0.0 0.0 0.0 0.0 0.0 12.5
    CLRS 0.0 0.0 0.0 25.0 0.0 0.0
    BabyLM 0.0 0.0 0.0 0.0 0.0 0.0
    Average 26.0 2.9 1.0 37.5 13.5 33.7

    The proposed structured agent outperforms AutoGPT (2.9%2.9\% vs 26.0%26.0\% with GPT-4-turbo; 13.5%13.5\% vs 37.5%37.5\% with Claude v3 Opus) and LangChain (1.0%1.0\% with GPT-4-turbo; 33.7%33.7\% with Claude v3 Opus). LangChain with Claude v3 Opus remained competitive primarily due to its simplicity, avoiding submission format modifications.

  7. Knowl 7 — Average Metric Percentage Improvement Across Language Models

    data/table

    The average percentage improvement of the task performance metric over the starter baseline was measured across all runs yielding a valid final submission file at termination.

    Task GPT-4 GPT-4-turbo Claude v1.0 Claude v2.1 Claude v3 Opus Gemini Pro Mixtral Baseline
    cifar10 9.2 5.3 -3.1 5.1 18.5 -36.4 6.5 0.0
    imdb 86.4 86.2 0.0 0.0 82.0 0.0 0.0 0.0
    ogbn-arxiv 48.9 38.6 10.7 19.8 49.5 7.3 -2.2 0.0
    house-price 100.0 100.0 100.0 100.0 100.0 100.0 100.0 0.0
    spaceship-titanic 45.8 45.0 48.4 40.5 44.8 45.4 0.0 0.0
    parkinsons-disease -0.0 0.0 -0.1 -13.3 -0.1 -0.2 -0.1 0.0
    fathomnet 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
    feedback 78.0 68.1 0.0 32.8 74.5 0.0 0.0 0.0
    identify-contrails 143.3 114.9 -48.9 24.1 0.0 -98.8 0.0 0.0
    llama-inference -1.3 -0.3 8.1 18.5 0.8 -23.0 10.7 0.0
    vectorization 0.0 -6.8 0.0 -10.0 -18.7 -11.9 -3.9 0.0
    CLRS 26.5 -24.2 0.6 -22.1 -11.6 -28.7 -6.6 0.0
    BabyLM 0.0 -0.0 0.0 -0.0 -0.5 0.0 -0.0 0.0
    Average 41.3 32.8 8.9 15.0 26.1 -3.6 8.0 0.0

    While Claude v3 Opus had the highest success rate (frequency of exceeding 10%10\% gain), GPT-4 achieved the highest average percentage improvement (41.3%41.3\%) among valid submissions, heavily influenced by extreme performance gains on identify-contrails (143.3%143.3\%) and imdb (86.4%86.4\%).

  8. Knowl 8 — Error Modes and Horizon Degradation in Autonomous ML Agents

    empirical result

    Analysis of agent traces across benchmark runs reveals five dominant failure modes:

    1. Hallucination: The agent fabricates experimental results or assumes accuracy improvements without executing the modified training code (e.g., asserting a 10%10\% gain when test accuracy actually fell from 51.80%51.80\% to 26.35%26.35\%).
    2. Bad Plan: The agent formulates an ineffective strategy early on (such as dropping essential input features before evaluating predictive relevance) and fails to recover in later steps.
    3. Response Format Error: The model produces malformed JSON or fails to adhere to the required tool formatting syntax.
    4. Submission Format Error: The agent modifies the columns, structure, or encoding of submission.csv, rendering the file unparseable by the automated evaluator despite good predictions.
    5. Small Improvement: The agent successfully implements code and achieves metric improvement, but the magnitude is under the 10%10\% success threshold.

    Evaluating model accuracy across all intermediate execution steps indicates that running longer action sequences degrades average performance metrics for most models (GPT-4-turbo, Claude v1.0, Claude v2.1, Gemini Pro), whereas Claude v3 Opus is uniquely resilient to degradation over extended horizons.

  9. Knowl 9 — Token Consumption, Execution Efficiency, and Economic Cost Analysis

    empirical result

    Benchmarking efficiency across agents demonstrates substantial variation in token consumption and execution runtime:

    • Token Consumption: GPT-4-turbo is the most token-efficient agent, consuming 51.0%51.0\% fewer tokens than the average agent by completing tasks and submitting early. Conversely, Claude v3 Opus consumes the most tokens and longest wall-clock time due to running longer iterative training experiments and API latency.
    • Pass Cost vs. Expected Completion Cost: Executing the complete 13-task benchmark suite once with GPT-4-turbo consumes approximately 6 million tokens, costing approximately $60 in API fees (a few dollars per individual task run). However, factoring in the 26.0%26.0\% average success rate, the expected economic cost to achieve a successful task outcome is $231, underscoring that agent reliability is the primary economic bottleneck.

Coverage note — Omitted specific verbatim prompt traces from Appendix F, as they are concrete code instances of the structured prompting architecture and CIFAR-10 experiment already detailed in the core knowls.

References

  1. 1.Significant-gravitas/auto-gpt: An experimental open-source attempt to make gpt-4 fully autonomous. https://github.com/Significant-Gravitas/Auto-GPT, 2023.
  2. 2.Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mane, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viegas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X. TensorFlow: Large-scale machine learning on heterogeneous systems, 2015. URL https://www.tensorflow.org/. Software available from tensorflow.org.
  3. 3.Adam-Bourdarios, C., Cowan, G., Germain, C., Guyon, I. M., Kegl, B., and Rousseau, D. How machine learning won the higgs boson challenge. In The European Symposium on Artificial Neural Networks, 2016.
  4. 4.Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Ho, D., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jang, E., Ruano, R. J., Jeffrey, K., Jesmonth, S., Joshi, N. J., Julian, R. C., Kalashnikov, D., Kuang, Y., Lee, K.-H., Levine, S., Lu, Y., Luu, L., Parada, C., Pastor, P., Quiambao, J., Rao, K., Rettinghouse, J., Reyes, D. M., Sermanet, P., Sievers, N., Tan, C., Toshev, A., Vanhoucke, V., Xia, F., Xiao, T., Xu, P., Xu, S., and Yan, M. Do as i can, not as i say: Grounding language in robotic affordances. In Conference on Robot Learning, 2022.
  5. 5.Anil, G. T. G. R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., Silver, D., Petrov, S., Johnson, M., Antonoglou, I., Schrittwieser, J., Glaese, A., Chen, J., Pitler, E., Lillicrap, T. P., Lazaridou, A., Firat, O., Molloy, J., Isard, M., Barham, P., Hennigan, T., Lee, B., Viola, F., Reynolds, M., Xu, Y., Doherty, R., Collins, E., Meyer, C., Rutherford, E., Moreira, E., Ayoub, K. W., Goel, M., Tucker, G., Piqueras, E., Krikun, M., Barr, I., Savinov, N., Danihelka, I., Roelofs, B., White, A., Andreassen, A., von Glehn, T., Yagati, L. N., Kazemi, M., Gonzalez, L., Khalman, M., Sygnowski, J., Frechette, A., Smith, C., Culp, L., Proleev, L., Luan, Y., Chen, X., Lottes, J., Schucher, N., Lebron, F., Rrustemi, A., Clay, N., Crone, P., Kocisky, T., Zhao, J., Perz, B., Yu, D., Howard, H., Bloniarz, A., Rae, J. W., Lu, H., Sifre, L., Maggioni, M., Alcober, F., Garrette, D. H., Barnes, M., Thakoor, S., Austin, J., Barth-Maron, G., Wong, W., Joshi, R., Chaabouni, R., Fatiha, D., Ahuja, A., Liu, R., Li, Y., Cogan, S., Chen, J., Jia, C., Gu, C., Zhang, Q., Grimstad, J., Hartman, A. J., Chadwick, M., Tomar, G. S., Garcia, X., Senter, E., Taropa, E., Pillai, T. S., Devlin, J., Laskin, M., de Las Casas, D., Valter, D., Tao, C., Blanco, L., Badia, A. P., Reitter, D., Chen, M., Brennan, J., Rivera, C., Brin, S., Iqbal, S., de Castro Surita, G., Labanowski, J., Rao, A., Winkler, S., Parisotto, E., Gu, Y., Olszewska, K., Zhang, Y., Addanki, R., Miech, A., Louis, A., Shafey, L. E., Teplyashin, D., Brown, G., Catt, E., Attaluri, N., Balaguer, J., Xiang, J., Wang, P., Ashwood, Z. C., Briukhov, A., Webson, A., Ganapathy, S., Sanghavi, S., Kannan, A., Chang, M.-W., Stjerngren, A., Djolonga, J., Sun, Y., Bapna, A., Aitchison, M., Pejman, P., Michalewski, H., Yu, T., Wang, C., Love, J. C., Ahn, J., Bloxwich, D., Han, K., Humphreys, P., Sellam, T., Bradbury, J., Godbole, V., Samangooei, S., Damoc, B., Kaskasoli, A., Arnold, S. M. R., Vasudevan, V., Agrawal, S., Riesa, J., Lepikhin, D., Tanburn, R., Srinivasan, S., Lim, H., Hodkinson, S., Shyam, P., Ferret, J., Hand, S., Garg, A., Paine, T. L., Li, J., Li, Y., Giang, M., Neitz, A., Abbas, Z., York, S., Reid, M., Cole, E., Chowdhery, A., Das, D., Rogozi’nska, D., Nikolaev, V., Sprechmann, P., Nado, Z., Zilka, L., Prost, F., He, L., Monteiro, M., Mishra, G., Welty, C. A., Newlan, J., Jia, D., Allamanis, M., Hu, C. H., de Liedekerke, R., Gilmer, J., Saroufim, C., Rijhwani, S., Hou, S., Shrivastava, D., Baddepudi, A., Goldin, A., Ozturel, A., Cassirer, A., Xu, Y., Sohn, D., Sachan, D. S., Amplayo, R. K., Swanson, C., Petrova, D., Narayan, S., Guez, A., Brahma, S., Landon, J., Patel, M., Zhao, R., Villela, K., Wang, L., Jia, W., Rahtz, M., Gim’enez, M., Yeung, L., Lin, H., Keeling, J., Georgiev, P., Mincu, D., Wu, B., Haykal, S., Saputro, R., Vodrahalli, K., Qin, J., Cankara, Z., Sharma, A., Fernando, N., Hawkins, W., Neyshabur, B., Kim, S., Hutter, A., Agrawal, P., Castro-Ros, A., van den Driessche, G., Wang, T., Yang, F., yiin Chang, S., Komarek, P., McIlroy, R., Luvci’c, M., Zhang, G., Farhan, W., Sharman, M., Natsev, P., Michel, P., Cheng, Y., Bansal, Y., Qiao, S., Cao, K., Shakeri, S., Butterfield, C., Chung, J., Rubenstein, P. K., Agrawal, S., Mensch, A., Soparkar, K., Lenc, K., Chung, T., Pope, A., Maggiore, L., Kay, J., Jhakra, P., Wang, S., Maynez, J., Phuong, M., Tobin, T., Tacchetti, A., Trebacz, M., Robinson, K., Katariya, Y., Riedel, S., Bailey, P., Xiao, K., Ghelani, N., Aroyo, L., Slone, A., Houlsby, N., Xiong, X., Yang, Z., Gribovskaya, E., Adler, J., Wirth, M., Lee, L., Li, M., Kagohara, T., Pavagadhi, J., Bridgers, S., Bortsova, A., Ghemawat, S., Ahmed, Z., Liu, T., Powell, R., Bolina, V., Iinuma, M., Zablotskaia, P., Besley, J., Chung, D.-W., Dozat, T., Comanescu, R., Si, X., Greer, J., Su, G., Polacek, M., Kaufman, R. L., Tokumine, S., Hu, H., Buchatskaya, E., Miao, Y., Elhawaty, M., Siddhant, A., Tomasevic, N., Xing, J., Greer, C., Miller, H., Ashraf, S., Roy, A., Zhang, Z., Ma, A., Filos, A., Besta, M., Blevins, R., Klimenko, T., Yeh, C.-K., Changpinyo, S., Mu, J., Chang, O., Pajarskas, M., Muir, C., Cohen, V., Lan, C. L., Haridasan, K. S., Marathe, A., Hansen, S., Douglas, S., Samuel, R., Wang, M., Austin, S., Lan, C., Jiang, J., Chiu, J., Lorenzo, J. A., Sjosund, L. L., Cevey, S., Gleicher, Z., Avrahami, T., Boral, A., Srinivasan, H., Selo, V., May, R., Aisopos, K., Hussenot, L., Soares, L. B., Baumli, K., Chang, M. B., Recasens, A., Caine, B., Pritzel, A., Pavetic, F., Pardo, F., Gergely, A., Frye, J., Ramasesh, V. V., Horgan, D., Badola, K., Kassner, N., Roy, S., Dyer, E., Campos, V., Tomala, A., Tang, Y., Badawy, D. E., White, E., Mustafa, B., Lang, O., Jindal, A., Vikram, S., Gong, Z., Caelles, S., Hemsley, R., Thornton, G., Feng, F., Stokowiec, W., Zheng, C., Thacker, P., cCauglar Unlu, Zhang, Z., Saleh, M., Svensson, J., Bileschi, M. L., Patil, P., Anand, A., Ring, R., Tsihlas, K., Vezer, A., Selvi, M., Shevlane, T., Rodriguez, M., Kwiatkowski, T., Daruki, S., Rong, K., Dafoe, A., FitzGerald, N., Gu-Lemberg, K., Khan, M., Hendricks, L. A., Pellat, M., Feinberg, V., Cobon-Kerr, J., Sainath, T. N., Rauh, M., Hashemi, S. H., Ives, R., Hasson, Y., Li, Y., Noland, E., Cao, Y., Byrd, N., Hou, L., Wang, Q., Sottiaux, T., Paganini, M., Lespiau, J.-B., Moufarek, A., Hassan, S., Shivakumar, K., van Amersfoort, J. R., Mandhane, A., Joshi, P. M., Goyal, A., Tung, M., Brock, A., Sheahan, H., Misra, V., Li, C., Raki’cevi’c, N., Dehghani, M., Liu, F., Mittal, S., Oh, J., Noury, S., Sezener, E., Huot, F., Lamm, M., Cao, N. D., Chen, C., Elsayed, G., hsin Chi, E. H., Mahdieh, M., Tenney, I., Hua, N., Petrychenko, I., Kane, P., Scandinaro, D., Jain, R., Uesato, J., Datta, R., Sadovsky, A., Bunyan, O., Rabiej, D., Wu, S., Zhang, J., Vasudevan, G., Leurent, E., Alnahlawi, M., Georgescu, I.-R., Wei, N., Zheng, I., Chan, B., Rabinovitch, P. G., Stanczyk, P., Zhang, Y., Steiner, D., Naskar, S., Azzam, M., Johnson, M., Paszke, A., Chiu, C.-C., Elias, J. S., Mohiuddin, A., Muhammad, F., Miao, J., Lee, A., Vieillard, N., Potluri, S., Park, J., Davoodi, E., Zhang, J., Stanway, J., Garmon, D., Karmarkar, A., Dong, Z., Lee, J., Kumar, A., Zhou, L., Evens, J., Isaac, W., Chen, Z., Jia, J., Levskaya, A., Zhu, Z., Gorgolewski, C. F., Grabowski, P., Mao, Y., Magni, A., Yao, K., Snaider, J., Casagrande, N., Suganthan, P., Palmer, E., Irving, G., Loper, E., Faruqui, M., Arkatkar, I., Chen, N., Shafran, I., Fink, M., Castano, A., Giannoumis, I., Kim, W., Rybi’nski, M., Sreevatsa, A., Prendki, J., Soergel, D. G., Goedeckemeyer, A., Gierke, W., Jafari, M., Gaba, M., Wiesner, J., Wright, D. G., Wei, Y., Vashisht, H., Kulizhskaya, Y., Hoover, J., Le, M., Li, L., Iwuanyanwu, C., Liu, L., Ramirez, K., Khorlin, A. Y., Cui, A., Lin, T., Georgiev, M., Wu, M., Aguilar, R., Pallo, K., Chakladar, A., Repina, A., Wu, X., van der Weide, T., Ponnapalli, P., Kaplan, C., Simsa, J., Li, S., Dousse, O., Piper, J., Ie, N., Lui, M., Pasumarthi, R. K., Lintz, N., Vijayakumar, A., Thiet, L. N., Andor, D., Valenzuela, P., Paduraru, C., Peng, D., Lee, K., Zhang, S., Greene, S., Nguyen, D. D., Kurylowicz, P., Velury, S., Krause, S., Hardin, C., Dixon, L., Janzer, L., Choo, K., Feng, Z., Zhang, B., Singhal, A., Latkar, T., Zhang, M., Le, Q. V., Abellan, E. A., Du, D., McKinnon, D., Antropova, N., Bolukbasi, T., Keller, O., Reid, D., Finchelstein, D. F., Raad, M. A., Crocker, R., Hawkins, P., Dadashi, R., Gaffney, C., Lall, S., Franko, K., Filonov, E., Bulanova, A., Leblond, R., Yadav, V., Chung, S., Askham, H., Cobo, L. C., Xu, K., Fischer, F., Xu, J., Sorokin, C., Alberti, C., Lin, C.-C., Evans, C., Zhou, H., Dimitriev, A., Forbes, H., Banarse, D. S., Tung, Z., Liu, J., Omernick, M., Bishop, C., Kumar, C., Sterneck, R., Foley, R., Jain, R., Mishra, S., Xia, J., Bos, T., Cideron, G., Amid, E., Piccinno, F., Wang, X., Banzal, P., Gurita, P., Noga, H., Shah, P., Mankowitz, D. J., Polozov, O., Kushman, N., Krakovna, V., Brown, S. M., Bateni, M., Duan, D., Firoiu, V., Thotakuri, M., Natan, T., Mohananey, A., Geist, M., Mudgal, S., Girgin, S., Li, H., Ye, J., Roval, O., Tojo, R., Kwong, M., Lee-Thorp, J., Yew, C., Yuan, Q., Bagri, S., Sinopalnikov, D., Ramos, S., Mellor, J. F. J., Sharma, A., Severyn, A., Lai, J., Wu, K., Cheng, H.-T., Miller, D., Sonnerat, N., Vnukov, D., Greig, R., Beattie, J., Caveness, E., Bai, L., Eisenschlos, J. M., Korchemniy, A., Tsai, T., Jasarevic, M., Kong, W., Dao, P., Zheng, Z., Liu, F., Zhu, R., Geller, M., Teh, T. H., Sanmiya, J., Gladchenko, E., Trdin, N., Sozanschi, A., Toyama, D., Rosen, E., Tavakkol, S., Xue, L., Elkind, C., Woodman, O., Carpenter, J., Papamakarios, G., Kemp, R., Kafle, S., Grunina, T., Sinha, R., Talbert, A., Goyal, A., Krishna, K., Wu, D., Owusu-Afriyie, D., Du, C., Thornton, C., Pont-Tuset, J., Narayana, P., Li, J., Fatehi, S., Wieting, J. M., Ajmeri, O., Uria, B., Zhu, T., Ko, Y., Knight, L., H’eliou, A., Niu, N., Gu, S., Pang, C., Tran, D., Li, Y., Levine, N., Stolovich, A., Kalb, N., Santamaria-Fernandez, R., Goenka, S., Yustalim, W., Strudel, R., Elqursh, A., Lakshminarayanan, B., Deck, C., Upadhyay, S., Lee, H., Dusenberry, M., Li, Z., Wang, X., Levin, K., Hoffmann, R., Holtmann-Rice, D. N., Bachem, O., Yue, S., Arora, S., Malmi, E., Mirylenka, D., Tan, Q., Koh, C., Yeganeh, S. H., Poder, S., Zheng, S., Pongetti, F., Tariq, M., Sun, Y., Ionita, L., Seyedhosseini, M., Tafti, P. D., Kotikalapudi, R., Liu, Z., Gulati, A., Liu, J., Ye, X., Chrzaszcz, B., Wang, L., Sethi, N., Li, T., Brown, B., Singh, S., Fan, W., Parisi, A., Stanton, J., Kuang, C., Koverkathu, V., Choquette-Choo, C. A., Li, Y., Lu, T., Ittycheriah, A., Shroff, P., Sun, P., Varadarajan, M., Bahargam, S., Willoughby, R., Gaddy, D., Dasgupta, I., Desjardins, G., Cornero, M., Robenek, B., Mittal, B., Albrecht, B., Shenoy, A., Moiseev, F., Jacobsson, H., Ghaffarkhah, A., Riviere, M., Walton, A., Crepy, C., Parrish, A., Liu, Y., Zhou, Z., Farabet, C., Radebaugh, C., Srinivasan, P., van der Salm, C., Fidjeland, A. Ø., Scellato, S., Latorre-Chimoto, E., Klimczak-Pluci’nska, H., Bridson, D., de Cesare, D., Hudson, T., Mendolicchio, P., Walker, L., Morris, A., Penchev, I., Mauger, M., Guseynov, A., Reid, A., Odoom, S., Loher, L., Cotruta, V., Yenugula, M., Grewe, D., Petrushkina, A., Duerig, T., Sanchez, A., Yadlowsky, S., Shen, A., Globerson, A., Kurzrok, A., Webb, L., Dua, S., Li, D., Lahoti, P., Bhupatiraju, S., Hurt, D., Qureshi, H., Agarwal, A., Shani, T., Eyal, M., Khare, A., Belle, S., Wang, L., Tekur, C., Kale, M., Wei, J., Sang, R., Saeta, B., Liechty, T., Sun, Y., Zhao, Y., Lee, S., Nayak, P., Fritz, D., Vuyyuru, M. R., Aslanides, J., Vyas, N., Wicke, M., Ma, X., Bilal, T., Eltyshev, E., Balle, D., Martin, N., Cate, H., Manyika, J., Amiri, K., Kim, Y., Xiong, X., Kang, K., Luisier, F., Tripuraneni, N., Madras, D., Guo, M., Waters, A., Wang, O., Ainslie, J., Baldridge, J., Zhang, H., Pruthi, G., Bauer, J., Yang, F., Mansour, R., Gelman, J., Xu, Y., Polovets, G., Liu, J., Cai, H., Chen, W., Sheng, X., Xue, E., Ozair, S., Yu, A. W., Angermueller, C., Li, X., Wang, W., Wiesinger, J., Koukoumidis, E., Tian, Y., Iyer, A., Gurumurthy, M., Goldenson, M., Shah, P., Blake, M., Yu, H., Urbanowicz, A., Palomaki, J., Fernando, C., Brooks, K., Durden, K., Mehta, H., Momchev, N., Rahimtoroghi, E., Georgaki, M. E., Raul, A., Ruder, S., Redshaw, M., Lee, J., Jalan, K., Li, D., Perng, G., Hechtman, B. A., Schuh, P., Nasr, M., Chen, M., Milan, K., Mikulik, V., Strohman, T., Franco, J., Green, T., Hassabis, D., Kavukcuoglu, K., Dean, J., and Vinyals, O. Gemini: A family of highly capable multimodal models. ArXiv, abs/2312.11805, 2023. URL https://api.semanticscholar.org/CorpusID:266361876.
  6. 6.Anna Montoya, D. House prices - advanced regression techniques, 2016. URL https://kaggle.com/competitions/house-prices-advanced-regression-techniques.
  7. 7.Anthropic. Introducing claude, 2023. URL https://www.anthropic.com/index/introducing-claude.
  8. 8.Berens, P., Cranmer, K., Lawrence, N. D., von Luxburg, U., and Montgomery, J. Ai for science: An emerging agenda. ArXiv, abs/2303.04217, 2023.
  9. 9.Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q. JAX: composable transformations of Python+NumPy programs, 2018. URL http://github.com/google/jax.
  10. 10.Chollet, F. et al. Keras, 2015. URL https://github.com/fchollet/keras.
  11. 11.Elsken, T., Metzen, J. H., and Hutter, F. Neural architecture search: a survey. J. Mach. Learn. Res., 20(1):1997–2017, jan 2019. ISSN 1532-4435.
  12. 12.Franklin, A., Maggie, Benner, M., Rambis, N., Baffour, P., Holbrook, R., Crossley, S., and ulrichboser. Feedback prize - english language learning, 2022. URL https://kaggle.com/competitions/feedback-prize-english-language-learning.
  13. 13.He, X., Zhao, K., and Chu, X. Automl: A survey of the state-of-the-art. Knowledge-Based Systems, 212:106622, 2021. ISSN 0950-7051. doi: https://doi.org/10.1016/j.knosys.2020.106622. URL https://www.sciencedirect.com/science/article/pii/S0950705120307516.
  14. 14.Howard, A., Chow, A., and Holbrook, R. Spaceship titanic, 2022. URL https://kaggle.com/competitions/spaceship-titanic.
  15. 15.Hu, W., Fey, M., Zitnik, M., Dong, Y., Ren, H., Liu, B., Catasta, M., and Leskovec, J. Open graph benchmark: Datasets for machine learning on graphs. ArXiv, abs/2005.00687, 2020.
  16. 16.Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de Las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. Mixtral of experts. ArXiv, abs/2401.04088, 2024. URL https://api.semanticscholar.org/CorpusID:266844877.
  17. 17.Jumper, J. M., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Z´ıdek, A., Potapenko, A., Bridgland, A., Meyer, C., Kohl, S. A. A., Ballard, A., Cowie, A., Romera-Paredes, B., Nikolov, S., Jain, R., Adler, J., Back, T., Petersen, S., Reiman, D. A., Clancy, E., Zielinski, M., Steinegger, M., Pacholska, M., Berghammer, T., Bodenstein, S., Silver, D., Vinyals, O., Senior, A. W., Kavukcuoglu, K., Kohli, P., and Hassabis, D. Highly accurate protein structure prediction with alphafold. Nature, 596:583 – 589, 2021.
  18. 18.King, R. D., Whelan, K. E., Jones, F. M., Reiser, P. G. K., Bryant, C. H., Muggleton, S. H., Kell, D. B., and Oliver, S. G. Functional genomic hypothesis generation and experimentation by a robot scientist. Nature, 427:247–252, 2004.
  19. 19.King, R. D., Rowland, J. J., Oliver, S. G., Young, M., Aubrey, W., Byrne, E., Liakata, M., Markham, M., Pir, P., Soldatova, L. N., Sparkes, A., Whelan, K. E., and Clare, A. The automation of science. Science, 324:85 – 89, 2009.
  20. 20.Kinniment, M., Sato, L. J. K., Du, H., Goodrich, B., Hasin, M., Chan, L., Miles, L. H., Lin, T. R., Wijk, H., Burget, J., Ho, A., Barnes, E., and Christiano, P. F. Evaluating language-model agents on realistic autonomous tasks. ArXiv, abs/2312.11671, 2023. URL https://api.semanticscholar.org/CorpusID:260472392.
  21. 21.Kirsch, L., Dane, S., Adam, S., and Dardov, V. Amp®-parkinson’s disease progression prediction, 2023. URL https://kaggle.com/competitions/amp-parkinsons-disease-progression-prediction.
  22. 22.Kitano, H. Nobel turing challenge: creating the engine for scientific discovery. NPJ Systems Biology and Applications, 7, 2021.
  23. 23.Kramer, S., Cerrato, M., Dzeroski, S., and King, R. D. Automated scientific discovery: From equation discovery to autonomous discovery systems. ArXiv, abs/2305.02251, 2023.
  24. 24.Krizhevsky, A. Learning multiple layers of features from tiny images. 2009.
  25. 25.Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., Yang, K., Zhang, S., Deng, X., Zeng, A., Du, Z., Zhang, C., Shen, S., Zhang, T., Su, Y., Sun, H., Huang, M., Dong, Y., and Tang, J. Agentbench: Evaluating LLMs as agents. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=zAdUB0aCTQ.
  26. 26.Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pp. 142–150, Portland, Oregon, USA, June 2011. Association for Computational Linguistics. URL http://www.aclweb.org/anthology/P11-1015.
  27. 27.Nakano, R., Hilton, J., Balaji, S., Wu, J., Long, O., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., Jiang, X., Cobbe, K., Eloundou, T., Krueger, G., Button, K., Knight, M., Chess, B., and Schulman, J. Webgpt: Browser-assisted question-answering with human feedback. ArXiv, abs/2112.09332, 2021. URL https://api.semanticscholar.org/CorpusID:245329531.
  28. 28.OpenAI. Gpt-4 technical report. ArXiv, abs/2303.08774, 2023.
  29. 29.Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, UIST ’23, New York, NY, USA, 2023. Association for Computing Machinery. ISBN 9798400701320. doi: 10.1145/3586183.3606763. URL https://doi.org/10.1145/3586183.3606763.
  30. 30.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32, pp. 8024–8035. Curran Associates, Inc., 2019. URL http://papers.neurips.cc/paper/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf.
  31. 31.Sarna, A., Elkin, C., inversion, Ng, J., Maggie, and Reade, W. Google research - identify contrails to reduce global warming, 2023. URL https://kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming.
  32. 32.Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T. Toolformer: Language models can teach themselves to use tools. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=Yacmpz84TH.
  33. 33.Schwaller, P., Gaudin, T., Lanyi, D., Bekas, C., and Laino, T. “found in translation”: predicting outcomes of complex organic chemistry reactions using neural sequence-to-sequence models† †electronic supplementary information (esi) available: Time-split test set and example predictions, together with attention weights, confidence and token probabilities. see do. Chemical Science, 9:6091 – 6098, 2017.
  34. 34.Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K. R., and Yao, S. Reflexion: language agents with verbal reinforcement learning. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=vAElhFcKW6.
  35. 35.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. Llama: Open and efficient foundation language models. ArXiv, abs/2302.13971, 2023.
  36. 36.Velivckovi’c, P., Badia, A. P., Budden, D., Pascanu, R., Banino, A., Dashevskiy, M., Hadsell, R., and Blundell, C. The clrs algorithmic reasoning benchmark. In International Conference on Machine Learning, 2022. URL https://api.semanticscholar.org/CorpusID:249210177.
  37. 37.Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L. J., and Anandkumar, A. Voyager: An open-ended embodied agent with large language models. ArXiv, abs/2305.16291, 2023a.
  38. 38.Wang, Q., Downey, D., Ji, H., and Hope, T. Scimon: Scientific inspiration machines optimized for novelty. arXiv preprint arXiv:2305.14259, 2023b.
  39. 39.Warstadt, A., Choshen, L., Mueller, A., Williams, A., Wilcox, E. G., and Zhuang, C. Call for papers - the babylm challenge: Sample-efficient pretraining on a developmentally plausible corpus. ArXiv, abs/2301.11796, 2023.
  40. 40.Woodward, B., eor123, GenevievePatterson, and Carlsen, L. Fathomnet 2023, 2023. URL https://kaggle.com/competitions/fathomnet-out-of-sample-detection.
  41. 41.Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR), 2023.
  42. 42.Zhang, L., Zhang, Y., Ren, K., Li, D., and Yang, Y. Mlcopilot: Unleashing the power of large language models in solving machine learning tasks. ArXiv, abs/2304.14979, 2023a. URL https://api.semanticscholar.org/CorpusID:258418182.
  43. 43.Zhang, M., Qamar, M., Kang, T., Jung, Y., Zhang, C., Bae, S.-H., and Zhang, C. A survey on graph diffusion models: Generative ai in science for molecule, protein and material. ArXiv, abs/2304.01565, 2023b.
  44. 44.Zhang, S., Gong, C., Wu, L., Liu, X., and Zhou, M. Automl-gpt: Automatic machine learning with gpt. ArXiv, abs/2305.02499, 2023c. URL https://api.semanticscholar.org/CorpusID:258480269.
  45. 45.Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Bisk, Y., Fried, D., Alon, U., and Neubig, G. Webarena: A realistic web environment for building autonomous agents. ArXiv, abs/2307.13854, 2023. URL https://api.semanticscholar.org/CorpusID:260164780.

Citation

MLA
Huang, Q., et al. “MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation”. arXiv, 2023, http://arxiv.org/abs/2310.03302v2.
APA
Huang, Q., Vora, J., Liang, P., & Leskovec, J. (2023). MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation. arXiv. http://arxiv.org/abs/2310.03302v2
Chicago
Huang, Q., J. Vora, P. Liang, and J. Leskovec. 2023. “MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation”. arXiv. http://arxiv.org/abs/2310.03302v2.
Harvard
Huang, Q. et al. (2023) “MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.03302v2.
Vancouver
1. Huang Q, Vora J, Liang P, Leskovec J (2023) MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation. arXiv

BibTeX

@article{huang2023mlagentbench,
  title = {MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation},
  author = {Huang, Qian and Vora, Jian and Liang, Percy and Leskovec, Jure},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.03302v2},
  eprint = {2310.03302}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/