AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?

Ori YoranSamuel Joseph AmouyalChaitanya MalaviyaBen BoginOfir PressJonathan Berant

article2024EMNLP102 citations

Introduces AssistantBench, an automatically evaluated benchmark of realistic and time-consuming web tasks that exposes severe limitations in existing language models, alongside SeePlanAct, a new agent architecture designed to improve multi-step web execution.

Listen

Automated artificial intelligence assistants have significant potential to help users handle routine, time-consuming internet research, such as market monitoring or multi-step local searches. However, existing benchmarks primarily evaluate these tools on single websites or restricted sandbox environments rather than the open web. Consequently, current performance metrics fail to reflect whether automated systems can reliably navigate multiple live websites, plan dynamic browsing paths, and synthesize real-world information.

The article introduces ASSISTANTBENCH, a new evaluation benchmark designed to measure how effectively language models and autonomous web agents solve realistic, multi-step web tasks. The authors also present and evaluate a new web agent architecture, SEEPLANACT (SPA), created to improve browsing performance through explicit planning and memory mechanisms.

The benchmark comprises 214 realistic, automatically verifiable tasks spanning diverse domains collected from 53 human contributors, including 35 domain experts. Completing these tasks requires navigating more than 525 distinct webpages across 258 websites. The authors conducted comparative experiments using leading closed-book models, retrieval-augmented models with search engine access, existing state-of-the-art web agents, and their proposed SPA agent, using both GPT-4-Turbo and Claude-3.5-Sonnet engines.

The investigation produced several key findings. First, all evaluated systems struggled severely, with no standalone model exceeding 26% accuracy. Second, standalone autonomous web agents achieved poor overall results due to frequent navigation failures; the baseline SEEACT agent scored only 4.1% accuracy. Third, while the proposed SPA agent outperformed the baseline by 7 points in accuracy and reached 11.1% on its own, closed-book models achieved higher raw accuracy (up to 22.2%) simply because they abstained less often. However, closed-book models suffered from severe unreliability, hallucinating facts in 85% of their error cases. Fourth, an ensemble pairing the SPA agent with a closed-book fallback reached the top overall test accuracy of 25.2% to 26.4%. Finally, web agents exhibited high failure rates on both very short and very long interaction trajectories, with error rates peaking sharply on tasks requiring more than 15 browsing actions.

These findings indicate that relying on current AI models for open-web assistance introduces substantial operational and informational risks. Pure language models frequently present plausible but fabricated data, while retrieval-augmented systems routinely fail to retrieve complete context from complex web interfaces. Autonomous web agents currently suffer from navigation loops and interaction grounding errors, making them unreliable for unsupervised deployment in professional or high-stakes environments.

Organizations developing or deploying AI-driven web assistants should not rely on current standalone agents for end-to-end autonomous research without human verification. Developers should prioritize hybrid architectures that integrate explicit step-by-step planning and structured memory, while establishing robust fallback mechanisms when agents encounter navigation uncertainty. Further research must focus on improving open-web navigation training and developing reliable evaluation protocols for time-sensitive, dynamic web content.

Confidence in these findings is reinforced by realistic multi-domain task sourcing and consistent performance trends across multiple advanced language model backbones. Nevertheless, readers should account for certain limitations: the benchmark contains a relatively compact set of 214 tasks restricted to verifiable static outputs, and evaluation was bounded to proprietary commercial models due to the high computational costs of multi-turn web browsing.

arXiv: 2407.15711oriyor/assistantbench
Cover for AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?

Abstract

Language agents, built on top of language models (LMs), are systems that can interact with complex environments, such as the open web. In this work, we examine whether such agents can perform realistic and time-consuming tasks on the web, e.g., monitoring real-estate markets or locating relevant nearby businesses. We introduce AssistantBench, a challenging new benchmark consisting of 214 realistic tasks that can be automatically evaluated, covering different scenarios and domains. We find that AssistantBench exposes the limitations of current systems, including language models and retrieval-augmented language models, as no model reaches an accuracy of more than 26 points. While closed-book LMs perform well in terms of accuracy, they exhibit low precision and tend to hallucinate facts. State-of-the-art web agents reach a score of near zero. Additionally, we introduce SeePlanAct (SPA), a new web agent that significantly outperforms previous agents, and an ensemble of SPA and closed-book models reaches the best overall performance. Moreover, we analyze failures of current systems and highlight that open web navigation remains a major challenge.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 AssistantBench
  • 3.1 Criteria for AssistantBench Tasks
  • 3.2 Data Collection
  • 3.3 Data Statistics
  • 3.4 Automatic Evaluation
  • 4 SPA: See-Plan-Act
  • 5 Experiments
  • 5.1 Models
  • 5.2 Info-seeking Tasks over Wikipedia
  • 5.3 Results
  • 6 Analysis
  • 6.1 When is AssistantBench Challenging?
  • 6.2 Why is AssistantBench Challenging for Current Systems?
  • 6.3 Can ChatGPT with Web Search solve AssistantBench?
  • 6.4 Time-dependency of tasks in AssistantBench
  • 7 Related Work
  • 8 Conclusion
  • 9 Limitations
  • 10 Ethical Implications and Broader Impact
  • References
  • A Appendix
  • A.1 Detailed Comparison to Recent Benchmarks
  • A.2 Seed Data Collection
  • A.3 Expanding the Seed Set with Crowd-workers
  • A.4 Collecting Examples from Domain-experts
  • A.5 AssistantBench Development Set
  • A.6 Models
  • A.7 FanoutQA
  • A.8 Results
  • A.9 Examples and Analysis
  • A.10 Prompts

Knowls

  1. Knowl 1 — ASSISTANTBENCH Benchmark for Realistic Multi-Page Web Tasks

    definition

    ASSISTANTBENCH is an evaluation benchmark consisting of 214 realistic, time-consuming web navigation tasks designed to test autonomous agents browsing the open web without multimodal media processing tools (such as video or audio processing). The tasks require navigating across 525 webpages spanning 258 distinct websites.

    Every task in ASSISTANTBENCH is curated to satisfy three core criteria:

    1. Realistic: Addresses genuine human information needs.
    2. Time-consuming: Requires at least several minutes of human browsing and multi-page interaction to complete.
    3. Automatically verifiable: Has a closed-form answer that is verifiable programmatically and is stable over time.

    The benchmark consists of 33 development tasks and 181 test tasks collected via three pipelines:

    • General Seed Set: 70 tasks submitted by 18 individual users documenting actual recent multi-step web searches.
    • Expanded Set: 102 tasks generated by crowdworkers using the seed tasks as structural templates.
    • Expert Set: 42 tasks contributed by 35 domain experts recruited across more than 15 specialized fields (e.g., Biology, Law, Medicine) involving professional websites and databases.

    In terms of temporal stability, 43.5% of tasks are strictly static (enforced via explicit date constraints), 15.4% are historically stable over multi-year horizons, and 41.1% are business or real-world details unlikely to change within a one-year window.

  2. Knowl 2 — ASSISTANTBENCH Evaluation Protocol and Scoring Metrics

    equation

    ASSISTANTBENCH evaluates predictions against reference answers supporting three distinct output types: strings, numbers, and semi-structured dictionaries (allowing up to 5 answers per task).

    1. String Evaluation: Evaluated using word-level F1F_1 score between predicted and reference words. For list-valued answers, predicted and reference lists are aligned based on maximal word-level similarity.

    2. Numerical Evaluation: To provide partial credit for predictions close to the reference value AA, numerical predictions A′A' are evaluated using a bounded order-of-magnitude metric: Scorenum(A,A′)=max⁡{0,1−log⁡10(max⁡(A,A′)min⁡(A,A′))}\text{Score}_{\text{num}}(A, A') = \max\left\{0, 1 - \log_{10}\left(\frac{\max(A, A')}{\min(A, A')}\right)\right\} where A,A′>0A, A' > 0. The score is strictly bounded in [0,1][0, 1] and assigns 00 whenever the prediction differs from the reference by an order of magnitude or more.

    3. Dictionary / Semi-Structured Evaluation: For tasks requiring structured key-value outputs (e.g., entity name and corresponding tuition fee), models receive the expected JSON key schema in the prompt. Matching keys between the predicted and reference dictionaries are scored using word-level F1F_1 for string values and Scorenum\text{Score}_{\text{num}} for numerical values; missing keys receive a score of 0. Precision and recall over key-value pairs are computed and aggregated into an F1F_1 score.

    4. Task-Level Metrics:

    • Accuracy (Acc.): Mean score across all benchmark tasks, where abstained tasks (empty answers) receive a score of 0.
    • Answer Rate (Ans. %): Percentage of tasks where the model generated a non-empty response.
    • Precision (Prec.): Mean accuracy computed exclusively over the subset of tasks where the model generated a non-empty answer.
    • Exact Match (EM): Percentage of tasks where the predicted output exactly matches the gold reference.
  3. Knowl 3 — SEEPLANACT (SPA) Web Agent Architecture

    model/method

    SEEPLANACT (SPA) is an autonomous multimodal web agent designed for multi-hop, multi-page open-web information seeking that builds upon the SEEACT architecture by incorporating explicit planning, inter-step memory, and extended browser navigation actions.

    At each execution step tt (up to a budget of 30 steps), SPA maintains a state consisting of the current webpage visual screenshot, an accumulated text memory buffer MtM_t, and an execution plan PtP_t. Each step executes two successive multimodal language model calls:

    1. Step Analysis, Memory, and Re-Planning:

      • The model observes the current rendered page screenshot alongside the preceding action history.
      • Memory Update: The agent extracts task-relevant factual information visible on the page and appends it to MtM_t. If no new relevant data is present, it notes this state.
      • Plan Refinement: The agent evaluates progress against plan Pt−1P_{t-1} and produces a refined plan PtP_t, declaring termination if the information need is fulfilled.
      • Action Proposal: It describes the next desired user action and target interface element in natural language.
    2. Action Grounding:

      • The natural language action description is mapped to a concrete browser operation grounded in the page's HTML structure or an agent action.
      • Action Space: In addition to standard single-page interactions (CLICK, SELECT, TYPE, PRESS ENTER, TERMINATE), SPA includes open-web navigation primitives:
        • SCROLL: Scrolls down to 75% of the viewport or up to the top of the page.
        • GOTO: Directly loads a specified destination URL.
        • SEARCH: Issues a search query directly into Google in a single operation.
        • GOBACK: Navigates to the preceding page in browser history.
  4. Knowl 4 — Performance of Web Agents, Retrieval-Augmented Models, and Closed-Book LMs on ASSISTANTBENCH

    data/table

    Evaluation of closed-book models (CB-INST, CB-1S), retrieval-augmented models using Google Search (RALM-INST, RALM-1S), web agents (SEEACT, SPA), and fallback ensembles on the ASSISTANTBENCH test set (181 tasks) across GPT-4-Turbo and Claude-3.5-Sonnet engines.

    Engine Model Acc. Ans. % Prec. EM
    GPT-4-Turbo CB-INST 16.5 53.6 30.7 6.1
    GPT-4-Turbo CB-1S 22.2 89.5 24.8 8.3
    GPT-4-Turbo RALM-INST 11.8 60.2 19.5 5.5
    GPT-4-Turbo RALM-1S 10.7 48.1 22.4 3.9
    GPT-4-Turbo SEEACT 4.1 15.5 26.3 2.2
    GPT-4-Turbo SPA (ours) 11.1 35.9 30.9 5.5
    GPT-4-Turbo RALM-INST →\to CB 18.7 93.9 19.9 6.6
    GPT-4-Turbo RALM-1S →\to CB 19.5 92.8 21.0 6.1
    GPT-4-Turbo SEEACT →\to CB 23.4 89.5 26.1 9.4
    GPT-4-Turbo SPA →\to CB (ours) 25.2 91.7 27.5 9.9
    Claude-3.5-Sonnet CB-INST 17.7 69.1 25.6 6.1
    Claude-3.5-Sonnet CB-1S 21.9 76.2 28.8 6.6
    Claude-3.5-Sonnet RALM-INST 11.5 43.1 26.7 5.0
    Claude-3.5-Sonnet RALM-1S 11.0 42.5 25.9 3.3
    Claude-3.5-Sonnet SEEACT 2.2 13.8 15.8 1.7
    Claude-3.5-Sonnet SPA (ours) 12.9 34.3 37.7 8.8
    Claude-3.5-Sonnet RALM-INST →\to CB 22.5 79.6 28.3 8.3
    Claude-3.5-Sonnet RALM-1S →\to CB 21.6 82.3 26.3 6.6
    Claude-3.5-Sonnet SEEACT →\to CB 22.3 76.2 29.3 7.7
    Claude-3.5-Sonnet SPA →\to CB (ours) 26.4 81.8 32.2 13.8

    Key takeaways:

    • Standalone web agents exhibit very low answer rates (13.8%--35.9%) due to navigation breakdowns.
    • SPA outperforms SEEACT by 7.0 accuracy points on GPT-4-Turbo and 10.7 points on Claude-3.5-Sonnet, while achieving higher precision.
    • Closed-book models achieve competitive accuracy primarily by guessing frequently (answer rates up to 89.5%), but at lower precision.
    • Average evaluation costs per test example are 0.006−−0.006--0.007 for closed-book LMs, 0.049−−0.049--0.069 for RALMs, and 2.094−−2.094--2.472 for multimodal web agents.
  5. Knowl 5 — Fallback Ensembling Strategy for Web Agents

    model/method

    Because web agents and retrieval-augmented language models frequently fail to complete open-web navigation paths and abstain (returning an empty response), a fallback ensemble mechanism couples the precision of web interaction with the coverage of parametric language models.

    In this framework (denoted as Agent →\to CB or RALM →\to CB):

    1. The primary agent (e.g., SPA, SEEACT, or RALM) attempts to solve the task using interactive web browsing or search tools.
    2. If the agent successfully terminates with a non-empty prediction, that prediction is returned as the final output.
    3. If the primary agent abstains (due to navigation failure, loop detection, grounding failure, or reaching the 30-step execution limit without an answer), execution falls back to a one-shot closed-book language model (CB-1S) prompted to answer using its parametric weights.

    On the ASSISTANTBENCH test set, this ensemble achieves the highest overall accuracy across both model backbones: SPA→CB\text{SPA} \to \text{CB} achieves 25.2% accuracy and 9.9% EM on GPT-4-Turbo, and 26.4% accuracy and 13.8% EM on Claude-3.5-Sonnet, outperforming every standalone model.

  6. Knowl 6 — Error Taxonomy and Distribution for Web Agents on ASSISTANTBENCH

    data/table

    A manual qualitative error analysis on the ASSISTANTBENCH development set reveals why web agents fail on multi-step tasks. For both SPA and SEEACT, the vast majority of errors occur because the agent abstains from providing an answer (80.0% of failures for SPA and 97.0% for SEEACT).

    Error Cause SPA SEEACT
    Navigation error 36.7% 63.6%
    Grounding failure 23.3% 18.2%
    Technical issue 16.7% 9.1%
    No answer (answer visible on page but not generated) 3.3% 6.1%
    Wrong answer generated 20.0% 3.0%

    Description of error classes:

    • Navigation Error: The agent chooses irrelevant links, follows incorrect navigational paths, or gets trapped in repetitive scrolling/action loops without reaching the gold webpage.
    • Grounding Failure: The agent generates a valid natural language action but fails to correctly identify or bind to the corresponding HTML element or bounding box.
    • Technical Issue: Failures in browser instrumentation, such as Playwright screenshot capture crashes on dynamic web pages.
    • No Answer: The agent successfully navigates to and displays the webpage containing the gold answer, but fails to extract and state it.
    • Wrong Answer: The agent generates an incorrect final answer upon termination.
  7. Knowl 7 — Failure Modes of Retrieval-Augmented and Closed-Book Models on Multi-Step Tasks

    empirical result

    Qualitative error analysis on ASSISTANTBENCH reveals distinct failure mechanics across system classes:

    1. Retrieval-Augmented Language Models (RALMs): 80% of errors stem from retrieval failures, which fall into three distinct types:

      • Irrelevant Context (50.0%): Standard search engine query results return noise or superficial keyword matches lacking required facts.
      • Tool-Dependent Needs (38.5%): Tasks require interacting with interactive web tools (e.g., computing walking distances via Google Maps or filtering dynamic databases) which plain text snippet search engines cannot execute.
      • Partial Information (11.5%): The retriever retrieves only incomplete page fragments when answering requires aggregating data across large tables or full pages. The remaining 20% of RALM errors occur when the model fails to utilize relevant retrieved information or follows flawed reasoning chains.
    2. Closed-Book Language Models: When closed-book models output incorrect answers, 85% of errors are factual hallucinations and 15% are due to outdated parametric memory. When closed-book models abstain (57.6% of CB-INST cases), 90% of abstentions consist of the model outputting a step-by-step procedural plan for how a human could find the data rather than providing an answer.

    3. Commercial Web-Search Chatbots: When tested on the development set, ChatGPT (GPT-4o with web search and code execution) fails on over 90% of tasks. Common failure modes include over-relying on top search snippets to generate incorrect answers, and hallucinating synthetic/non-factual numerical data inside Python code interpreter executions to perform calculations.

  8. Knowl 8 — Non-Monotonic Relationship Between Trajectory Length and Agent Success

    empirical result

    On ASSISTANTBENCH, the accuracy of web agents (SPA, SEEACT) and multi-step retrieval-augmented models (RALM-INST, RALM-1S) is a non-monotonic function of execution trajectory length:

    • Short Trajectories (1--6 steps): Accuracy is near zero because ASSISTANTBENCH tasks require multi-hop exploration across multiple websites that cannot be resolved in very few actions.
    • Intermediate Trajectories (7--15 steps): Model accuracy peaks. For SPA, accuracy reaches its maximum of approximately 0.35 on trajectories spanning 10--12 steps.
    • Long Trajectories (> 15 steps): Accuracy rapidly degrades to near zero. Agents and retrieval systems operating past 15 steps invariably succumb to error compounding, getting stuck in unrecoverable cyclical loops (e.g., repeatedly scrolling the same page or re-querying search engines) or drifting into irrelevant parts of the web.
  9. Knowl 9 — Model Performance Across Task Difficulty and Data Source Categories

    data/table

    Tasks in the ASSISTANTBENCH test set are categorized by difficulty based on closed-book model success (Easy: both GPT-4-Turbo and Claude-3.5-Sonnet CB-1S achieve ≥0.5\ge 0.5 accuracy; Medium: exactly one achieves ≥0.5\ge 0.5; Hard: neither achieves ≥0.5\ge 0.5) and by data collection source.

    Model Easy Medium Hard Total Seed Expanded Expert
    (n=9n=9) (n=56n=56) (n=116n=116) (n=181n=181) (n=58n=58) (n=81n=81) (n=42n=42)
    CB-INST 51.2 40.2 2.3 16.5 15.8 19.6 11.3
    CB-1S 82.1 49.7 4.2 22.2 29.7 18.7 19.5
    RALM-INST 50.2 17.0 6.2 11.8 13.4 12.7 9.7
    RALM-1S 80.0 10.5 5.5 10.7 7.9 13.3 11.4
    SEEACT 28.9 2.5 2.9 4.1 2.9 1.9 9.7
    SPA 29.5 12.3 9.1 11.1 12.9 11.8 7.6
    SPA →\to CB 80.7 42.7 12.4 25.2 - - -

    Key observations:

    • 64.1% (116/181116/181) of tasks are Hard; on these tasks, closed-book models fail completely (4.2% accuracy), whereas web agents provide genuine utility.
    • On Domain-Expert tasks, 70% of gold answers reside within a single target URL (compared to only 20% in the general seed set). Consequently, single-page web agent SEEACT performs better on the expert set (9.7%) than on the multi-site seed set (2.9%).
  10. Knowl 10 — Evaluation of Multi-Hop Navigation on FANOUTQA Benchmark

    data/table

    To assess web agent performance in an environment restricted to a single information domain (Wikipedia), models were evaluated on a filtered development split of 31 multi-hop, multi-document dictionary-format tasks from FANOUTQA where the closed-book model CB-1S had accuracy <0.5< 0.5.

    Model Acc. Ans. % Prec. EM
    CB-INST 34.8 93.5 37.2 0.0
    CB-1S 40.9 100.0 40.9 0.0
    RALM-INST 9.6 93.5 10.2 0.0
    RALM-1S 27.3 93.5 29.2 0.0
    SEEACT 7.5 16.1 46.4 0.0
    SPA (ours) 30.0 61.3 48.9 9.7

    On Wikipedia multi-hop aggregation, SPA improves accuracy over SEEACT by 22.5 points (30.0% vs. 7.5%) and raises the answer rate from 16.1% to 61.3% while maintaining the highest precision (48.9%) of all evaluated models. SPA is the only model that achieves exact match answers on this set (9.7%, solving 3 of 31 tasks exactly).

Coverage note — Deliberately omitted verbatim prompt text templates from Appendix A.10 and survey UI screenshots from Appendix A.2-A.4, as the operational prompt structure, action set, and benchmark design are fully captured in the knowls.

References

  1. 1.Anthropic. 2024. Claude 3.5: Technical specifications and usage. Technical report, Anthropic.
  2. 2.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021. Evaluating large language models trained on code. Preprint, arXiv:2107.03374.
  3. 3.Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su. 2023. Mind2web: Towards a generalist agent for the web. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  4. 4.Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste. 2024. Workarena: How capable are web agents at solving common knowledge work tasks? Preprint, arXiv:2403.07718.
  5. 5.Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2368–2378, Minneapolis, Minnesota. Association for Computational Linguistics.
  6. 6.Hiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo, Aleksandra Faust, Shixiang Shane Gu, and Izzeddin Gur. 2024. Multimodal web navigation with instruction-finetuned foundation models. In The Twelfth International Conference on Learning Representations.
  7. 7.Izzeddin Gur, Hiroki Furuta, Austin V Huang, Mustafa Safdari, Yutaka Matsuo, Douglas Eck, and Aleksandra Faust. 2024. A real-world webagent with planning, long context understanding, and program synthesis. In The Twelfth International Conference on Learning Representations.
  8. 8.Izzeddin Gur, Ofir Nachum, Yingjie Miao, Mustafa Safdari, Austin Huang, Aakanksha Chowdhery, Sharan Narang, Noah Fiedel, and Aleksandra Faust. 2023. Understanding HTML with large language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 2803–2821, Singapore. Association for Computational Linguistics.
  9. 9.Izzeddin Gur, Ulrich Rückert, Aleksandra Faust, and Dilek Hakkani-Tür. 2019. Learning to navigate the web. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  10. 10.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. Realm: Retrieval-augmented language model pre-training. In Proceedings of the 37th International Conference on Machine Learning, ICML’20. JMLR.org.
  11. 11.Shibo Hao, Tianyang Liu, Zhen Wang, and Zhiting Hu. 2023. ToolkenGPT: Augmenting frozen language models with massive tools via tool embeddings. In Thirty-seventh Conference on Neural Information Processing Systems.
  12. 12.Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024. Webvoyager: Building an end-to-end web agent with large multimodal models. Preprint, arXiv:2401.13919.
  13. 13.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  14. 14.Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. 2022. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 9118–9147. PMLR.
  15. 15.Peter C. Humphreys, David Raposo, Tobias Pohlen, Gregory Thornton, Rachita Chhaparia, Alistair Muldal, Josh Abramson, Petko Georgiev, Adam Santoro, and Timothy P. Lillicrap. 2022. A data-driven approach for learning to control computers. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 9466–9482. PMLR.
  16. 16.Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2024. Atlas: few-shot learning with retrieval augmented language models. J. Mach. Learn. Res., 24(1).
  17. 17.Ashwin Kalyan, Abhinav Kumar, Arjun Chandrasekaran, Ashish Sabharwal, and Peter Clark. 2021. How much coffee was consumed during EMNLP 2019? fermi problems: A new reasoning challenge for AI. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7318–7328, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  18. 18.Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, and Minjoon Seo. 2024a. Prometheus: Inducing fine-grained evaluation capability in language models. In The Twelfth International Conference on Learning Representations.
  19. 19.Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Yuchen Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2024b. The biggen bench: A principled benchmark for fine-grained evaluation of language models with language models. Preprint, arXiv:2406.05761.
  20. 20.Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2024c. Prometheus 2: An open source language model specialized in evaluating other language models. Preprint, arXiv:2405.01535.
  21. 21.Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried. 2024. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks. ArXiv preprint, abs/2401.13649.
  22. 22.Yann LeCun. 2022. A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27. Open Review, 62(1).
  23. 23.Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  24. 24.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland. Association for Computational Linguistics.
  25. 25.Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang. 2018. Reinforcement learning on web interfaces using workflow-guided exploration. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net.
  26. 26.Fangyu Liu, Julian Eisenschlos, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Wenhu Chen, Nigel Collier, and Yasemin Altun. 2023a. DePlot: One-shot visual language reasoning by plot-to-table translation. In Findings of the Association for Computational Linguistics: ACL 2023, pages 10381–10399, Toronto, Canada. Association for Computational Linguistics.
  27. 27.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023b. Visual instruction tuning. In Thirty-seventh Conference on Neural Information Processing Systems.
  28. 28.Junpeng Liu, Yifan Song, Bill Yuchen Lin, Wai Lam, Graham Neubig, Yuanzhi Li, and Xiang Yue. 2024a. Visualwebbench: How far have multimodal llms evolved in web page understanding and grounding? Preprint, arXiv:2404.05955.
  29. 29.Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang. 2024b. Agentbench: Evaluating LLMs as agents. In The Twelfth International Conference on Learning Representations.
  30. 30.Xing Han Lù, Zdeněk Kasner, and Siva Reddy. 2024. Weblinx: Real-world website navigation with multi-turn dialogue. Preprint, arXiv:2402.05930.
  31. 31.Chaitanya Malaviya, Priyanka Agrawal, Kuzman Ganchev, Pranesh Srinivasan, Fantine Huot, Jonathan Berant, Mark Yatskar, Dipanjan Das, Mirella Lapata, and Chris Alberti. 2024a. Dolomites: Domain-specific long-form methodical tasks. Preprint, arXiv:2405.05938.
  32. 32.Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, and Dan Roth. 2024b. ExpertQA: Expert-curated questions and attributed answers. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 3025–3045, Mexico City, Mexico. Association for Computational Linguistics.
  33. 33.Grégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun, and Thomas Scialom. 2024. GAIA: a benchmark for general AI assistants. In The Twelfth International Conference on Learning Representations.
  34. 34.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2022. Webgpt: Browser-assisted question-answering with human feedback. Preprint, arXiv:2112.09332.
  35. 35.Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena. 2022. Show your work: Scratchpads for intermediate computation with language models.
  36. 36.OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. 2024. Gpt-4 technical report. Preprint, arXiv:2303.08774.
  37. 37.Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin. 2024. tinybenchmarks: evaluating llms with fewer examples. Preprint, arXiv:2402.14992.
  38. 38.Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah Smith, and Mike Lewis. 2023. Measuring and narrowing the compositionality gap in language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 5687–5711, Singapore. Association for Computational Linguistics.
  39. 39.Yujia Qin, Shengding Hu, Yankai Lin, Weize Chen, Ning Ding, Ganqu Cui, Zheni Zeng, Yufei Huang, Chaojun Xiao, Chi Han, Yi Ren Fung, Yusheng Su, Huadong Wang, Cheng Qian, Runchu Tian, Kunlun Zhu, Shihao Liang, Xingyu Shen, Bokai Xu, Zhen Zhang, Yining Ye, Bowen Li, Ziwei Tang, Jing Yi, Yuzhang Zhu, Zhenning Dai, Lan Yan, Xin Cong, Yaxi Lu, Weilin Zhao, Yuxiang Huang, Junxi Yan, Xu Han, Xian Sun, Dahai Li, Jason Phang, Cheng Yang, Tongshuang Wu, Heng Ji, Zhiyuan Liu, and Maosong Sun. 2023. Tool learning with foundation models. Preprint, arXiv:2304.08354.
  40. 40.Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. In-context retrieval-augmented language models. Transactions of the Association for Computational Linguistics, 11:1316–1331.
  41. 41.Christopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz, Gabrielle Lau, Marybeth Fair, Alice Li, William Bishop, Wei Li, Folawiyo Campbell-Ajala, Daniel Toyama, Robert Berry, Divya Tyamagundlu, Timothy Lillicrap, and Oriana Riva. 2024. Androidworld: A dynamic benchmarking environment for autonomous agents. Preprint, arXiv:2405.14573.
  42. 42.Christopher Rawles, Alice Li, Daniel Rodriguez, Oriana Riva, and Timothy P Lillicrap. 2023. Androidinthewild: A large-scale dataset for android device control. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  43. 43.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5418–5426, Online. Association for Computational Linguistics.
  44. 44.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kam-yar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. 2022. Photorealistic text-to-image diffusion models with deep language understanding. In Advances in Neural Information Processing Systems, volume 35, pages 36479–36494. Curran Associates, Inc.
  45. 45.Timo Schick, Jane Dwivedi-Yu, Roberto Dessi, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. In Thirty-seventh Conference on Neural Information Processing Systems.
  46. 46.Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni. 2020. Green ai. Commun. ACM, 63(12):54–63.
  47. 47.Peter Shaw, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, and Kristina Toutanova. 2023. From pixels to UI actions: Learning to follow instructions via graphical user interfaces. In Thirty-seventh Conference on Neural Information Processing Systems.
  48. 48.Tianlin Shi, Andrej Karpathy, Linxi Fan, Jonathan Hernandez, and Percy Liang. 2017. World of bits: An open-domain platform for web-based agents. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, volume 70 of Proceedings of Machine Learning Research, pages 3135–3144. PMLR.
  49. 49.Qiaoyu Tang, Ziliang Deng, Hongyu Lin, Xianpei Han, Qiao Liang, Boxi Cao, and Le Sun. 2023. Toolalpaca: Generalized tool learning for language models with 3000 simulated cases. ArXiv preprint, abs/2306.05301.
  50. 50.Daniel Toyama, Philippe Hamel, Anita Gergely, Gheorghe Comanici, Amelia Glaese, Zafarali Ahmed, Tyler Jackson, Shibl Mourad, and Doina Precup. 2021. Androidenv: A reinforcement learning platform for android. Preprint, arXiv:2105.13231.
  51. 51.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2023. Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 10014–10037, Toronto, Canada. Association for Computational Linguistics.
  52. 52.Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Jirong Wen. 2024. A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6).
  53. 53.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  54. 54.Yijun Xiao and William Yang Wang. 2021. On hallucination and predictive uncertainty in conditional language generation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2734–2744, Online. Association for Computational Linguistics.
  55. 55.Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. 2024. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Preprint, arXiv:2404.07972.
  56. 56.Kevin Xu, Yeganeh Kordi, Kate Sanders, Yizhong Wang, Adam Byerly, Jack Zhang, Benjamin Van Durme, and Daniel Khashabi. 2024. Tur[k]ingbench: A challenge benchmark for web agents. Preprint, arXiv:2403.11905.
  57. 57.John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. 2024. Swe-agent: Agent-computer interfaces enable automated software engineering. Preprint, arXiv:2405.15793.
  58. 58.Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2022. Webshop: Towards scalable real-world web interaction with grounded language agents. In Advances in Neural Information Processing Systems, volume 35, pages 20744–20757. Curran Associates, Inc.
  59. 59.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations.
  60. 60.Ori Yoran, Tomer Wolfson, Ben Bogin, Uri Katz, Daniel Deutch, and Jonathan Berant. 2023. Answering questions by meta-reasoning over multiple chains of thought. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5942–5966, Singapore. Association for Computational Linguistics.
  61. 61.Ori Yoran, Tomer Wolfson, Ori Ram, and Jonathan Berant. 2024. Making retrieval-augmented language models robust to irrelevant context. In The Twelfth International Conference on Learning Representations.
  62. 62.Ziniu Zhang, Shulin Tian, Liangyu Chen, and Ziwei Liu. 2024. Mmina: Benchmarking multihop multimodal internet agents. Preprint, arXiv:2404.09992.
  63. 63.Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. 2024. Gpt-4v(ision) is a generalist web agent, if grounded. ArXiv preprint, abs/2401.01614.
  64. 64.Longtao Zheng, Rundong Wang, Xinrun Wang, and Bo An. 2023. Synapse: Trajectory-as-exemplar prompting with memory for computer control. In NeurIPS 2023 Foundation Models for Decision Making Workshop.
  65. 65.Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Yonatan Bisk, Daniel Fried, Uri Alon, et al. 2024. Webarena: A realistic web environment for building autonomous agents. ICLR.
  66. 66.Andrew Zhu, Alyssa Hwang, Liam Dugan, and Chris Callison-Burch. 2024. Fanoutqa: Multi-hop, multi-document question answering for large language models. Preprint, arXiv:2402.14116.
  67. 67.Xizhou Zhu, Yuntao Chen, Hao Tian, Chenxin Tao, Weijie Su, Chenyu Yang, Gao Huang, Bin Li, Lewei Lu, Xiaogang Wang, Y. Qiao, Zhaoxiang Zhang, and Jifeng Dai. 2023. Ghost in the minecraft: Generally capable agents for open-world environments via large language models with text-based knowledge and memory. ArXiv preprint, abs/2305.17144.

Citation

MLA
Yoran, O., et al. “AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 8938–68, https://doi.org/10.18653/v1/2024.emnlp-main.505.
APA
Yoran, O., Amouyal, S. J., Malaviya, C., Bogin, B., Press, O., & Berant, J. (2024). AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 8938–8968. https://doi.org/10.18653/v1/2024.emnlp-main.505
Chicago
Yoran, O., S. J. Amouyal, C. Malaviya, B. Bogin, O. Press, and J. Berant. 2024. “AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 8938–68. https://doi.org/10.18653/v1/2024.emnlp-main.505.
Harvard
Yoran, O. et al. (2024) “AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 8938–8968. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.505.
Vancouver
1. Yoran O, Amouyal SJ, Malaviya C, Bogin B, Press O, Berant J (2024) AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 8938–8968

BibTeX

@inproceedings{yoran-etal-2024-assistantbench,
    title = "{A}ssistant{B}ench: Can Web Agents Solve Realistic and Time-Consuming Tasks?",
    author = "Yoran, Ori  and
      Amouyal, Samuel Joseph  and
      Malaviya, Chaitanya  and
      Bogin, Ben  and
      Press, Ofir  and
      Berant, Jonathan",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.505/",
    doi = "10.18653/v1/2024.emnlp-main.505",
    pages = "8938--8968"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/