WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?

Alexandre DrouinMaxime GasseMassimo CacciaIssam H. LaradjiManuel Del VermeTom MartyDavid VázquezNicolas ChapadosAlexandre Lacoste

article2024ICML177 citations

Introduces WorkArena and the BrowserGym environment to evaluate web agents on realistic enterprise workflows in ServiceNow, revealing critical automation limitations and a wide performance gap between open- and closed-source language models.

Listen

Enterprise software systems often prioritize deep functionality over user experience, leading to complex interfaces that require steep learning curves and impose burdensome, repetitive tasks on knowledge workers. While autonomous artificial intelligence agents capable of interacting directly with web browsers could significantly enhance workforce productivity and digital accessibility, most evaluation benchmarks have focused on toy tasks or consumer websites rather than complex workplace tools.

The article aims to evaluate how effectively state-of-the-art large language models can perform everyday knowledge work tasks in realistic enterprise environments. To achieve this, the authors introduce WorkArena, a benchmark based on the widely used ServiceNow platform, alongside BrowserGym, a modular development and testing environment for web agents.

The authors designed 33 representative enterprise tasks encompassing 19,912 unique instances, including form filling, list filtering and sorting, dashboard data retrieval, service catalog ordering, and knowledge base searches. Using BrowserGym to supply agents with web accessibility trees, coordinate spaces, and interactive chat capabilities, the study evaluated leading commercial models (GPT-4o and GPT-3.5) and a top open-source model (Llama 3-70B) in a zero-shot setup with a cap of 15 execution steps per episode.

The findings reveal a substantial gap between current model capabilities and autonomous task execution. GPT-4o demonstrated the highest overall success rate on WorkArena at 42.7%, while Llama 3-70B achieved 17.9% and GPT-3.5 scored only 6.1%. Specialized dynamic interfaces presented severe hurdles; for example, list-filtering tasks involving non-standard controls yielded a 0% success rate across all models. The ablation analysis showed that step-by-step reasoning via chain-of-thought prompting is critical, whereas incorporating visual screenshots or overloading prompts with extensive action histories and descriptions frequently degraded performance or provided negligible benefit.

These results indicate that while foundation models are not yet ready for unattended, end-to-end automation in enterprise software, they are viable for guided assistance on well-defined operations like knowledge retrieval and catalog ordering. Deploying current models without human supervision introduces substantial operational risks, especially on tasks with complex forms or data table modifications. However, the rapid progress seen from older open-source models to current versions suggests that model capability is growing steadily.

Decision-makers should focus near-term investments on partial automation and supervised human-in-the-loop assistants rather than fully autonomous agents for enterprise workflows. Technology teams should prioritize developing agents with better long-context comprehension of web document structures and fine-tune vision models on user interface data before attempting mission-critical deployments.

The study's conclusions are constrained by an episode limit of 15 actions, a zero-shot prompting setup, and evaluation centered on a single software platform. While confidence is high that modern web agents struggle with non-standard enterprise interfaces, performance may improve under customized training regimes, larger action budgets, or domain-adapted models.

No sufficiently relevant recommendations were found.

Cover for WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?

Abstract

We study the use of large language model-based agents for interacting with software via web browsers. Unlike prior work, we focus on measuring the agents' ability to perform tasks that span the typical daily work of knowledge workers utilizing enterprise software systems. To this end, we propose WorkArena, a remote-hosted benchmark of 33 tasks based on the widely-used ServiceNow platform. We also introduce BrowserGym, an environment for the design and evaluation of such agents, offering a rich set of actions as well as multimodal observations. Our empirical evaluation reveals that while current agents show promise on WorkArena, there remains a considerable gap towards achieving full task automation. Notably, our analysis uncovers a significant performance disparity between open and closed-source LLMs, highlighting a critical area for future exploration and development in the field.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 WorkArena – An Enterprise Benchmark
  • 3.1 WorkArena Tasks
  • 3.2 Challenges: the World Wild Web of Work
  • 3.3 Availability
  • 4 BrowserGym
  • 4.1 Capabilities
  • 4.2 An Ideal Experimental Framework
  • 5 Experiments
  • 5.1 Agent Design
  • 5.2 Experimental Protocol
  • 5.3 Results
  • 5.4 Ablation Study
  • 6 Conclusion
  • References
  • A WorkArena – Additional Details
  • A.1 Tasks
  • A.2 Task User Interface Examples
  • A.3 Knowledge Base Tasks – Additional Details
  • B BrowserGym – Additional Details
  • B.1 Action Space
  • B.2 MiniWoB

Knowls

  1. Knowl 1 — WorkArena covers common enterprise-software workflows

    definition

    WorkArena is a benchmark of 33 ServiceNow tasks with 19,912 unique instances. Its tasks are grouped into six categories:

    • Lists: 12 tasks and 6,900 instances. Six tasks build filters with 1–5 conditions; six sort by up to three columns. Each group covers different data tables.
    • Forms: 5 tasks and 5,000 instances for creating records, with between 1 and 26 fields to fill.
    • Knowledge base: 1 task and 1,000 instances for finding information in articles and answering a question.
    • Service catalogs: 9 tasks and 3,550 instances for ordering products with specified quantities and options.
    • Dashboards: 4 tasks and 1,862 instances for retrieving chart values or extrema, sometimes across multiple charts.
    • Menus: 2 tasks and 1,600 instances for navigating to a module or impersonating a user.

    Task goals describe these operations in natural language. Collectively, the categories test list manipulation, record creation, information retrieval, purchasing, chart reading, and application navigation.

  2. Knowl 2 — WorkArena tasks use generated goals, validators, and executable oracles

    model/method

    Each WorkArena instance is created from a human-designed natural-language template populated with predefined values, such as a field value, menu name, or product specification. The goal explicitly supplies the information needed to carry out the task.

    A task-specific validator checks whether the goal was fulfilled. Validation may inspect the interface or URL, query the database for a created record or order, or check the answer sent in chat. Some validators also provide immediate feedback about errors, ranging from missing required fields to invalid database entries.

    Tasks can also provide a hand-written Playwright oracle that completes the task automatically. The oracle establishes that the task is feasible, can provide ground-truth demonstrations for learning agents, and helps identify tasks affected by platform updates.

  3. Knowl 3 — BrowserGym provides multimodal browser observations and flexible actions

    model/method

    BrowserGym is an OpenAI Gym-style browser environment implemented with Chromium, the Chrome DevTools Protocol, and Playwright. It models browser interaction as a partially observable decision process: at each step, the agent receives information about the current view rather than an automatically maintained history.

    An observation includes chat messages, open-page URLs, any error from the last action, and a multimodal representation of the active page: a DOM snapshot, an accessibility tree, and a screenshot. Page elements are augmented with a unique browser element identifier (bid), bounding-box coordinates, and visibility and clickability flags.

    The configurable action space includes identifier-based actions such as clicking or filling an element, coordinate-based mouse and keyboard actions, tab and navigation controls, chat messages, and Python code that can use Playwright. The environment supports multiple pages, including tabs and popups, as well as nested iframes and shadow DOMs.

  4. Knowl 4 — BrowserGym makes new task benchmarks implementable through a small task interface

    model/method

    A BrowserGym task is implemented with setup(), teardown(), and validate() functions, plus an optional cheat() oracle. setup() prepares the task, such as by creating data, authenticating, or navigating to a starting URL. teardown() cleans up resources. validate() checks whether the goal was achieved and returns a reward, an optional chat message, and a completion flag. The optional cheat() function is a hard-coded Playwright solution that can establish task feasibility.

    This common interface is intended to let researchers compare different agents on the same tasks while allowing agent design to vary, including which observations to use, how to handle history, and which action types to select. BrowserGym supports WorkArena, MiniWoB, and WebArena. In its MiniWoB port, the task goal is moved from the page HTML into the chat interface.

  5. Knowl 5 — Evaluation shows a large gap between closed- and open-source LLM agents

    empirical result

    The table reports success rate (SR) and standard error (SE), both in percentage points, for the selected agents on WorkArena, MiniWoB, and WebArena. GPT-4o-V is GPT-4o augmented with a screenshot and Set-of-Mark visual prompting. Subset rows report scores for the named portion of each benchmark.

    Task category (number of tasks) GPT-4o GPT-4o-V GPT-3.5 Llama3
    WorkArena (33) 42.7 ±\pm 1.5 41.8 ±\pm 1.7 6.1 ±\pm 1.3 17.9 ±\pm 1.5
    Dashboard (4) 62.5 ±\pm 6.8 72.5 ±\pm 6.0 20.0 ±\pm 4.8 37.5 ±\pm 6.0
    Form (5) 40.0 ±\pm 5.9 34.0 ±\pm 4.8 2.0 ±\pm 2.5 32.0 ±\pm 4.6
    Knowledge (1) 80.0 ±\pm 12.2 70.0 ±\pm 13.9 0.0 ±\pm 4.3 30.0 ±\pm 12.3
    List-filter (6) 0.0 ±\pm 1.6 0.0 ±\pm 1.7 0.0 ±\pm 1.6 0.0 ±\pm 1.8
    List-sort (6) 10.0 ±\pm 3.8 13.3 ±\pm 4.0 8.3 ±\pm 3.7 1.7 ±\pm 2.5
    Menu (2) 60.0 ±\pm 8.0 90.0 ±\pm 6.0 5.0 ±\pm 4.7 0.0 ±\pm 2.9
    Service catalog (9) 77.8 ±\pm 3.2 65.6 ±\pm 3.6 5.6 ±\pm 2.3 26.7 ±\pm 3.4
    MiniWoB (125) 66.1 ±\pm 1.0 67.7 ±\pm 1.0 38.9 ±\pm 1.1 62.6 ±\pm 0.6
    WebGum subset (56) 82.9 ±\pm 1.5 83.2 ±\pm 1.5 53.6 ±\pm 1.4 80.5 ±\pm 1.0
    WebArena (812) 23.5 ±\pm 0.7 24.0 ±\pm 0.6 6.7 ±\pm 0.6 11.0 ±\pm 0.6
    Content-and-config (411) 25.8 ±\pm 1.0 26.8 ±\pm 0.9 8.8 ±\pm 0.8 12.7 ±\pm 0.9
    Information-seeking (325) 22.5 ±\pm 1.0 22.5 ±\pm 0.9 4.3 ±\pm 0.9 9.8 ±\pm 1.1
    Navigation (76) 15.8 ±\pm 2.2 15.8 ±\pm 1.8 5.3 ±\pm 1.9 6.6 ±\pm 1.9

    GPT-4o has the highest overall WorkArena score at 42.7%, compared with 17.9% for Llama3 and 6.1% for GPT-3.5. The agents perform much better on MiniWoB than on WorkArena or WebArena, and all four score 0% on WorkArena list-filter tasks. Adding visual input yields no consistent overall gain: GPT-4o-V scores slightly below GPT-4o on WorkArena and slightly above it on MiniWoB and WebArena.

  6. Knowl 6 — Enterprise interface complexity is a central WorkArena challenge

    empirical result

    The ServiceNow pages used by WorkArena contain dynamic and sometimes nonstandard interfaces: fields can become required or hidden depending on other inputs, menus can appear through unusual interactions, and pages may combine nested iframes, shadow DOMs, and proprietary elements. After basic cleaning, page HTML can still range from 40,000 to 500,000 tokens. In the experiments, the researchers therefore used accessibility trees rather than HTML for WorkArena and WebArena.

    The task results indicate that difficulty is not simply a matter of high-level task intent. All evaluated agents achieved 0% on list-filter tasks, which require operating a nonstandard list widget; sorting tasks also remained difficult, with the highest score only 13.3%. The paper identifies long observations and complex real-world UI interactions as major obstacles for current agents.

  7. Knowl 7 — Agent evaluation uses zero-shot reasoning under bounded context and interaction budgets

    experimental setup

    The agents use a common codebase and chain-of-thought prompting, with one generic example of the expected reasoning and action format rather than task-specific demonstrations. The evaluated language models are GPT-4o, GPT-3.5, and Llama3-70B-Instruct; a vision-augmented GPT-4o variant is also evaluated. Model configurations are selected by random search on MiniWoB and WorkArena, then evaluated with a separate seed.

    For WorkArena and WebArena, agents receive accessibility trees rather than HTML because the HTML is too large; MiniWoB agents receive both. Maximum prompt lengths are 40K tokens for GPT-4o, 15K for GPT-3.5, and 8K for Llama3, with the end of the page representation truncated when necessary. Episodes are capped at 15 steps, and malformed model outputs can be reparsed by prompting the agent up to four times. WorkArena and MiniWoB results use 10 seeds per task; WebArena uses one seed per task. The reported standard errors are computed from 1,000 stratified-bootstrap samples.

  8. Knowl 8 — Chain-of-thought helps, while extra agent features have benchmark-dependent effects

    empirical result

    Ablations show that removing chain-of-thought reasoning reduces performance for the tested models. For example, on WorkArena GPT-3.5 falls from 8.5% to 6.1%, and Llama3 falls from 20.0% to 8.5%; on MiniWoB, their scores fall from 41.3% to 30.2% and from 59.8% to 48.6%, respectively. These ablation runs use a different seed from the main evaluation.

    Other features do not have uniformly positive effects. Adding coordinate observations and actions raises GPT-4o's MiniWoB score from 68.2% to 72.6% in the ablation, but lowers its WorkArena score from 45.5% to 41.2%. Adding thought history lowers GPT-4o's WorkArena score from 45.5% to 42.4%; the authors report that agents can persist with early mistakes rather than self-correct. They also observe that longer prompts and nonessential prompt features can hurt weaker models, plausibly because the already-large page representation leaves less context for task-relevant information.

  9. Knowl 9 — WorkArena knowledge-base articles and questions are generated from controlled facts

    model/method

    The knowledge-base task uses 100 facts, each pairing an item to be looked up with its value. GPT-4 generates one HTML article per fact, ensuring that the article states the fact-value relation. For each fact, GPT-4 is prompted to create 10 alternative questions and formatting instructions. GPT-3.5 is then used to check whether each question can be answered correctly from the generated article; questions it cannot answer are revised with GPT-4 and checked again. This independent checking model is used to reduce the chance of GPT-4 producing questions that are ambiguous but answerable by itself.

    For evaluation, GPT-4 generates alternative answer formats from the expected value. The validator accepts answers from a set of formats rather than requiring exact string equality. This process supplies 1,000 WorkArena knowledge-base instances while allowing minor wording and formatting variations in responses.

  10. Knowl 10 — The benchmark does not yet evaluate compositional end-to-end work workflows

    limitation

    WorkArena's task goals are deliberately explicit and provide the information needed to complete an instance. Although the benchmark covers operations that can form realistic work trajectories, its tasks are not yet composed into longer workflows that chain multiple skills or subtasks. The paper identifies compositional workflows involving capabilities such as retrieval, memorization, visual perception, and advanced reasoning as future expansions. It also notes that the 15-step episode limit may be insufficient for some WorkArena tasks unless an agent can issue multiple actions in one step.

Coverage note — The complete list of primitive action signatures and the per-task oracle-action counts are omitted as implementation inventories; the core action families, task categories, and evaluation findings are included.

References

  1. 1.Assouel, R., Marty, T., Caccia, M., Laradji, I., Drouin, A., Rajeswar, S., Palacios, H., Cappart, Q., Vazquez, D., Chapados, N., Gasse, M., and Lacoste, A. The unsolved challenges of LLMs in open-ended web tasks: A case study. In NeurIPS 2023 Foundation Models for Decision Making Workshop, 2023. URL https://openreview.net/forum?id=jt3il4fC5B.
  2. 2.Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. OpenAI gym, 2016.
  3. 3.Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y. Mind2Web: Towards a generalist agent for the web. arXiv, abs/2306.06070, 2023.
  4. 4.Furuta, H., Nachum, O., Lee, K.-H., Matsuo, Y., Gu, S. S., and Gur, I. Multimodal web navigation with instruction-finetuned foundation models. arXiv, abs/2305.11854, 2023. URL https://arxiv.org/abs/2305.11854.
  5. 5.Google. Chrome devtools protocol, 2023. URL https://chromedevtools.github.io/devtools-protocol/.
  6. 6.Gur, I., Furuta, H., Huang, A., Safdari, M., Matsuo, Y., Eck, D., and Faust, A. A real-world webagent with planning, long context understanding, and program synthesis. arXiv preprint arXiv:2307.12856, 2023a.
  7. 7.Gur, I., Furuta, H., Huang, A., Safdari, M., Matsuo, Y., Eck, D., and Faust, A. A real-world WebAgent with planning, long context understanding, and program synthesis. arXiv, abs/2307.12856, 2023b. URL https://arxiv.org/abs/2307.12856.
  8. 8.He, H., Yao, W., Ma, K., Yu, W., Dai, Y., Zhang, H., Lan, Z., and Yu, D. WebVoyager: Building an end-to-end web agent with large multimodal models. arXiv, abs/2401.13919, 2024. URL https://arxiv.org/abs/2401.13919.
  9. 9.Huang, F., Li, G., Li, T., and Li, Y. Automatic macro mining from interaction traces at scale. In Proceedings of the CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY, USA, 2024. Association for Computing Machinery. ISBN 9798400703300. doi: 10.1145/3613904.3642074. URL https://doi.org/10.1145/3613904.3642074.
  10. 10.Humphreys, P. C., Raposo, D., Pohlen, T., Thornton, G., Chhaparia, R., Muldal, A., Abramson, J., Georgiev, P., Santoro, A., and Lillicrap, T. A data-driven approach for learning to control computers. In International Conference on Machine Learning (ICML), 2022.
  11. 11.Kim, G., Baldi, P., and McAleer, S. Language models can solve computer tasks. arXiv, abs/2303.17491, 2023. URL https://arxiv.org/abs/2303.17491.
  12. 12.Li, Y., He, J., Zhou, X., Zhang, Y., and Baldridge, J. Mapping natural language instructions to mobile ui action sequences. In Annual Conference of the Association for Computational Linguistics (ACL 2020), 2020. URL https://www.aclweb.org/anthology/2020.acl-main.729.pdf.
  13. 13.Liu, E. Z., Guu, K., Pasupat, P., Shi, T., and Liang, P. Reinforcement learning on web interfaces using workflow-guided exploration. In International Conference on Learning Representations (ICLR), 2018.
  14. 14.Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., Yang, K., Zhang, S., Deng, X., Zeng, A., Du, Z., Zhang, C., Shen, S., Zhang, T., Su, Y., Sun, H., Huang, M., Dong, Y., and Tang, J. AgentBench: Evaluating LLMs as agents. arXiv, abs/2308.03688, 2023a. URL https://arxiv.org/abs/2308.03688.
  15. 15.Liu, Z., Yao, W., Zhang, J., Xue, L., Heinecke, S., Murthy, R., Feng, Y., Chen, Z., Niebles, J. C., Arpit, D., Xu, R., Mui, P., Wang, H., Xiong, C., and Savarese, S. BOLAA: Benchmarking and orchestrating LLM-augmented autonomous agents. arXiv, abs/2308.05960, 2023b.
  16. 16.Lu, X. H., Kasner, Z., and Reddy, S. Weblinx: Real-world website navigation with multi-turn dialogue. arXiv preprint arXiv:2402.05930, 2024.
  17. 17.Maas, M. Knowledge 2020: “The digital workflow revolution has just begun”. Technical report, Sprinklr, 2020. URL https://www.linkedin.com/pulse/knowledge-2020-digital-workflow-revolution-has-just-begun-maas/.
  18. 18.Mastantuono, G. ServiceNow joins the prestigious Fortune 500 list. https://www.servicenow.com/blogs/2023/servicenow-joins-fortune-500-list.html, 2023. Accessed: 2024-01-29.
  19. 19.Meta. Llama 3: Meta’s latest large language model. https://github.com/meta-llama/llama3, 2024. Accessed: 2024-06-03.
  20. 20.Microsoft. Playwright for Python documentation, 2023. URL https://playwright.dev/python/.
  21. 21.Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., Jiang, X., Cobbe, K., Eloundou, T., Krueger, G., Button, K., Knight, M., Chess, B., and Schulman, J. WebGPT: Browser-assisted question-answering with human feedback. arXiv, abs/2112.09332, 2021. URL https://arxiv.org/abs/2112.09332.
  22. 22.OpenAI. GPT-4 technical report. ArXiv, abs/2303.08774, 2023. URL https://arxiv.org/abs/2303.08774.
  23. 23.Rawles, C., Li, A., Rodriguez, D., Riva, O., and Lillicrap, T. Androidinthewild: A large-scale dataset for android device control. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in Neural Information Processing Systems, volume 36, pp. 59708–59728, 2023.
  24. 24.SAE. Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles. Technical report, Society of Automotive Engineers (SAE), 04 2021. URL https://doi.org/10.4271/J3016 202104.
  25. 25.ServiceNow. Vancouver release notes. Online, 2023. Available at: https://docs.servicenow.com/bundle/vancouver-release-notes/.
  26. 26.Shi, T., Karpathy, A., Fan, L., Hernandez, J., and Liang, P. World of bits: An open-domain platform for web-based agents. In International Conference on Machine Learning (ICML), 2017a.
  27. 27.Shi, T., Karpathy, A., Fan, L., Hernandez, J., and Liang, P. World of bits: An open-domain platform for web-based agents. ICML, 2017b.
  28. 28.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Canton-Ferrer, C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288, 2023.
  29. 29.van der Meer, J. A journey into the future of the translation industry, 2021. URL https://www.taus.net/resources/blog/a-journey-into-the-future-of-the-translation-industry. Accessed: 2024-02-01.
  30. 30.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022a.
  31. 31.Wei, J., Wang, X., Schuurmans, D., Bosma, M., ichter, b., Xia, F., Chi, E., Le, Q. V., and Zhou, D. Chain-of-thought prompting elicits reasoning in large language models. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), Advances in Neural Information Processing Systems, volume 35, pp. 24824–24837. Curran Associates, Inc., 2022b. URL https://proceedings.neurips.cc/paper_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf.
  32. 32.Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T. J., Cheng, Z., Shin, D., Lei, F., Liu, Y., Xu, Y., Zhou, S., Savarese, S., Xiong, C., Zhong, V., and Yu, T. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments, 2024.
  33. 33.Yang, J., Zhang, H., Li, F., Zou, X., Li, C., and Gao, J. Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v. arXiv preprint arXiv:2310.11441, 2023.
  34. 34.Yao, S., Chen, H., Yang, J., and Narasimhan, K. WebShop: Towards scalable real-world web interaction with grounded language agents. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  35. 35.Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. ReAct: Synergizing reasoning and acting in language models. arXiv, abs/2210.03629, 2023. URL https://arxiv.org/abs/2210.03629.
  36. 36.Zeng, A., Liu, M., Lu, R., Wang, B., Liu, X., Dong, Y., and Tang, J. Agenttuning: Enabling generalized agent abilities for llms. arXiv preprint arXiv:2310.12823, 2023.
  37. 37.Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Bisk, Y., Fried, D., Alon, U., and Neubig, G. Webarena: A realistic web environment for building autonomous agents. ArXiv, abs/2307.13854, 2023. URL https://arxiv.org/abs/2307.13854.

Citation

MLA
Drouin, A., et al. “WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?”. arXiv, 2024, http://arxiv.org/abs/2403.07718v5.
APA
Drouin, A., Gasse, M., Caccia, M., Laradji, I. H., Verme, M. D., Marty, T., Boisvert, L., Thakkar, M., Cappart, Q., Vazquez, D., Chapados, N., & Lacoste, A. (2024). WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?. arXiv. http://arxiv.org/abs/2403.07718v5
Chicago
Drouin, A., M. Gasse, M. Caccia, et al. 2024. “WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?”. arXiv. http://arxiv.org/abs/2403.07718v5.
Harvard
Drouin, A. et al. (2024) “WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.07718v5.
Vancouver
1. Drouin A, Gasse M, Caccia M, et al (2024) WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?. arXiv

BibTeX

@article{drouin2024workarena,
  title = {WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?},
  author = {Drouin, Alexandre and Gasse, Maxime and Caccia, Massimo and Laradji, Issam H. and Verme, Manuel Del and Marty, Tom and Boisvert, Léo and Thakkar, Megh and Cappart, Quentin and Vazquez, David and Chapados, Nicolas and Lacoste, Alexandre},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.07718v5},
  eprint = {2403.07718}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/