WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?
Alexandre DrouinMaxime GasseMassimo CacciaIssam H. LaradjiManuel Del VermeTom MartyDavid VázquezNicolas ChapadosAlexandre Lacoste
Introduces WorkArena and the BrowserGym environment to evaluate web agents on realistic enterprise workflows in ServiceNow, revealing critical automation limitations and a wide performance gap between open- and closed-source language models.
Enterprise software systems often prioritize deep functionality over user experience, leading to complex interfaces that require steep learning curves and impose burdensome, repetitive tasks on knowledge workers. While autonomous artificial intelligence agents capable of interacting directly with web browsers could significantly enhance workforce productivity and digital accessibility, most evaluation benchmarks have focused on toy tasks or consumer websites rather than complex workplace tools.
The article aims to evaluate how effectively state-of-the-art large language models can perform everyday knowledge work tasks in realistic enterprise environments. To achieve this, the authors introduce WorkArena, a benchmark based on the widely used ServiceNow platform, alongside BrowserGym, a modular development and testing environment for web agents.
The authors designed 33 representative enterprise tasks encompassing 19,912 unique instances, including form filling, list filtering and sorting, dashboard data retrieval, service catalog ordering, and knowledge base searches. Using BrowserGym to supply agents with web accessibility trees, coordinate spaces, and interactive chat capabilities, the study evaluated leading commercial models (GPT-4o and GPT-3.5) and a top open-source model (Llama 3-70B) in a zero-shot setup with a cap of 15 execution steps per episode.
The findings reveal a substantial gap between current model capabilities and autonomous task execution. GPT-4o demonstrated the highest overall success rate on WorkArena at 42.7%, while Llama 3-70B achieved 17.9% and GPT-3.5 scored only 6.1%. Specialized dynamic interfaces presented severe hurdles; for example, list-filtering tasks involving non-standard controls yielded a 0% success rate across all models. The ablation analysis showed that step-by-step reasoning via chain-of-thought prompting is critical, whereas incorporating visual screenshots or overloading prompts with extensive action histories and descriptions frequently degraded performance or provided negligible benefit.
These results indicate that while foundation models are not yet ready for unattended, end-to-end automation in enterprise software, they are viable for guided assistance on well-defined operations like knowledge retrieval and catalog ordering. Deploying current models without human supervision introduces substantial operational risks, especially on tasks with complex forms or data table modifications. However, the rapid progress seen from older open-source models to current versions suggests that model capability is growing steadily.
Decision-makers should focus near-term investments on partial automation and supervised human-in-the-loop assistants rather than fully autonomous agents for enterprise workflows. Technology teams should prioritize developing agents with better long-context comprehension of web document structures and fine-tune vision models on user interface data before attempting mission-critical deployments.
The study's conclusions are constrained by an episode limit of 15 actions, a zero-shot prompting setup, and evaluation centered on a single software platform. While confidence is high that modern web agents struggle with non-standard enterprise interfaces, performance may improve under customized training regimes, larger action budgets, or domain-adapted models.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). WebArena establishes the realistic, multi-step web-agent benchmark approach that WorkArena adapts to enterprise knowledge-work tasks.
- Paper: Mind2Web: Towards a Generalist Agent for the Web, Xiang Deng et al. (2023). Mind2Web provides an earlier benchmark for generalist agents on real websites, clarifying the prior evaluation landscape WorkArena addresses.
No sufficiently relevant recommendations were found.
