Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents
Dongjun LeeJuyong LeeKyuyoung KimJihoon TackJinwoo ShinYee Whye TehKimin Lee
Introduces LCoW, a framework that decouples web page comprehension from action planning by training a specialized contextualization module, boosting LLM agent success rates on WorkArena by up to 23.7% and outperforming human experts on WebShop.
Automating routine web tasks using artificial intelligence has become an important priority across many industries seeking operational efficiency. However, even leading large language models struggle to navigate everyday websites accurately because real-world web pages contain dense, complex structures that obscure critical user interface elements and create confusing clutter.
The article evaluates a framework called LCoW (Learning to Contextualize Web Pages), designed to demonstrate how separating web page interpretation from task execution improves an artificial intelligence agent's decision-making accuracy.
The researchers developed an auxiliary language model that filters out irrelevant page information, highlights essential components, and explains interactive elements before passing the refined view to the main decision-making agent. To train this module without relying solely on manual prompt design, the authors used an iterative optimization process across hundreds of web tasks from benchmarks representing online shopping, enterprise workflows, and general web browsing. The module generates candidate page summaries, evaluates them based on whether multiple autonomous agents successfully predict the correct subsequent action, and refines the contextualizer using the highest-scoring examples.
The evaluation revealed substantial performance gains across models of varying sizes. First, adding the contextualization module increased task success rates on enterprise workflows by an average of 15.6% for premier proprietary models and 23.7% for open-source models, while raising an 8-billion-parameter open-source model from a 1.2% baseline success rate to 37.0%. Second, on an online retail benchmark, an agent powered by this framework achieved a 62.8% success rate, surpassing human expert performance of 59.6% and outperforming previous automation methods by more than 12 percentage points. Third, the system demonstrated successful transfer to unseen task types and external websites, generating a 4.3% improvement on completely novel sites by recognizing universal interface components like search fields and filters.
These findings indicate that the primary bottleneck in autonomous web navigation lies in processing cluttered page observations rather than in the underlying reasoning capabilities of the models. For organizations, adopting a specialized contextualization layer enables smaller, cost-effective open-source models to perform tasks previously achievable only by larger proprietary systems, while reducing repetitive actions and improving execution speed. It also provides a practical mechanism to steer closed-source commercial models without requiring costly fine-tuning of the primary decision-makers.
Organizations developing digital automation should consider deploying modular architectures that separate raw interface interpretation from core action planning. When deploying such agents, teams should test performance against target interface elements and consider using lightweight acceleration techniques, such as speculative decoding, to offset the computational latency introduced by the additional processing step.
While confidence in the framework's core performance gains is high across tested environments, the article notes key limitations: the contextualizer struggles to generalize to entirely unfamiliar categories that introduce novel interface mechanisms not covered in initial successful demonstration trajectories. Stakeholders should therefore pilot implementations on representative organizational workflows to ensure sufficient demonstration data exists before full-scale deployment.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). Introduces WebArena, the foundational realistic web environment and benchmark that established the difficulties LLM agents face with complex web structures.
- Paper: WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents, Shunyu Yao et al. (2022). Presents WebShop, a core benchmark directly targeted and evaluated in the source paper to test web decision-making agents.
- Paper: Mind2Web: Towards a Generalist Agent for the Web, Xiang Deng et al. (2023). Pioneers the two-stage paradigm of filtering and simplifying complex webpage structures for LLM agents via Mind2Web and MindAct.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Provides a comprehensive architectural taxonomy of LLM-based autonomous agents, detailing how perception and action modules interact.
- Paper: WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models, Hongliang He et al. (2024). Examines the challenge of transforming complex interactive webpage elements into comprehensible multimodal representations for autonomous agents.
- Paper: Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models, Andy Zhou et al. (2024). Explores how decision making and tree search can be integrated into LLM agents across environments like WebShop.
- Paper: WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration, Yao Zhang et al. (2025). Builds on web environment reasoning by employing multi-agent strategic exploration and reflection-guided tree search on WebArena.
- Paper: Agent Workflow Memory, Zora Zhiruo Wang et al. (2025). Extends web agent execution capabilities by inducing and reusing abstract procedural workflows across WebArena and Mind2Web tasks.
- Paper: Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic Environments, Hongjin Su et al. (2025). Develops a data-centric backward construction framework to adapt agents to dynamic, realistic software and web environments without human labeling.
- Paper: DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments, Yuxiang Zheng et al. (2025). Scales autonomous web navigation and research to live, dynamic web environments using outcome-driven reinforcement learning.
- Paper: Large Language Model-Brained GUI Agents: A Survey, Chaoyun Zhang et al. (2025). Surveys the broader ecosystem of LLM-brained GUI and web agents, contextualizing techniques for perception, grounding, and decision making.
- Paper: SafeArena: Evaluating the Safety of Autonomous Web Agents, Ada Defne Tur et al. (2025). Evaluates the safety and misuse vulnerabilities of capable web agents operating in realistic interactive web environments.
