Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents

Dongjun LeeJuyong LeeKyuyoung KimJihoon TackJinwoo ShinYee Whye TehKimin Lee

article2025ICLR23 citations

Introduces LCoW, a framework that decouples web page comprehension from action planning by training a specialized contextualization module, boosting LLM agent success rates on WorkArena by up to 23.7% and outperforming human experts on WebShop.

Listen

Automating routine web tasks using artificial intelligence has become an important priority across many industries seeking operational efficiency. However, even leading large language models struggle to navigate everyday websites accurately because real-world web pages contain dense, complex structures that obscure critical user interface elements and create confusing clutter.

The article evaluates a framework called LCoW (Learning to Contextualize Web Pages), designed to demonstrate how separating web page interpretation from task execution improves an artificial intelligence agent's decision-making accuracy.

The researchers developed an auxiliary language model that filters out irrelevant page information, highlights essential components, and explains interactive elements before passing the refined view to the main decision-making agent. To train this module without relying solely on manual prompt design, the authors used an iterative optimization process across hundreds of web tasks from benchmarks representing online shopping, enterprise workflows, and general web browsing. The module generates candidate page summaries, evaluates them based on whether multiple autonomous agents successfully predict the correct subsequent action, and refines the contextualizer using the highest-scoring examples.

The evaluation revealed substantial performance gains across models of varying sizes. First, adding the contextualization module increased task success rates on enterprise workflows by an average of 15.6% for premier proprietary models and 23.7% for open-source models, while raising an 8-billion-parameter open-source model from a 1.2% baseline success rate to 37.0%. Second, on an online retail benchmark, an agent powered by this framework achieved a 62.8% success rate, surpassing human expert performance of 59.6% and outperforming previous automation methods by more than 12 percentage points. Third, the system demonstrated successful transfer to unseen task types and external websites, generating a 4.3% improvement on completely novel sites by recognizing universal interface components like search fields and filters.

These findings indicate that the primary bottleneck in autonomous web navigation lies in processing cluttered page observations rather than in the underlying reasoning capabilities of the models. For organizations, adopting a specialized contextualization layer enables smaller, cost-effective open-source models to perform tasks previously achievable only by larger proprietary systems, while reducing repetitive actions and improving execution speed. It also provides a practical mechanism to steer closed-source commercial models without requiring costly fine-tuning of the primary decision-makers.

Organizations developing digital automation should consider deploying modular architectures that separate raw interface interpretation from core action planning. When deploying such agents, teams should test performance against target interface elements and consider using lightweight acceleration techniques, such as speculative decoding, to offset the computational latency introduced by the additional processing step.

While confidence in the framework's core performance gains is high across tested environments, the article notes key limitations: the contextualizer struggles to generalize to entirely unfamiliar categories that introduce novel interface mechanisms not covered in initial successful demonstration trajectories. Stakeholders should therefore pilot implementations on representative organizational workflows to ensure sufficient demonstration data exists before full-scale deployment.

Cover for Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents

Abstract

Recent advances in large language models (LLMs) have led to a growing interest in developing LLM-based agents for automating web tasks. However, these agents often struggle with even simple tasks on real-world websites due to their limited capability to understand and process complex web page structures. In this work, we introduce LCoW, a framework for Learning language models to Contextualize complex Web pages into a more comprehensible form, thereby enhancing decision making by LLM agents. LCoW decouples web page understanding from decision making by training a separate contextualization module to transform complex web pages into comprehensible format, which are then utilized by the decision-making agent. We demonstrate that our contextualization module effectively integrates with LLM agents of various scales to significantly enhance their decision-making capabilities in web automation tasks. Notably, LCoW improves the success rates of closed-source LLMs (e.g., Gemini-1.5-flash, GPT-4o, Claude-3.5-Sonnet) by an average of 15.6%, and demonstrates a 23.7% average improvement in success rates for open-source LMs (e.g., Llama-3.1-8B, Llama-3.1-70B) on the WorkArena benchmark. Moreover, the Gemini-1.5-flash agent with LCoW achieves state-of-the-art results on the WebShop benchmark, outperforming human experts. The relevant code materials are available at our project page: this https URL.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Method
  • 3.1 Contextualization module
  • 3.2 Algorithm for training the contextualization module
  • 4 Experiments
  • 4.1 Benchmarks & Evaluation setup
  • 4.2 Main results
  • 4.3 Analysis
  • 5 Related work
  • 6 Limitations and future directions
  • 7 Conclusion
  • References
  • A Qualitative analysis
  • A.1 Examples of contextualized Web Pages
  • A.2 Analysis on decision making process
  • B Experimental Details & Additional Experiments
  • B.1 Main experiment setup
  • B.2 Generalization experiment setup
  • B.3 Additional generalization experiment results
  • C Prompts
  • C.1 WorkArena & WebArena
  • C.2 WebShop
  • C.3 Action matching evaluation prompt

Knowls

  1. Knowl 1 — LCoW Architecture for Decoupled Web Observation Contextualization

    model/method

    Learning to Contextualize complex Web pages (LCoW) is a framework that decouples web page perception/understanding from sequential action decision-making in web automation tasks.

    In standard language-agent web browsing, the agent policy π\pi directly receives lengthy, complex web observations (such as raw HTML or accessibility trees oto_t) alongside the task instruction [TASK][\text{TASK}] and action history a<t=(a1,…,at−1)a_{<t} = (a_1, \dots, a_{t-1}) to select action ata_t. LCoW introduces an intermediate contextualization language model fθf_\theta. At each time step tt, fθf_\theta transforms the raw observation oto_t into a refined, concise observation otcoo_t^{co} containing filtered UI element hierarchies and verbal explanations of element functionalities:

    otco=fθ([TASK],a<t,ot)o_t^{co} = f_\theta([\text{TASK}], a_{<t}, o_t)

    The downstream decision-making agent π\pi (which can be any closed-source or open-source LLM) then predicts the next action conditioned on otcoo_t^{co}:

    at=π([TASK],a<t,otco)a_t = \pi([\text{TASK}], a_{<t}, o_t^{co})

  2. Knowl 2 — Iterative Training Algorithm for Observation Contextualization Modules

    algorithm

    The LCoW training procedure iteratively optimizes the parameters θ\theta of a contextualization language model fθf_\theta through trajectory rollouts, consensus-based candidate sampling, and supervised fine-tuning across MM iterations.

    Input: Initial contextualizer fθ(0)f_{\theta}^{(0)}, policy agent π\pi, ensemble of reward evaluator agents Π={πk}k=1K\Pi = \{\pi_k\}_{k=1}^K, training task set GtrG_{tr}, number of candidate samples NN, trajectory buffer T\mathcal{T}, empty dataset buffer D\mathcal{D}
    Output: Trained contextualization module fθ(M)f_{\theta}^{(M)}
    for iteration i=0i = 0 to M−1M-1 do
        // Step 1: Trajectory collection
        for each [TASK] in GtrG_{tr} do
            Roll out episode τ=(o1,a1,…,oT,aT)\tau = (o_1, a_1, \dots, o_T, a_T) and terminal reward R∼π(⋅∣[TASK],fθ(i))R \sim \pi(\cdot \mid [\text{TASK}], f_{\theta}^{(i)})
            if R==1.0R == 1.0 then
                T\mathcal{T}.append(τ\tau)
            end if
        end for
        // Step 2: Sampling optimal contextualizations
        for each step (ot,at)(o_t, a_t) in T\mathcal{T} do
            for n=1n = 1 to NN do
                Sample candidate ot,nco∼fθ(i)(⋅∣[TASK],a<t,ot)o_{t,n}^{co} \sim f_{\theta}^{(i)}(\cdot \mid [\text{TASK}], a_{<t}, o_t)
                Compute reward rt,n=∑π′∈ΠActionMatchingScore(π′([TASK],a<t,ot,nco),at)r_{t,n} = \sum_{\pi' \in \Pi} \text{ActionMatchingScore}(\pi'([\text{TASK}], a_{<t}, o_{t,n}^{co}), a_t)
            end for
            if max⁡n(rt,n)==0\max_n(r_{t,n}) == 0 then
                // Retry with ground-truth action context
                for n=1n = 1 to NN do
                    Sample candidate ot,nco∼fθ(i)(⋅∣[TASK],a<t,ot,at)o_{t,n}^{co} \sim f_{\theta}^{(i)}(\cdot \mid [\text{TASK}], a_{<t}, o_t, a_t)
                    Compute reward rt,n=∑π′∈ΠActionMatchingScore(π′([TASK],a<t,ot,nco),at)r_{t,n} = \sum_{\pi' \in \Pi} \text{ActionMatchingScore}(\pi'([\text{TASK}], a_{<t}, o_{t,n}^{co}), a_t)
                end for
            end if
            Select target ot,∗co=arg⁡max⁡ot,ncort,no_{t,*}^{co} = \arg\max_{o_{t,n}^{co}} r_{t,n}
            D\mathcal{D}.append((([TASK],a<t,ot),ot,∗co)(([\text{TASK}], a_{<t}, o_t), o_{t,*}^{co}))
        end for
        // Step 3: Supervised fine-tuning parameter update
        θ(i+1)←arg⁡max⁡θE([TASK],a<t,ot,ot,∗co)∼D[log⁡fθ(ot,∗co∣[TASK],a<t,ot)]\theta^{(i+1)} \leftarrow \arg\max_\theta \mathbb{E}_{([\text{TASK}], a_{<t}, o_t, o_{t,*}^{co}) \sim \mathcal{D}} [\log f_\theta(o_{t,*}^{co} \mid [\text{TASK}], a_{<t}, o_t)]
    end for
    return fθ(M)f_{\theta}^{(M)}
  3. Knowl 3 — Action-Matching Consensus Reward for Observation Selection

    model/method

    In LCoW, candidate contextualized observations ot,ncoo_{t,n}^{co} are evaluated using a multi-agent action-matching reward. Rather than optimizing the contextualizer for a single decision model, the candidate is fed to an ensemble of KK distinct LLM agents Π={π1,π2,…,πK}\Pi = \{\pi_1, \pi_2, \dots, \pi_K\}.

    The action-matching reward rt,nr_{t,n} is defined as:

    rt,n=∑π∈ΠActionMatchingScore(π([TASK],a<t,ot,nco),at)r_{t,n} = \sum_{\pi \in \Pi} \text{ActionMatchingScore}(\pi([\text{TASK}], a_{<t}, o_{t,n}^{co}), a_t)

    where ata_t is the ground-truth action from a successful trajectory, and ActionMatchingScore(a^,at)∈{0,1}\text{ActionMatchingScore}(\hat{a}, a_t) \in \{0, 1\} determines whether the action predicted by agent π\pi matches the ground-truth action ata_t. In open-ended web environments where exact string parsing is insufficient, the matching score is evaluated using an LLM evaluator (e.g., GPT-4o) prompted to judge semantic equivalence.

    Optimizing against this multi-agent consensus reward prevents fθf_\theta from overfitting to idiosyncratic biases of any individual agent, enabling the produced observations to generalize across arbitrary open- and closed-source LLM backbones.

  4. Knowl 4 — Evaluation on WebShop Benchmark

    data/table

    On the WebShop shopping simulation benchmark (500 evaluation tasks), LCoW iteratively improves both task success rates (SR, %) and average reward scores (Rew ∈[0,1]\in [0, 1]) across multiple LLM backbones. Fine-tuning Phi-3-mini-Instruct as the contextualization module over three iterations enables Gemini-1.5-flash to achieve a 62.8% success rate, exceeding the expert human benchmark of 59.6%.

    Method GPT-4o Gemini-1.5-flash Claude-3.5-Sonnet Llama-3.1-70B (Unseen)
    SR (%) Rew SR (%) Rew SR (%) Rew SR (%) Rew
    Raw observation 34.8 0.496 43.6 0.693 26.6 0.336 34.2 0.590
    Self-ctx 26.2 0.459 46.4 0.608 12.4 0.146 40.2 0.547
    LCoW (iter 1) 27.8 0.545 46.4 0.705 39.4 0.600 39.2 0.666
    LCoW (iter 2) 46.0 0.647 58.2 0.796 58.8 0.780 55.0 0.781
    LCoW (iter 3) 50.6 0.666 62.8 0.803 59.8 0.771 59.6 0.803

    For comparison, prior methods achieve: WebN-T5: 29.8%, ASH: 30.2%, ReAct: 40.0%, WebGUM: 45.0%, LASER: 50.0%, AgentQ: 50.5%, Average Human: 50.0%, Human Expert: 59.6%.

  5. Knowl 5 — Evaluation on WorkArena Enterprise Benchmark

    data/table

    On the WorkArena enterprise web navigation benchmark (evaluated across 165 task instances covering 33 task types), contextualizing observations using a Llama-3.1-8B-Instruct model trained for one iteration of LCoW consistently improves success rates across all agent sizes, including models never used in the training reward ensemble.

    Method GPT-4o Gemini-1.5-flash Claude-3.5-Sonnet Llama-3.1-70B (Unseen) Llama-3.1-8B (Unseen)
    Raw observation 38.2% 11.5% 44.8% 26.1% 1.2%
    Self-ctx 43.0% 12.7% 50.3% 29.1% 7.3%
    LCoW (iter 1) 44.2% 41.2% 55.8% 40.0% 37.0%

    The largest absolute improvement occurs on the smallest agent (Llama-3.1-8B), which rises from 1.2% on raw accessibility tree observations to 37.0% with LCoW contextualization.

  6. Knowl 6 — Observation Contextualization vs. Direct Policy Behavior Cloning

    empirical result

    Using demonstration trajectories to train a web page contextualization module yields substantially higher task success rates than using the same demonstrations to directly fine-tune an agent policy via behavior cloning (BC).

    When Llama-3.1-8B is fine-tuned directly on 264 seed demonstrations from WorkArena via behavior cloning, it achieves a task success rate of 23.6%. In contrast, when the same 264 seed demonstrations are used to train a Llama-3.1-8B contextualization module via LCoW, pairing that contextualization module with a zero-shot Llama-3.1-8B decision agent achieves a 37.0% success rate (+13.4% over BC). When the same contextualizer is paired with larger LLM agents (Claude-3.5-Sonnet, GPT-4o, Llama-3.1-70B), success rates reach 40.0% to 55.8%.

  7. Knowl 7 — Generalization Across Unseen Task Types and Unseen Websites

    empirical result

    LCoW contextualization modules exhibit zero-shot transfer capabilities across task templates and unseen websites:

    1. Unseen Task Types in WebArena: When evaluated on 165 WebArena tasks split into 117 seen-template tasks and 48 unseen-template tasks, a GPT-4o agent using LCoW improves performance over raw observations from 35.9% to 41.9% on seen types, and from 14.6% to 20.8% on unseen types (+6.2% improvement).
    2. Unseen Task Types in WorkArena: On 70 unseen-type tasks, LCoW improves GPT-4o success rate from 35.7% to 42.9% (+7.2%) and Gemini-1.5-flash success rate from 14.5% to 37.1% (+22.6%).
    3. Unseen Website Transfer in WebArena: When the contextualization module is trained exclusively on 5 websites (GitLab, Reddit, Map, CMS, Wikipedia) and evaluated on the held-out Shopping website, GPT-4o agent performance increases from 17.4% to 21.7% (+4.3%). Shared UI components (such as search boxes and date pickers) allow transferable semantic grounding across domains.
  8. Knowl 8 — Generalization Breakdown on Unseen UI Categories in LCoW

    limitation

    LCoW fails to improve agent performance when evaluated on task categories whose fundamental UI interaction mechanisms were completely absent from the training set.

    In the WorkArena benchmark, when evaluating on the 30 tasks of the Filter-List category as an unseen-category test (where training tasks contained no filter-related dropdown elements), both raw observation baselines and LCoW contextualization achieved 0.0% success rate on GPT-4o and Gemini-1.5-flash backbones.

    This breakdown occurs because the contextualization module omits the hidden dropdown UI elements required to manipulate list filters from its output. Because the model did not encounter filter functionality during training, it fails to identify and extract the necessary interaction elements.

  9. Knowl 9 — Reduction in Agent Decision Steps and Extraneous Actions

    empirical result

    LCoW contextualization reduces the total number of interaction steps required by LLM agents to complete tasks by eliminating inadmissible actions and redundant exploration.

    Qualitative and rollout step distributions demonstrate that raw observations frequently cause agents to repeatedly scroll, attempt invalid click sequences on un-interactable UI containers, or execute unnecessary pre-clearing commands (e.g., calling clear() on already empty input fields). By providing explicit text explanations of UI state (e.g., confirming a field is already empty or specifying the exact button ID for sorting/searching), LCoW shifts the distribution of step counts downward for successful task executions on both Claude-3.5-Sonnet and Llama-3.1-70B agents.

Coverage note — Omitted verbatim prompt templates from Appendix C and exhaustive seed demonstration counts per individual WorkArena sub-task (Table 5/Table 6) as they are implementation details that support the main methods and results.

References

  1. 1.Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems, 36, 2024.
  2. 2.Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H Laradji, Manuel Del Verme, Tom Marty, Leo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, et al. Workarena: How capable are web agents at solving common knowledge work tasks? arXiv preprint arXiv:2403.07718, 2024.
  3. 3.Hiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo, Aleksandra Faust, Shixiang Shane Gu, and Izzeddin Gur. Multimodal web navigation with instruction-finetuned foundation models. arXiv preprint arXiv:2305.11854, 2023.
  4. 4.Izzeddin Gur, Ofir Nachum, Yingjie Miao, Mustafa Safdari, Austin Huang, Aakanksha Chowdhery, Sharan Narang, Noah Fiedel, and Aleksandra Faust. Understanding html with large language models, 2023.
  5. 5.Izzeddin Gur, Hiroki Furuta, Austin Huang, Mustafa Safdari, Yutaka Matsuo, Douglas Eck, and Aleksandra Faust. A real-world webagent with planning, long context understanding, and program synthesis. In International Conference on Learning Representations, 2024.
  6. 6.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213, 2022.
  7. 7.Hanyu Lai, Xiao Liu, Iat Long Iong, Shuntian Yao, Yuxuan Chen, Pengbo Shen, Hao Yu, Hanchen Zhang, Xiaohan Zhang, Yuxiao Dong, et al. Autowebglm: Bootstrap and reinforce a large language model-based web navigating agent. arXiv preprint arXiv:2404.03648, 2024.
  8. 8.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. In International Conference on Learning Representations, 2024.
  9. 9.Xiao Liu, Tianjie Zhang, Yu Gu, Iat Long Iong, Yifan Xu, Xixuan Song, Shudan Zhang, Hanyu Lai, Xinyi Liu, Hanlin Zhao, et al. Visualagentbench: Towards large multimodal models as visual foundation agents. arXiv preprint arXiv:2408.06327, 2024.
  10. 10.Kaixin Ma, Hongming Zhang, Hongwei Wang, Xiaoman Pan, Wenhao Yu, and Dong Yu. Laser: Llm agent with state-space exploration for web navigation, 2024.
  11. 11.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. Self-refine: Iterative refinement with self-feedback, 2023.
  12. 12.Oscar Manas, Pietro Astolfi, Melissa Hall, Candace Ross, Jack Urbanek, Adina Williams, Aish- ˜warya Agrawal, Adriana Romero-Soriano, and Michal Drozdzal. Improving text-to-image consistency via automatic prompt optimization. arXiv preprint arXiv:2403.17804, 2024.
  13. 13.OpenAI. https://openai.com/index/hello-gpt-4o/, 2024.
  14. 14.Jiayi Pan, Yichi Zhang, Nicholas Tomlin, Yifei Zhou, Sergey Levine, and Alane Suhr. Autonomous evaluation and refinement of digital agents. In First Conference on Language Modeling, 2024.
  15. 15.Pranav Putta, Edmund Mills, Naman Garg, Sumeet Motwani, Chelsea Finn, Divyansh Garg, and Rafael Rafailov. Agent q: Advanced reasoning and learning for autonomous ai agents. arXiv preprint arXiv:2408.07199, 2024.
  16. 16.Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, 2020.
  17. 17.Paloma Sodhi, S. R. K. Branavan, Yoav Artzi, and Ryan McDonald. Step: Stacked llm policies for web actions, 2024.
  18. 18.Abishek Sridhar, Robert Lo, Frank F. Xu, Hao Zhu, and Shuyan Zhou. Hierarchical prompting assists large language model on web navigation, 2023.
  19. 19.Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried, and Graham Neubig. Agent workflow memory, 2024.
  20. 20.Fangyuan Xu, Weijia Shi, and Eunsol Choi. Recomp: Improving retrieval-augmented lms with compression and selective augmentation. In International Conference on Learning Representations, 2024.
  21. 21.Ke Yang, Yao Liu, Sapana Chaudhary, Rasool Fakoor, Pratik Chaudhari, George Karypis, and Huzefa Rangwala. Agentoccam: A simple yet strong baseline for llm-based web agents. arXiv preprint arXiv:2410.13825, 2024.
  22. 22.Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. Webshop: Towards scalable real-world web interaction with grounded language agents. Advances in Neural Information Processing Systems, 35:20744–20757, 2022a.
  23. 23.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629, 2022b.
  24. 24.Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al. Webarena: A realistic web environment for building autonomous agents. arXiv preprint arXiv:2307.13854, 2023a.
  25. 25.Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. Large language models are human-level prompt engineers. In International Conference on Learning Representations, 2023b.

Citation

MLA
Lee, D., et al. “Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents”. arXiv, 2025, http://arxiv.org/abs/2503.10689v2.
APA
Lee, D., Lee, J., Kim, K., Tack, J., Shin, J., Teh, Y. W., & Lee, K. (2025). Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents. arXiv. http://arxiv.org/abs/2503.10689v2
Chicago
Lee, D., J. Lee, K. Kim, et al. 2025. “Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents”. arXiv. http://arxiv.org/abs/2503.10689v2.
Harvard
Lee, D. et al. (2025) “Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2503.10689v2.
Vancouver
1. Lee D, Lee J, Kim K, Tack J, Shin J, Teh YW, Lee K (2025) Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents. arXiv

BibTeX

@article{lee2025learning,
  title = {Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents},
  author = {Lee, Dongjun and Lee, Juyong and Kim, Kyuyoung and Tack, Jihoon and Shin, Jinwoo and Teh, Yee Whye and Lee, Kimin},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2503.10689v2},
  eprint = {2503.10689}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/