From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces

Peter ShawMandar JoshiJames CohanJonathan BerantPanupong PasupatHexiang HuUrvashi KhandelwalKenton LeeKristina Toutanova

article2023NeurIPS88 citations

Introduces Pix2Act, a visual web agent that relies strictly on raw screenshots and generic mouse and keyboard actions to complete instruction-following tasks, demonstrating for the first time that pixel-only models can outperform human crowdworkers on MiniWob++.

Listen

Digital agents designed to complete tasks via graphical user interfaces (GUIs) have traditionally depended on structured underlying text, such as Document Object Model (DOM) trees or HTML code, paired with specialized, task-specific actions. However, these structured representations are frequently unavailable, heavily scripted, or misaligned with actual visual displays. Humans operate across diverse digital interfaces without accessing underlying code, using only visual perception and standard inputs like mouse clicks and keystrokes. Developing agents capable of navigating software using purely visual inputs and generic controls is essential for broader automation, accessibility, and digital assistance across diverse, real-world computing environments.

The article evaluates whether an automated agent can successfully interpret and complete multi-step digital tasks using only raw, pixel-level screenshot observations and low-level mouse and keyboard actions. It demonstrates the feasibility and effectiveness of this approach across standard web-based benchmark environments.

To accomplish this, the authors introduced an agent named PIX2ACT, built on a 282-million parameter Vision Transformer architecture pre-trained to parse web screenshots. The system receives visual screen captures with rendered instructions and generates textual tokens corresponding to discrete mouse movements, clicks, drags, scrolling, and keyboard actions. The training methodology combined behavioral cloning on human demonstrations with policy refinement using Monte Carlo Tree Search. The evaluation was conducted within a browser framework across 59 tasks from the MiniWob++ benchmark and an adapted version of the WebShop shopping benchmark.

The findings show that PIX2ACT achieves an average score of 96.2 out of 100 on the MiniWob++ benchmark, outperforming human crowdworkers (94.4) and matching state-of-the-art models (96.3) that have direct access to internal DOM structures. Pre-training on screenshot parsing proved vital; without it, performance dropped sharply from 66.5 to 17.1 on MiniWob++ and from 46.7 to 1.1 on WebShop under behavioral cloning alone. Furthermore, policy refinement through tree search improved the agent's greedy policy performance from 66.5 to 96.2. On the WebShop benchmark, PIX2ACT established the first visual-only baseline score of 46.7, though a gap remains compared to language models utilizing HTML code (67.5). Finally, the agent demonstrated zero-shot transfer capability, achieving a score of 28.3 on completely unseen interface tasks compared to 7.6 without pre-training.

These results establish that automated agents do not require internal software code to operate digital interfaces effectively. By interacting visually, systems can bypass technical hurdles like code obfuscation and sandboxing, enabling universal compatibility across software ecosystems. However, closing the remaining performance gap on complex, text-heavy tasks will likely require larger multimodal models or enhanced visual scaling comparable to recent text-based language models.

Before deploying pixel-based agents in real-world online services, organizations must address safety, abuse, and compliance risks. Because visual agents interact like human users, they could potentially bypass standard security defenses, create spam, or interact improperly with external platforms. Deployments require rigorous behavioral guardrails, adherence to terms of service, and protections against privacy risks associated with processing screen captures. Further research should focus on refining reward modeling for environments where deterministic task resets are unavailable.

Confidence in these findings is high for structured, deterministic browser tasks, as evidenced by consistent performance across extensive evaluation seeds. However, key limitations remain: the current system excludes tasks requiring complex real-time animations or specific interactions like manual text-highlighting drag actions. Additionally, the policy improvement mechanism relied on deterministic environment resets and automated reward signals, which may not readily exist in arbitrary production applications without supplementary reward modeling.

arXiv: 2306.00245
Cover for From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces

Abstract

Much of the previous work towards digital agents for graphical user interfaces (GUIs) has relied on text-based representations (derived from HTML or other structured data sources), which are not always readily available. These input representations have been often coupled with custom, task-specific action spaces. This paper focuses on creating agents that interact with the digital world using the same conceptual interface that humans commonly use — via pixel-based screenshots and a generic action space corresponding to keyboard and mouse actions. Building upon recent progress in pixel-based pretraining, we show, for the first time, that it is possible for such agents to outperform human crowdworkers on the MiniWob++ benchmark of GUI-based instruction following tasks.

Table of Contents

  • 1 Introduction
  • 2 Environment
  • 3 Proposed Agent
  • 3.1 Training
  • 4 Benchmarks and Demonstrations
  • 4.1 MiniWob++
  • 4.2 WebShop
  • 5 Experiments and Analysis
  • 5.1 Training Details
  • 5.2 Main Results
  • 5.3 Ablations and Analysis
  • 6 Related Work
  • 7 Limitations and Discussion
  • Acknowledgments
  • References
  • A Additional Dataset Details
  • A.1 MiniWob++ Supported Tasks
  • A.2 MiniWob++ Rendering Differences
  • B Additional Technical Details
  • B.1 Beam Search
  • B.2 Tree Search
  • B.3 Compute Details
  • C Additional Results
  • C.1 Variance Estimates
  • C.2 MiniWob++ Results Per Task

Knowls

  1. Knowl 1 — PIX2ACT Model Architecture and Observation Pipeline

    model/method

    PIX2ACT is a multimodal agent that interacts with graphical user interfaces (GUIs) using raw pixel observations and a generic keyboard/mouse action space, without accessing structured Document Object Model (DOM) trees or HTML code.

    The model is initialized from the base variant of PIX2STRUCT (282 million parameters: 12 encoder layers, 12 decoder layers, hidden dimension 768), which is pretrained on web screenshots to predict visually situated HTML markup. The observation at time step tt is an RGB screenshot of the browser window processed with three visual augmentations:

    1. The natural language instruction is rendered onto the top of the image canvas if not already visible.
    2. A mouse pointer cursor icon is rendered at the current cursor coordinates (x,y)(x, y) because standard headless browser captures omit operating system cursors.
    3. An indicator of the mouse button state (pressed down versus released) is rendered to support dragging interactions.

    The augmented image is dynamically scaled to extract the maximum number of fixed-size image patches that fit within the model's sequence length limit (input resolutions are 160×210160 \times 210 pixels for MiniWob++ and 800×600800 \times 600 pixels for WebShop). The image Transformer encoder encodes the visual patches, and the autoregressive text Transformer decoder predicts low-level action sequences as text tokens.

  2. Knowl 2 — Low-Level Action Space for Pixel-Based GUI Interaction

    definition

    The action space in PIX2ACT operates on basic mouse and keyboard commands parameterized by discrete spatial and token arguments:

    • click X Y: Moves the mouse pointer to discrete horizontal bin XX and discrete vertical bin YY, executes a mouse press, and releases.
    • begin_drag X Y: Moves the pointer to coordinate (X,Y)(X, Y) and presses and holds the primary mouse button down.
    • end_drag X Y: Moves the pointer to coordinate (X,Y)(X, Y) with the button held, releases the primary mouse button, and completes a drag-and-drop operation.
    • key [M] K: Presses key or string sequence KK, optionally modified by modifier key MM (e.g., shift).
    • scroll Z: Scrolls the viewport vertically by a discrete displacement bin ZZ.

    Coordinates XX and YY are discretized into uniform spatial buckets per axis (for example, 32 bins along each axis in MiniWob++). Actions are emitted autoregressively as text strings and decoded into environment execution primitives.

  3. Knowl 3 — Greedy Action Decoding with Cycle Prevention

    model/method

    At inference time, the action sequence is decoded autoregressively from the text decoder using beam search. For an observation screenshot ss, beam search yields the top-kk candidate action strings with normalized log-probabilities using a length normalization exponent of 0.60.6 (with k=8k=8 for MiniWob++ and k=10k=10 for WebShop).

    To prevent the agent from becoming trapped in infinite operational loops (such as repeatedly clicking an unclickable element or oscillating between two states), PIX2ACT maintains an execution history buffer that tracks actions previously taken for each distinct observation state ss. The greedy policy πθ(s)\pi_\theta(s) selects the highest-probability candidate action from the beam that has not been executed previously from state ss during the current episode.

  4. Knowl 4 — Surrogate Step Penalty Reward and Hybrid Leaf State Evaluation

    equation

    To train reinforcement learning value networks and guide lookahead tree search in GUI environments, PIX2ACT defines a surrogate reward r(s)r(s) and a hybrid leaf evaluation v~(sL)\tilde{v}(s_L).

    The surrogate reward penalizes trajectory length to encourage efficient completion without redundant steps:

    r(s)=αS+rt(s)r(s) = \alpha_S + r^t(s)

    where αS=−130\alpha_S = -\frac{1}{30} is a per-step penalty, and rt(s)r^t(s) is defined as the raw terminal reward rraw(s)r_{\text{raw}}(s) if ss is a terminal state and rraw(s)>0.8r_{\text{raw}}(s) > 0.8, and 00 otherwise.

    When Monte Carlo Tree Search reaches a non-terminal leaf state sL=f(st,a)s_L = f(s_t, a), its estimated value v~(sL)\tilde{v}(s_L) is computed as a convex combination of a learned value network and a greedy policy rollout:

    v~(sL)=λv^ϕ(sL)+(1−λ)vπθ(sL)\tilde{v}(s_L) = \lambda \hat{v}_\phi(s_L) + (1 - \lambda) v^{\pi_\theta}(s_L)

    where v^ϕ(sL)\hat{v}_\phi(s_L) is the value predicted by a PIX2STRUCT-initialized value network (mapping returns to discrete integer tokens, approximated via top-3 probability-weighted beam search), vπθ(sL)v^{\pi_\theta}(s_L) is the accumulated surrogate return from a greedy rollout following policy πθ\pi_\theta for up to 20 steps (clipped to a minimum of 00), and λ=0.1\lambda = 0.1 is the mixing weight.

  5. Knowl 5 — Monte Carlo Tree Search for GUI Policy Execution

    algorithm

    PIX2ACT uses Monte Carlo Tree Search (MCTS) with neural network priors and value estimation to plan actions in deterministic web environments.

    Input: Root environment state ss, policy network πθ\pi_\theta, value network v^ϕ\hat{v}_\phi, search rounds K=16K = 16, exploration weight c=0.1c = 0.1, beam size k=8k = 8, mix weight λ=0.1\lambda = 0.1, step penalty αS=−1/30\alpha_S = -1/30
    Output: Selected action a∗a^*
    Initialize root node with state ss, visit count N(s)=0N(s) = 0
    for round =1= 1 to KK do
        st←ss_t \leftarrow s
        path ←[]\leftarrow []
        while sts_t is in search tree and not terminal do
            Identify candidate actions Atop-k=top-k(πθ(a∣st))\mathcal{A}_{\text{top-}k} = \text{top-k}(\pi_\theta(a \mid s_t))
            for a∈Atop-ka \in \mathcal{A}_{\text{top-}k} do
                U(st,a)←c⋅pθ(a∣st)⋅N(st)1+n(st,a)U(s_t, a) \leftarrow c \cdot p_\theta(a \mid s_t) \cdot \frac{\sqrt{N(s_t)}}{1 + n(s_t, a)}
                Score(st,a)←Q(st,a)+U(st,a)(s_t, a) \leftarrow Q(s_t, a) + U(s_t, a)
            end for
            at←arg⁡max⁡a∈Atop-kScore(st,a)a_t \leftarrow \arg\max_{a \in \mathcal{A}_{\text{top-}k}} \text{Score}(s_t, a)
            Append (st,at)(s_t, a_t) to path
            if edge (st,at)(s_t, a_t) not visited before then
                sL←f(st,at)s_L \leftarrow f(s_t, a_t)
                Add sLs_L to search tree
                if sLs_L is terminal then
                    R←r(sL)R \leftarrow r(s_L)
                else
                    R←λv^ϕ(sL)+(1−λ)vπθ(sL)R \leftarrow \lambda \hat{v}_\phi(s_L) + (1 - \lambda) v^{\pi_\theta}(s_L)
                end if
                Q(sL,a)←R+αSQ(s_L, a) \leftarrow R + \alpha_S for all a∈Aa \in \mathcal{A}
                Break to backup phase
            end if
            st←f(st,at)s_t \leftarrow f(s_t, a_t)
        end while
        for each (s′,a′)(s', a') in path in reverse order do
            n(s′,a′)←n(s′,a′)+1n(s', a') \leftarrow n(s', a') + 1
            N(s′)←N(s′)+1N(s') \leftarrow N(s') + 1
            Q(s′,a′)←Q(s′,a′)+R−Q(s′,a′)n(s′,a′)Q(s', a') \leftarrow Q(s', a') + \frac{R - Q(s', a')}{n(s', a')}
        end for
    end for
    return a∗=arg⁡max⁡an(s,a)a^* = \arg\max_{a} n(s, a)
  6. Knowl 6 — Iterative Policy Improvement via MCTS Trajectory Distillation

    model/method

    To improve upon behavioral cloning without unstable reinforcement learning updates, PIX2ACT iteratively distills search trajectories into the policy network:

    1. Starting with an initial policy network πθ\pi_\theta trained via behavioral cloning on human demonstrations, construct the MCTS search policy πθ∗\pi^*_\theta.
    2. Execute πθ∗\pi^*_\theta across procedurally generated tasks and random environment seeds to collect full interaction trajectories.
    3. Filter out unsuccessful episodes by retaining only trajectories achieving a raw reward rraw≥0.8r_{\text{raw}} \ge 0.8.
    4. Fine-tune the policy network parameters θ\theta via standard supervised maximum likelihood on the filtered action trajectories generated by πθ∗\pi^*_\theta. The value network parameters ϕ\phi remain fixed.

    The improved policy network πθ\pi_\theta provides better action priors pθ(a∣s)p_\theta(a \mid s) and rollout evaluations for subsequent MCTS iterations, yielding higher-quality trajectories.

  7. Knowl 7 — PIX2ACT Task Performance on MiniWob++ and WebShop

    empirical result

    PIX2ACT achieves state-of-the-art results among pixel-only agents and matches human crowdworkers on GUI instruction following benchmarks.

    On MiniWob++ (evaluated across 59 tasks, 100 seeds per task, raw reward mapped to [0,100][0, 100]):

    • PIX2ACT (Behavioral Cloning + 2 iterations of MCTS policy improvement): achieves an average score of 96.2±0.1596.2 \pm 0.15.
    • Human Crowdworkers: achieve an average score of 94.494.4.
    • Prior pixel-only baseline CC-Net (Humphreys et al., 2022) without DOM inputs: achieves an estimated average score of 24.124.1.
    • CC-Net with DOM access (trained with BC + RL): achieves 96.396.3.
    • PIX2ACT with Behavioral Cloning only: achieves 66.566.5, outperforming CC-Net with BC only (38.738.7).

    On WebShop (1,566 demonstrations, evaluated on 500 test instructions):

    • PIX2ACT achieves a Task Score of 46.746.7, establishing the first pixel-only, low-level action baseline.
    • Text/HTML-based model WebGUM (using Flan-T5-XL and custom high-level domain actions) achieves 67.567.5.
  8. Knowl 8 — Ablation on the Necessity of Visual Pretraining for GUI Agents

    empirical result

    Pretraining on web screenshot parsing via PIX2STRUCT is essential for GUI instruction following from pixels alone under behavioral cloning:

    • On the 59 MiniWob++ tasks, the behavioral cloning model initialized from PIX2STRUCT achieves an average score of 66.566.5, whereas the identical architecture initialized randomly without pretraining achieves only 17.117.1.
    • On WebShop, the behavioral cloning model initialized from PIX2STRUCT achieves a Task Score of 46.746.7, whereas random initialization yields 1.11.1.

    These results demonstrate that learning to visually parse webpage layouts and renderable text provides the representation needed to map raw screenshots directly to interface coordinates and keyboard actions.

  9. Knowl 9 — MiniWob++ Policy Score Progression Across MCTS Iterations

    data/table

    Iterative training on MCTS trajectories improves both the search policy and the distilled greedy policy over successive rounds on MiniWob++:

    Policy Iteration 0 Iteration 1 Iteration 2
    Greedy Policy 66.5 93.1 96.2
    Tree Search Policy 91.7 98.4 —

    At Iteration 0 (model trained only on human demonstrations), MCTS lookahead improves performance from 66.566.5 (greedy) to 91.791.7. Training on 826K826\text{K} episodes collected by this search policy produces Iteration 1, where the greedy policy reaches 93.193.1 and tree search reaches 98.498.4. A second round of distillation (Iteration 2) brings the standalone greedy policy to 96.296.2, matching human-level performance at inference time without requiring test-time search or reward feedback.

  10. Knowl 10 — Leaf Evaluation Ablation in MCTS on Challenging Tasks

    data/table

    The effectiveness of hybrid leaf state evaluation in MCTS is evaluated across 12 challenging MiniWob++ tasks (including number-checkboxes, which require longer interaction trajectories), comparing rollout-only, value-network-only, and combined valuation:

    Greedy Policy Rollout Only (λ=0\lambda = 0) Hybrid (λ=0.1\lambda = 0.1) Value Only (λ=1\lambda = 1)
    28.8 74.2 78.3 57.4

    Evaluating leaf states purely with greedy rollouts (λ=0\lambda = 0) achieves 74.274.2, but is prone to uninformative terminal failures on hard tasks. Pure value estimation (λ=1.0\lambda = 1.0) achieves 57.457.4 due to out-of-distribution value inaccuracies. The hybrid combination (λ=0.1\lambda = 0.1) achieves the best performance (78.378.3), demonstrating that combining lightweight neural value estimation with environment rollouts improves search quality.

  11. Knowl 11 — Cross-Task Generalization and Pretraining Transfer

    data/table

    Visual pretraining enables PIX2ACT to generalize to entirely unseen GUI tasks:

    Initialization Included in Training Held-out Tasks
    PIX2STRUCT Pre-trained 65.5 28.3
    Random Initialization 11.0 7.6

    When evaluated on 9 held-out MiniWob++ tasks (click-checkboxes-large, click-color, click-tab-2, click-tab-2-hard, count-shape, drag-shapes, use-colorwheel-2, use-slider-2) after training on the remaining 50 tasks, the PIX2STRUCT-initialized model achieves an average score of 28.328.3, compared to 7.67.6 for a randomly initialized model.

    Furthermore, intermediate fine-tuning on MiniWob++ prior to fine-tuning on WebShop increases WebShop Task Score by +4.0+4.0 points compared to training directly from PIX2STRUCT.

  12. Knowl 12 — Limitations of Pixel-Only GUI Agents and Tree Search Scaling

    limitation

    The pixel-based GUI agent framework exhibits three key limitations:

    1. Environment Reset and Reward Requirements: MCTS policy improvement relies on deterministic environment transitions, exact environment state resets, and programmatic reward functions, which are often unavailable in live web services.
    2. Performance Gap with Text-Dominant LLMs: On text-dense web environments such as WebShop, pixel-only agents (46.746.7 score) lag behind large language models using HTML text inputs (67.567.5 score) due to LLM reasoning scale, pretraining data size, and structured text availability.
    3. Browser Automation Restrictions: Frameworks like Selenium do not natively support certain operating system interactions present in human demonstrations, including real-time sub-frame animations, mouseover hover states, native <select> dropdown handling, and drag-to-highlight copy-pasting.

Coverage note — None was omitted; all key contributions—including model architecture, action space, MCTS algorithm, empirical comparisons on MiniWob++ and WebShop, ablations, transfer experiments, and limitations—are represented.

References

  1. 1.James F. Allen, Nathanael Chambers, George Ferguson, Lucian Galescu, Hyuckchul Jung, Mary D. Swift, and William Taysom. Plow: A collaborative task learning agent. In AAAI Conference on Artificial Intelligence, 2007.
  2. 2.Thomas Anthony, Zheng Tian, and David Barber. Thinking fast and slow with deep learning and tree search. Advances in neural information processing systems, 30, 2017.
  3. 3.Dzmitry Bahdanau, Felix Hill, Jan Leike, Edward Hughes, Arian Hosseini, Pushmeet Kohli, and Edward Grefenstette. Learning to understand goal specifications by modelling reward. In International Conference on Learning Representations, 2018.
  4. 4.Bowen Baker, Ilge Akkaya, Peter Zhokov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune. Video pretraining (VPT): Learning to act by watching unlabeled online videos. Advances in Neural Information Processing Systems, 35:24639–24654, 2022.
  5. 5.S. R. K. Branavan, Luke Zettlemoyer, and Regina Barzilay. Reading between the lines: Learning to map high-level instructions to commands. In Annual Meeting of the Association for Computational Linguistics, 2010.
  6. 6.Andrea Burns, Deniz Arsan, Sanjna Agrawal, Ranjitha Kumar, Kate Saenko, and Bryan A Plummer. Interactive mobile app navigation with uncertain or under-specified natural language commands. arXiv preprint arXiv:2202.02312, 2022.
  7. 7.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022.
  8. 8.Remi Coulom. Efficient selectivity and backup operators in monte-carlo tree search. In Computers and Games, 2006.
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  10. 10.Yuqing Du, Ksenia Konyushkova, Misha Denil, Akhil Raju, Jessica Landon, Felix Hill, Nando de Freitas, and Serkan Cabi. Vision-language models as success detectors. arXiv preprint arXiv:2303.07280, 2023.
  11. 11.Hiroki Furuta, Ofir Nachum, Kuang-Huei Lee, Yutaka Matsuo, Shixiang Shane Gu, and Izzeddin Gur. Instruction-finetuned foundation models for multimodal web navigation. In Workshop on Reincarnating Reinforcement Learning at ICLR 2023, 2023a.
  12. 12.Hiroki Furuta, Ofir Nachum, Kuang-Huei Lee, Yutaka Matsuo, Shixiang Shane Gu, and Izzeddin Gurt. Instruction-finetuned foundation models for multimodal web navigation. In First Workshop on Multimodal Representation Learning at ICLR, 2023b.
  13. 13.Izzeddin Gur, Ulrich Rueckert, Aleksandra Faust, and Dilek Hakkani-Tur. Learning to navigate the web. arXiv preprint arXiv:1812.09195, 2018.
  14. 14.Izzeddin Gur, Ofir Nachum, Yingjie Miao, Mustafa Safdari, Austin Huang, Aakanksha Chowdhery, Sharan Narang, Noah Fiedel, and Aleksandra Faust. Understanding HTML with large language models. arXiv preprint 2210.03945, 2022.
  15. 15.Peter C Humphreys, David Raposo, Tobias Pohlen, Gregory Thornton, Rachita Chhaparia, Alistair Muldal, Josh Abramson, Petko Georgiev, Adam Santoro, and Timothy Lillicrap. A data-driven approach for learning to control computers. In International Conference on Machine Learning, pages 9466–9482. PMLR, 2022.
  16. 16.Sheng Jia, Jamie Ryan Kiros, and Jimmy Ba. Dom-q-net: Grounded rl on structured language. In International Conference on Learning Representations, 2019.
  17. 17.Geunwoo Kim, Pierre Baldi, and Stephen McAleer. Language models can solve computer tasks. arXiv preprint arXiv:2303.17491, 2023.
  18. 18.Kenton Lee, Mandar Joshi, Iulia Turc, Hexiang Hu, Fangyu Liu, Julian Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova. Pix2Struct: Screenshot parsing as pretraining for visual language understanding. arXiv preprint arXiv:2210.03347, 2022.
  19. 19.Gang Li and Yang Li. Spotlight: Mobile ui understanding using vision-language models with a focus. In The Eleventh International Conference on Learning Representations, 2022.
  20. 20.Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge. Mapping natural language instructions to mobile ui action sequences. arXiv preprint arXiv:2005.03776, 2020a.
  21. 21.Yang Li, Gang Li, Luheng He, Jingjie Zheng, Hong Li, and Zhiwei Guan. Widget captioning: Generating natural language description for mobile user interface elements. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5495–5510, Online, November 2020b. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.443. URL https://aclanthology.org/2020.emnlp-main.443.
  22. 22.Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang. Reinforcement learning on web interfaces using workflow-guided exploration. In International Conference on Learning Representations (ICLR), 2018. URL https://arxiv.org/abs/1802.08802.
  23. 23.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Kirkeby Fidjeland, Georg Ostrovski, Stig Petersen, Charlie Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. Human-level control through deep reinforcement learning. Nature, 518:529–533, 2015.
  24. 24.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332, 2021.
  25. 25.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67, 2020. URL https://arxiv.org/abs/1910.10683.
  26. 26.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. ArXiv, abs/1707.06347, 2017.
  27. 27.Noam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning, pages 4596–4604. PMLR, 2018.
  28. 28.Tianlin Shi, Andrej Karpathy, Linxi Fan, Jonathan Hernandez, and Percy Liang. World of bits: An open-domain platform for web-based agents. In Doina Precup and Yee Whye Teh, editors, Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 3135–3144. PMLR, 06–11 Aug 2017. URL https://proceedings.mlr.press/v70/shi17a.html.
  29. 29.David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al. Mastering the game of go without human knowledge. nature, 550(7676):354–359, 2017.
  30. 30.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  31. 31.Bryan Wang, Gang Li, Xin Zhou, Zhourong Chen, Tovi Grossman, and Yang Li. Screen2Words: Automatic mobile UI summarization with multimodal learning. In The 34th Annual ACM Symposium on User Interface Software and Technology, UIST ’21, page 498–510, New York, NY, USA, 2021. Association for Computing Machinery. ISBN 9781450386357. doi: 10.1145/3472749.3474765. URL https://doi.org/10.1145/3472749.3474765.
  32. 32.Ronald J. Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8:229–256, 2004.
  33. 33.Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144, 2016.
  34. 34.Shunyu Yao, Howard Chen, John Yang, and Karthik R Narasimhan. Webshop: Towards scalable real-world web interaction with grounded language agents. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=R9KnuFlvnU.
  35. 35.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR), 2023.
  36. 36.Luke S Zettlemoyer and Robert St. Amant. A visual medium for programmatic control of interactive applications. In Proceedings of the SIGCHI conference on Human Factors in Computing Systems, pages 199–206, 1999.

Citation

MLA
Shaw, P., et al. “From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 34354–70, https://proceedings.neurips.cc/paper_files/paper/2023/file/6c52a8a4fadc9129c6e1d1745f2dfd0f-Paper-Conference.pdf.
APA
Shaw, P., Joshi, M., Cohan, J., Berant, J., Pasupat, P., Hu, H., Khandelwal, U., Lee, K., & Toutanova, K. N. (2023). From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces. Advances in Neural Information Processing Systems, 36, 34354–34370. https://proceedings.neurips.cc/paper_files/paper/2023/file/6c52a8a4fadc9129c6e1d1745f2dfd0f-Paper-Conference.pdf
Chicago
Shaw, P., M. Joshi, J. Cohan, et al. 2023. “From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces”. Advances in Neural Information Processing Systems 36: 34354–70. https://proceedings.neurips.cc/paper_files/paper/2023/file/6c52a8a4fadc9129c6e1d1745f2dfd0f-Paper-Conference.pdf.
Harvard
Shaw, P. et al. (2023) “From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 34354–34370. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/6c52a8a4fadc9129c6e1d1745f2dfd0f-Paper-Conference.pdf.
Vancouver
1. Shaw P, Joshi M, Cohan J, Berant J, Pasupat P, Hu H, Khandelwal U, Lee K, Toutanova KN (2023) From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 34354–34370

BibTeX

@inproceedings{shaw2023from,
  title = {From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces},
  author = {Shaw, Peter and Joshi, Mandar and Cohan, James and Berant, Jonathan and Pasupat, Panupong and Hu, Hexiang and Khandelwal, Urvashi and Lee, Kenton and Toutanova, Kristina N},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {34354-34370},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/6c52a8a4fadc9129c6e1d1745f2dfd0f-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission