From Gaze to Guidance: Interpreting and Adapting to Users' Cognitive Needs with Multimodal Gaze-Aware AI Assistants

Valdemar DanryJavier HernandezAndrew D. WilsonPattie MaesJudith Amores

article2026arXiv2 citations

Demonstrates that equipping multimodal AI assistants with egocentric gaze tracking allows them to pinpoint user comprehension difficulties, significantly improving information recall while reducing interaction effort.

Listen

Modern artificial intelligence assistants rely heavily on explicit user prompts and remain blind to real-time behavioral cues that signal when a user is struggling. Because people often fail to recognize or articulate their own comprehension breakdowns, current conversational tools cannot easily provide targeted, proactive support. The article evaluates whether integrating first-person video with eye-tracking data into a multimodal large language model allows an assistant to infer cognitive difficulty and provide more effective retrospective guidance.

To test this approach, researchers developed a wearable prototype using smart glasses that stream egocentric video with projected gaze coordinates directly into a language model. The system decomposes the analysis into tracking timestamped visual behaviors and inferring specific trouble spots before delivering conversational voice support. The authors evaluated the system through a controlled within-subjects study involving 36 adult participants who read passages of varying complexity and received assistance from either the gaze-aware system or a text-only baseline assistant.

The study revealed three main findings. First, participants using the gaze-aware assistant achieved a statistically significant improvement in factual recall, scoring 96.3% compared to 88.9% with the baseline—a 7.4 percentage point gain. Second, users rated the gaze-informed diagnostic analysis as significantly more accurate and personalized, with nearly 64% preferring it over the text-only summary. Third, interactions were notably more efficient: users spoke approximately 31% fewer words with the gaze-aware assistant (57 words versus 83 words) and spent less effort steering the conversation or repairing misunderstandings. However, higher-order conceptual transfer and definition learning showed only slight, non-significant gains.

These findings indicate that grounding artificial intelligence in physical gaze traces successfully shifts the burden of diagnosing cognitive difficulties from the human to the machine. By pinpointing exact moments of hesitation or rereading, the assistant delivers relevant, low-friction help. Nevertheless, qualitative feedback highlighted that visual behavior remains inherently ambiguous: the model occasionally misidentified productive rereading, cross-referencing, or skimming as confusion, underscoring that surface gaze represents probabilistic evidence rather than ground truth.

Organizations developing multimodal assistants should treat gaze signals as hypotheses rather than definitive indicators. Systems should use hedged language, confirm user intent before offering explanations, and preserve user agency. Future research and development should focus on testing longer texts, integrating complementary physiological sensors such as pupil dilation, and building long-term user models. Practitioners must also establish robust privacy protections, as wearable video and eye tracking capture sensitive personal and environmental data.

The findings are supported by a rigorous experimental design and power analysis, providing strong confidence in the system's ability to improve local information retrieval and interaction efficiency. Readers should exercise caution regarding broader claims of deep conceptual learning or generalization beyond visually anchored reading tasks until real-time and multi-domain trials are conducted.

arXiv: 2604.08062
Cover for From Gaze to Guidance: Interpreting and Adapting to Users' Cognitive Needs with Multimodal Gaze-Aware AI Assistants

Abstract

Current LLM assistants are powerful at answering questions, but they have limited access to the behavioral context that reveals when and where a user is struggling. We present a gaze-grounded multimodal LLM assistant that uses egocentric video with gaze overlays to identify likely points of difficulty and target follow-up retrospective assistance. We instantiate this vision in a controlled study (n=36) comparing the gaze-aware AI assistant to a text-only LLM assistant. Compared to a conventional LLM assistant, the gaze-aware assistant was rated as significantly more accurate and personalized in its assessments of users' reading behavior and significantly improved people's ability to recall information. Users spoke significantly fewer words with the gaze-aware assistant, indicating more efficient interactions. Qualitative results underscored both perceived benefits in comprehension and challenges when interpretations of gaze behaviors were inaccurate. Our findings suggest that gaze-aware LLM assistants can reason about cognitive needs to improve cognitive outcomes of users.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Gaze as a signal for adaptive assistance
  • 2.2 Gaze-grounded multimodal assistants
  • 2.3 Positioning our work
  • 3 System Description
  • 3.1 Multimodal Scene Grounding
  • 3.2 Inference of Candidate Moments of Difficulty
  • 3.3 Conversational Assistance
  • 3.4 Design Rationale
  • 4 Study Design
  • 4.1 Participants and Data Collection
  • 4.2 Experimental Conditions
  • 4.3 Materials
  • 4.4 Procedure
  • 5 Analysis
  • 6 Quantitative Results
  • 6.1 The Gaze-aware Cognitive Assistant Increased Recall
  • 6.2 Eye Tracking Analyses Were Perceived as More Accurate and Personalized
  • 6.3 Perceived Effort and Frustration
  • 7 Exploratory Interaction Analysis
  • 7.1 Both assistants were highly aligned and consistently checked user needs
  • 7.2 Users appeared to steer less, and showed less confusion, with the gaze-aware AI assistant
  • 8 Qualitative Results
  • 8.0.1 Eye-Tracking Assistance Was Experienced as More Targeted, Adaptive, and Less Repetitive
  • 8.0.2 Eye-Tracking Analyses Felt Granular and Personally Relevant While Text-Only Analyses Felt Generic
  • 8.0.3 The System Reduced Cognitive Friction and Barriers to Seeking Help
  • 8.0.4 Misinterpretation of Reading Behaviors
  • 9 Discussion
  • 9.1 Limitations
  • 9.2 Design Implications for Gaze-aware AI Assistants
  • 10 Ethics, Consent, and Privacy
  • 11 Future Directions
  • 12 Conclusion
  • References
  • A Prompt Templates and Examples
  • A.1 Eye Tracking Analysis Prompt
  • A.2 Paragraph-Only Context Analysis Prompt
  • A.3 Real-Time Intervention Variant
  • A.4 Assistant Prompt
  • B LLM Analysis Samples
  • B.1 Example of Generic LLM Analysis
  • B.2 Example of Eye Tracking LLM Analysis
  • C Examples of Questions
  • D LLM-as-a-Judge - Classifier descriptions
  • E Additional Results

Knowls

  1. Knowl 1 — Gaze-Aware Multimodal Cognitive AI Assistant Architecture

    model/method

    The gaze-aware cognitive assistant system architecture captures first-person behavioral data and provides interactive conversational support through three stages:

    1. Multimodal Scene Grounding: Users wear Meta Aria glasses equipped with a forward-facing RGB camera (2880×28802880 \times 2880 resolution) and two inward-facing eye-tracking cameras. Eye-gaze vectors are estimated locally using the Project Aria gaze inference model and projected onto the egocentric frame as a semi-transparent marker. A 200×200200 \times 200 pixel crop centered on the gaze point is extracted in parallel to provide detailed local visual context alongside the global scene frame.

    2. Two-Stage Temporal Cognitive Need Inference: Every 0.5 s0.5\text{ s}, the full frame and gaze crop are sent to a multimodal large language model (GPT-4.1) to produce a structured description of the attended target:

    class GazeAnalysis(BaseModel):
        type: Literal["word", "object", "none"]
        content: str  # e.g., "Quantum"
        context: str  # e.g., "In the sentence: 'Quantum is...'"
    

    After a participant finishes reading a passage, the accumulated sequence of GazeAnalysis objects across the episode is processed retrospectively by GPT-4.1 to detect behavioral patterns (fixation durations, regressions, skips, and off-text glances) and map them to candidate points of cognitive struggle.

    1. Conversational Assistance: Inferred candidate need states and passage text are passed to a conversational agent implemented with the OpenAI Realtime API (gpt-4o-realtime-preview) streaming 24 kHz PCM16 mono audio. The assistant uses Socratic dialogue, explanations, and linguistic hedging to verify inferred difficulties with the user before explaining. Speech round-trip latency remains below 2 s2\text{ s}, with full pipeline latency from observation to assistant speech averaging 4.1 s4.1\text{ s}.
  2. Knowl 2 — Effect of Gaze-Aware Assistance on Reading Learning Outcomes

    empirical result

    In a within-subjects reading comprehension experiment (N=36N = 36), grounding conversational AI assistance in egocentric gaze behavior produced a selective benefit in factual recall while yielding non-significant positive differences in deeper comprehension metrics relative to an active text-only LLM baseline:

    • Literal Factual Recall: Gaze-aware assistance achieved significantly higher recall scores (96.3%±10.6%96.3\% \pm 10.6\%) than the control baseline (88.9%±15.9%88.9\% \pm 15.9\%), reflecting a mean paired increase of +7.4+7.4 percentage points (95% CI [1.3,13.5]95\%\text{ CI } [1.3, 13.5]; paired tt-test p=0.0187p = 0.0187; Wilcoxon signed-rank p=0.0209p = 0.0209; Cohen's dz=0.41d_z = 0.41).
    • Definition Probe Accuracy: Gaze-aware assistance resulted in 83.3%±37.8%83.3\% \pm 37.8\% accuracy compared to 77.8%±42.2%77.8\% \pm 42.2\% in control, a non-significant mean difference of +5.6+5.6 percentage points (95% CI [−10.5,21.6]95\%\text{ CI } [-10.5, 21.6]; paired tt-test p=0.4873p = 0.4873; Wilcoxon p=0.4795p = 0.4795; Cohen's dz=0.12d_z = 0.12).
    • Concept Inventory (Transfer) Score: Gaze-aware assistance resulted in 51.9%±27.0%51.9\% \pm 27.0\% accuracy versus 49.1%±30.3%49.1\% \pm 30.3\% in control, a non-significant mean difference of +2.8+2.8 percentage points (95% CI [−10.3,15.8]95\%\text{ CI } [-10.3, 15.8]; paired tt-test p=0.6679p = 0.6679; Wilcoxon p=0.7552p = 0.7552; Cohen's dz=0.07d_z = 0.07).
    • Self-Reported Understanding: Pre-to-post interaction changes in self-reported understanding did not differ between conditions, showing that objective recall gains were not reflected in participants' subjective comprehension ratings.
  3. Knowl 3 — Perceived Accuracy, Personalization, and User Preference for Gaze-Derived Analysis

    empirical result

    Evaluating the LLM-generated diagnostic summaries demonstrated that incorporating gaze behavior significantly enhanced user perception of diagnostic precision and personalization compared to text-only analysis (N=36N = 36):

    • Perceived Accuracy (1–7 Likert scale): Increased from 4.31±1.584.31 \pm 1.58 in the control condition to 5.50±1.065.50 \pm 1.06 in the gaze-aware condition (mean paired difference =1.19= 1.19, 95% CI [0.60,1.79]95\%\text{ CI } [0.60, 1.79]; paired tt-test p=0.00027p = 0.00027; Wilcoxon signed-rank p=0.00044p = 0.00044).
    • Perceived Personalization (1–7 Likert scale): Increased from 3.94±1.623.94 \pm 1.62 in control to 5.50±1.185.50 \pm 1.18 in the gaze-aware condition (mean paired difference =1.56= 1.56, 95% CI [0.99,2.12]95\%\text{ CI } [0.99, 2.12]; paired tt-test p<0.00001p < 0.00001; Wilcoxon signed-rank p=0.00005p = 0.00005).
    • Direct Summary Preference: In a side-by-side comparison on an additional passage, 63.9%63.9\% of participants (23/3623/36) preferred the eye-tracking-based summary, whereas 36.1%36.1\% (13/3613/36) preferred the text-only summary.
    • Difficulty Moderation Effect: Accuracy rating advantages for gaze-aware analysis diminished on more difficult technical passages (such as Superdeterminism), where baseline text-only prompts and gaze data converged on flagging the same complex terms.
  4. Knowl 4 — Conversational Efficiency and Interaction Dynamics Under Gaze-Aware Assistance

    empirical result

    Grounding conversational support in temporal gaze analyses reduced spoken user effort and conversational steering compared to text-only support (N=36N = 36):

    • Total User Spoken Words: Participants produced significantly fewer words during interactions with the gaze-aware assistant (57.28±66.5757.28 \pm 66.57 words) than with the text-only baseline (83.25±91.6283.25 \pm 91.62 words), representing a mean paired difference of −25.97-25.97 words (Wilcoxon signed-rank p=0.0469p = 0.0469, Cohen's dz=−0.29d_z = -0.29).
    • Conversational Turns: Turns showed a non-significant directional reduction from 8.39±3.598.39 \pm 3.59 in control to 7.22±3.847.22 \pm 3.84 in the gaze condition (Wilcoxon p=0.1067p = 0.1067).
    • NASA-TLX Workload Scores: Directional decreases favoring the gaze assistant were observed for mental demand (8.89±3.728.89 \pm 3.72 vs. 10.00±4.8110.00 \pm 4.81), effort (10.06±3.6710.06 \pm 3.67 vs. 11.00±4.5811.00 \pm 4.58), and frustration (4.78±3.964.78 \pm 3.96 vs. 5.17±4.255.17 \pm 4.25), though none reached statistical significance.
    • Dialogue Strategy Classification (LLM-as-a-Judge): Independent transcript evaluation using GPT-5.4 showed that users took the lead in steering the conversation significantly less often in the gaze condition (0.44±0.500.44 \pm 0.50) than in control (0.67±0.480.67 \pm 0.48; paired t=−2.26,p=0.030,dz=−0.38t = -2.26, p = 0.030, d_z = -0.38). Inferred user expressions of confusion were also lower in the gaze condition (0.39±0.490.39 \pm 0.49 vs. 0.61±0.490.61 \pm 0.49; paired t=−1.67,p=0.103,dz=−0.28t = -1.67, p = 0.103, d_z = -0.28), while assistant alignment with identified needs remained consistently high across both conditions (1.00±0.001.00 \pm 0.00 experimental vs. 0.94±0.230.94 \pm 0.23 control).
  5. Knowl 5 — Behavioral Misinterpretation and Semantic Ambiguity in Gaze Reasoning

    empirical result

    Qualitative analysis of semi-structured explicitation interviews identified two principal failure modes when multimodal LLMs interpret eye-gaze behavior:

    1. Misclassification of Normative and Strategic Reading: Eye movements such as rereading, backtracking, and skimming were often misclassified as confusion or comprehension breakdown. Users reported that rereading frequently served to consolidate concepts, connect ideas, or confirm well-understood statements, whereas skimming indicated existing familiarity with the topic. Explicitly prompting the LLM to search for struggles exacerbated false-positive interpretations of these productive behaviors.
    2. Invisibility of Latent Conceptual Disconnects: Gaze tracking captures visual attention on localized lexical tokens (words or phrases) but cannot directly measure abstract relationship synthesis. When participants struggled with high-level conceptual connections (such as the interaction between geological layers rather than the definitions of the terms themselves), the system focused on isolated vocabulary rather than the conceptual relationship.
  6. Knowl 6 — Controlled Experimental Evaluation Setup for Gaze-Grounded Reading Assistance

    experimental setup

    A within-subjects user study (N=36N = 36 analyzed; 5 excluded from N=41N = 41 due to device disconnections or completion under 10 minutes; age 18–44, 57% male, 43% female) compared a gaze-aware multimodal AI assistant to an active text-only baseline assistant across reading tasks:

    • Conditions: Participants completed two counterbalanced main blocks. In the experimental condition, the voice assistant was informed by both the passage text and an LLM analysis of the user's recorded egocentric gaze behavior. In the control condition, the assistant used the same prompt architecture operating solely on the text without gaze data.
    • Stimuli: Passages of controlled word count across varying Flesch-Kincaid Grade difficulty levels were selected from Wikipedia (including topics on Inflation, Plate Tectonics, Climate Change, Water Cycle, Superdeterminism, and Topological Quantum Computing). Phase 1 and 2 utilized fixed passages (Inflation and Plate Tectonics), while Phase 3 sampled passages across difficulties for side-by-side analysis comparison.
    • Instruments: Learning was measured via three-item multiple-choice factual retrieval probes, definition probes for undefined terms, and concept-inventory generalization items. User perception was assessed using 7-point Likert ratings of accuracy, confidence, and personalization, NASA-TLX workload scales, the Human Language Model Interaction Questionnaire (HLMIQ), behavioral word/turn logs, and post-study explicitation interviews.
  7. Knowl 7 — Design Principles for Gaze-Aware Multimodal Cognitive Assistants

    model/method

    Synthesizing empirical and interaction findings yields four design principles for gaze-aware multimodal assistants:

    1. Separate Need Detection from Assistance Delivery: Gaze signals primarily optimize targeting (identifying which concepts or passages caused hesitation) rather than explanation formulation itself. Systems should decouple trouble-spot localization from dialogic scaffolding, applying instructional strategies like Socratic questioning, self-explanation, or concept comparison during delivery.
    2. Model Gaze Cues as Probabilistic Hypotheses: Gaze patterns (fixations, saccades, regressions) are ambiguous and must not be treated as ground-truth indicators of comprehension failure. Systems should incorporate conversational hedging (e.g., framing observations as possibilities rather than conclusions) and allow users to confirm, reject, or redirect interpretations.
    3. Preserve User Agency and Minimize Disruption: Unsolicited real-time interventions risk disrupting immersion. Systems should utilize boundary-aligned intervention windows (e.g., post-task retrospective assistance) or lightweight opt-in prompts that let users control the engagement.
    4. Preserve Visual Context via Direct Multimodal Grounding: Ingesting full egocentric frames alongside gaze-centered crops preserves surrounding environmental and textual context, allowing the LLM to generalize beyond narrow OCR-based text pipelines to open-ended first-person tasks.
  8. Knowl 8 — Limitations of Retrospective Gaze-Grounded Cognitive Assistance

    limitation

    The study and system architecture are constrained by four key limitations:

    1. Task Generalization and Spatial Anchoring: Evaluation was restricted to reading comprehension, where attention is anchored to discrete visual tokens (text). Findings may not directly generalize to open-ended, dynamic, or socially complex environments lacking clearly localizable attentional targets.
    2. Stimulus Length and Ceiling Effects: Reading passages were short, yielding high baseline recall (88.9%88.9\% in control), which created potential ceiling effects that may have obscured larger between-condition differences on deeper transfer and definitional learning metrics.
    3. Retrospective Rather than Real-Time Support: Assistance was delivered retrospectively at reading boundaries because early pilot testing demonstrated that in-the-moment interruptions disrupted reader flow and caused frustration, leaving real-time triggering unmeasured.
    4. Conservative Active Baseline: The control condition utilized an LLM explicitly prompted to infer difficulties from text structure rather than a generic unprompted chatbot, isolating the specific incremental effect of gaze data but potentially underestimating the contrast relative to typical conversational agents.

Coverage note — No substantial contributed material was omitted from the extraction.

References

  1. 1.Yasmeen Abdrabou, Süleyman Özdel, Virmarie Maquiling, Efe Bozkir, and Enkelejda Kasneci. 2025. From gaze to data: Privacy and societal challenges of using eye-tracking data to inform GenAI models. In Proceedings of the 2025 Symposium on Eye Tracking Research and Applications. 1–9.
  2. 2.Dekel Abeles and Shlomit Yuval-Greenberg. 2017. Just look away: Gaze aversions as an overt attentional disengagement mechanism. Cognition 168 (2017), 99–109.
  3. 3.Rawan Alharbi, Tammy Stump, Nilofar Vafaie, Angela Pfammatter, Bonnie Spring, and Nabil Alshurafa. 2018. I can’t be myself: effects of wearable cameras on the capture of authentic behavior in the wild. Proceedings of the ACM on interactive, mobile, wearable and ubiquitous technologies 2, 3 (2018), 1–40.
  4. 4.Michael Argyle, Mark Cook, and Duncan Cramer. 1994. Gaze and mutual gaze. The British Journal of Psychiatry 165, 6 (1994), 848–850.
  5. 5.Riccardo Bovo, Steven Abreu, Karan Ahuja, Eric J. Gonzalez, Li-Te Cheng, and Mar Gonzalez-Franco. 2024. EmBARDiment: an Embodied AI Agent for Productivity in XR. arXiv:2408.08158 [cs.HC] https://arxiv.org/abs/2408.08158
  6. 6.Richard E Boyatzis. 1998. Transforming qualitative information: Thematic analysis and code development. Sage.
  7. 7.Jennifer Choe Bush, Peter Christopher Pantelis, Xavier Morin Duchesne, Sebastian Alexander Kagemann, and Daniel Patrick Kennedy. 2015. Viewing complex, dynamic scenes “through the eyes” of another person: The gaze-replay paradigm. PloS one 10, 8 (2015), e0134347.
  8. 8.Roser Cañigueral and Antonia F de C Hamilton. 2019. The role of eye gaze during natural social interactions in typical and autistic people. Frontiers in psychology 10 (2019), 560.
  9. 9.Michelene TH Chi, Nicholas De Leeuw, Mei-Hung Chiu, and Christian LaVancher. 1994. Eliciting self-explanations improves understanding. Cognitive science 18, 3 (1994), 439–477.
  10. 10.Michelene TH Chi and Ruth Wylie. 2014. The ICAP framework: Linking cognitive engagement to active learning outcomes. Educational psychologist 49, 4 (2014), 219–243.
  11. 11.John W Creswell and J David Creswell. 2017. Research design: Qualitative, quantitative, and mixed methods approaches. Sage publications.
  12. 12.Valdemar Danry, Pat Pataranutaporn, Yaoli Mao, and Pattie Maes. 2023. Don’t just tell me, ask me: Ai systems that intelligently frame explanations as questions improve human logical discernment accuracy over causal ai explanations. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–13.
  13. 13.Sidney K. D’Mello and Arthur C. Graesser. 2012. AutoTutor and Affective AutoTutor: Learning by Talking with Cognitively and Emotionally Intelligent Computers That Talk Back. ACM Transactions on Interactive Intelligent Systems 2, 4, Article 23 (2012). doi:10.1145/2395123.2395128
  14. 14.Sidney K. D’Mello, Andrew Olney, Charles Williams, and Peter Hays. 2012. Gaze Tutor: A Gaze-Reactive Intelligent Tutoring System. International Journal of Human-Computer Studies 70, 5 (2012), 377–398. doi:10.1016/j.ijhcs.2012.01.004
  15. 15.Gwyneth Doherty-Sneddon and Fiona G Phelps. 2005. Gaze aversion: A response to cognitive or social difficulty? Memory & cognition 33, 4 (2005), 727–733.
  16. 16.Sidney D’Mello, Blair Lehman, Reinhard Pekrun, and Art Graesser. 2014. Confusion can be beneficial for learning. Learning and Instruction 29 (2014), 153–170.
  17. 17.Selina N Emhardt, Margot van Wermeskerken, Katharina Scheiter, and Tamara van Gog. 2020. Inferring task performance and confidence from displays of eye movements. Applied Cognitive Psychology 34, 6 (2020), 1430–1443.
  18. 18.Alexandra Frischen, Andrew P Bayliss, and Steven P Tipper. 2007. Gaze cueing of attention: visual attention, social cognition, and individual differences. Psychological bulletin 133, 4 (2007), 694.
  19. 19.Holly Gorin, Jigna Patel, Qinyin Qiu, Alma Merians, Sergei Adamovich, and Gerard Fluet. 2024. A review of the use of gaze and pupil metrics to assess mental workload in gamified and simulated sensorimotor tasks. Sensors 24, 6 (2024), 1759.
  20. 20.Robert GM Hausmann and Kurt VanLehn. 2007. Explaining self-explaining: A contrast between content and generation. Frontiers in Artificial Intelligence and Applications 158 (2007), 417.
  21. 21.Javier Hernandez, Josh Lovejoy, Daniel McDuff, Jina Suh, Tim O’Brien, Arathi Sethumadhavan, Gretchen Greene, Rosalind W Picard, and Mary Czerwinski. 2021. Guidelines for Assessing and Minimizing Risks of Emotion Recognition Applications.. In ACII. 1–8.
  22. 22.Roy S Hessels, Antje Nuthmann, Marcus Nyström, Richard Andersson, Diederick C Niehorster, and Ignace TC Hooge. 2024. The fundamentals of eye tracking part 1: The link between theory and research question. Behavior Research Methods 57, 1 (2024), 16.
  23. 23.David Hestenes, Malcolm Wells, Gregg Swackhamer, et al. 1992. Force concept inventory. The physics teacher 30, 3 (1992), 141–158.
  24. 24.Roland S Johansson, Göran Westling, Anders Bäckström, and J Randall Flanagan. 2001. Eye–hand coordination in object manipulation. Journal of neuroscience 21, 17 (2001), 6917–6932.
  25. 25.Jeffrey D Karpicke and Henry L Roediger III. 2008. The critical importance of retrieval for learning. science 319, 5865 (2008), 966–968.
  26. 26.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2022. Decomposed prompting: A modular approach for solving complex tasks. arXiv preprint arXiv:2210.02406 (2022).
  27. 27.Peter König, Niklas Wilming, Tim C Kietzmann, Jose P Ossandón, Selim Onat, Benedikt V Ehinger, Ricardo R Gameiro, and Kai Kaspar. 2016. Eye movements as a window to cognitive processes. Journal of eye movement research 9, 5 (2016), 25.
  28. 28.Sébastien Lallé, Cristina Conati, and Giuseppe Carenini. 2016. Predicting Confusion in Information Visualization from Eye Tracking and Interaction Data.. In IJCAI. 2529–2535.
  29. 29.Moritz Langner, Peyman Toreini, and Alexander Maedche. 2023. Leveraging eye tracking technology for a situation-aware writing assistant. In Proceedings of the 2023 Symposium on Eye Tracking Research and Applications. 1–2.
  30. 30.Jaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu, Sebastian S Rodriguez, and Jon E Froehlich. 2024. GazePointAR: A context-aware multimodal voice assistant for pronoun disambiguation in wearable augmented reality. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 1–20.
  31. 31.Jia Zheng Lim, James Mountstephens, and Jason Teo. 2020. Emotion recognition using eye-tracking: taxonomy, review and current challenges. Sensors 20, 8 (2020), 2384.
  32. 32.Gwen Marchand and Ellen A Skinner. 2007. Motivational dynamics of children’s academic help-seeking and concealment. Journal of Educational Psychology 99, 1 (2007), 65.
  33. 33.Diane C Mézière, Niilo E Hautala, Timo T Heikkilä, and Johanna K Kaakinen. 2025. Eye-movement markers of mind wandering during reading: A meta-analysis. Memory & Cognition (2025), 1–26.
  34. 34.Chiara Mirandola, Alfonso Ciriello, Martina Gigli, and Cesare Cornoldi. 2018. Metacognitive monitoring of text comprehension: An investigation on postdictive judgments in typically developing children and children with reading comprehension difficulties. Frontiers in psychology 9 (2018), 2253.
  35. 35.William E Nagy, Patricia A Herman, and Richard C Anderson. 1985. Learning words from context. Reading research quarterly (1985), 233–253.
  36. 36.Paul Nation and David Beglar. 2007. A vocabulary size test. (2007).
  37. 37.Mariya Pachman, Amaël Arguel, Lori Lockyer, Gregor Kennedy, and Jason Lodge. 2016. Eye tracking and early detection of confusion in digital learning environments: Proof of concept. Australasian Journal of Educational Technology 32, 6 (2016).
  38. 38.Taiying Peng, Jiacheng Hua, Miao Liu, and Feng Lu. 2025. In the eye of mllm: Benchmarking egocentric video intent understanding with gaze-guided prompting. arXiv preprint arXiv:2509.07447 (2025).
  39. 39.Alexander Plopski, Teresa Hirzle, Nahal Norouzi, Long Qian, Gerd Bruder, and Tobias Langlotz. 2022. The eye in extended reality: A survey on gaze interaction and eye tracking in head-worn extended reality. ACM Computing Surveys (CSUR) 55, 3 (2022), 1–39.
  40. 40.Jun Rekimoto. 2025. GazeLLM: Multimodal LLMs incorporating human visual attention. In Proceedings of the Augmented Humans International Conference 2025. 302–311.
  41. 41.Jayasankar Santhosh, Andreas Dengel, and Shoya Ishimaru. 2024. Gaze-driven adaptive learning system with ChatGPT-generated summaries. IEEE Access 12 (2024), 173714–173733.
  42. 42.Gabriel Herbert Sarch, Balasaravanan Thoravi Kumaravel, Sahithya Ravi, Vibhav Vineet, and Andrew D Wilson. 2025. Grounding Task Assistance with Multimodal Cues from a Single Demonstration. In Findings of the Association for Computational Linguistics: ACL 2025. 12807–12833.
  43. 43.Katharina Scheiter, Carina Schubert, Anne Schüler, Holger Schmidt, Gottfried Zimmermann, Benjamin Wassermann, Marie-Christin Krebs, and Thérése Eder. 2019. Adaptive multimedia: Using gaze-contingent instructional guidance to provide personalized processing support. Computers & Education 139 (2019), 31–47.
  44. 44.Thimo Schulz, Chiara Krisam, and Julia Seitz. 2025. EyeGPT: A Cognitive Load-Adaptive GenAI Assistant with Eye Tracking for Programming Education. In NeuroIS Retreat. Springer, 45–55.
  45. 45.John L Sibert, Mehmet Gokturk, and Robert A Lavine. 2000. The reading assistant: eye gaze triggered auditory prompting for reading remediation. In Proceedings of the 13th annual ACM symposium on User interface software and technology. 101–107.
  46. 46.Julie Dangremond Stanton, Amanda J Sebesta, and John Dunlosky. 2021. Fostering metacognition to support student learning and performance. CBE—Life Sciences Education 20, 2 (2021), fe3.
  47. 47.Enkeleda Thaqi, Mohamed Omar Mantawy, and Enkelejda Kasneci. 2024. SARA: Smart AI reading assistant for reading comprehension. In Proceedings of the 2024 Symposium on Eye Tracking Research and Applications. 1–3.
  48. 48.Margot van Wermeskerken, Damien Litchfield, and Tamara van Gog. 2018. What am I looking at? Interpreting dynamic and static gaze displays. Cognitive science 42, 1 (2018), 220–252.
  49. 49.Pierre Vermersch. 1994. The explicitation interview. French original ESF (1994).
  50. 50.Ru Wang, Zach Potter, Yun Ho, Daniel Killough, Linxiu Zeng, Sanbrita Mondal, and Yuhang Zhao. 2024. GazePrompt: Enhancing Low Vision People’s Reading Experience with Gaze-Aware Augmentations. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 1–17.
  51. 51.Xin Wang, Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bugra Tekin, Felipe Vieira Frujeri, et al. 2023. Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 20270–20281.
  52. 52.Haowei Zhang, Jianzhe Liu, Zhen Han, Shuo Chen, Bailan He, Volker Tresp, Zhiqiang Xu, and Jindong Gu. 2024. Visual question decomposition on multimodal large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024. 1926–1949.
  53. 53.Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara, and Filip Ilievski. 2023. Visual cropping improves zero-shot question answering of multimodal large language models. In R0-FoMo: Robustness of Few-shot and Zero-shot Learning in Large Foundation Models.
  54. 54.Yichi Zhang, Xin Luna Dong, Zhaojiang Lin, Andrea Madotto, Anuj Kumar, Babak Damavandi, Joyce Chai, and Seungwhan Moon. 2025. Proactive assistant dialogue generation from streaming egocentric videos. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 12055–12079.
  55. 55.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems 36 (2023), 46595–46623.

Citation

MLA
Danry, V., et al. “From Gaze to Guidance: Interpreting and Adapting to Users' Cognitive Needs with Multimodal Gaze-Aware AI Assistants”. arXiv, 2026, http://arxiv.org/abs/2604.08062v1.
APA
Danry, V., Hernandez, J., Wilson, A., Maes, P., & Amores, J. (2026). From Gaze to Guidance: Interpreting and Adapting to Users' Cognitive Needs with Multimodal Gaze-Aware AI Assistants. arXiv. http://arxiv.org/abs/2604.08062v1
Chicago
Danry, V., J. Hernandez, A. Wilson, P. Maes, and J. Amores. 2026. “From Gaze to Guidance: Interpreting and Adapting to Users' Cognitive Needs with Multimodal Gaze-Aware AI Assistants”. arXiv. http://arxiv.org/abs/2604.08062v1.
Harvard
Danry, V. et al. (2026) “From Gaze to Guidance: Interpreting and Adapting to Users' Cognitive Needs with Multimodal Gaze-Aware AI Assistants”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2604.08062v1.
Vancouver
1. Danry V, Hernandez J, Wilson A, Maes P, Amores J (2026) From Gaze to Guidance: Interpreting and Adapting to Users' Cognitive Needs with Multimodal Gaze-Aware AI Assistants. arXiv

BibTeX

@article{danry2026from,
  title = {From Gaze to Guidance: Interpreting and Adapting to Users' Cognitive Needs with Multimodal Gaze-Aware AI Assistants},
  author = {Danry, Valdemar and Hernandez, Javier and Wilson, Andrew and Maes, Pattie and Amores, Judith},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2604.08062v1},
  eprint = {2604.08062}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/