Thinking Ahead: Prospection-Guided Retrieval of Memory with Language Models

Harshita ChopraKrishna ChintalapudiSuman NathRyen W. WhiteChirag Shah

article2026arXiv2 citations

Proposes Prospection-Guided Retrieval, a framework that expands user queries into simulated future actions to retrieve semantically distant interaction memories, nearly tripling recall compared to standard retrieval-augmented generation on long-horizon dialogue tasks.

Listen

Conversational artificial intelligence assistants often struggle with long-term personalization when required to retrieve relevant user history across extended interactions. Current retrieval systems rely primarily on retrospective search, querying memory databases using direct semantic or keyword similarity to a user's prompt. In practice, critical user context—such as past constraints, personal preferences, or schedule conflicts—often shares little direct textual similarity with a new request, leading conventional systems to overlook essential background information.

The article evaluates whether simulating future trajectories can improve memory recall and response personalization for AI assistants. It introduces Prospection-Guided Retrieval, an approach inspired by human cognitive prospection that decouples the memory retrieval process from how data is physically stored. Rather than executing a single retrospective query, the system simulates plausible future actions or subgoals, uses those imagined steps as search probes to surface relevant memories, and iteratively refines its simulations based on newly retrieved user context.

The evaluation compared this prospection approach against state-of-the-art memory baselines across 1,625 queries and 185 user profiles across three benchmarks, including a new multi-session dataset named MemoryQuest. The MemoryQuest benchmark specifically challenges retrieval systems by requiring assistants to identify three to five dated, semantically distant memories across long interaction histories. Assessments combined automated evaluations with large language models alongside blinded human evaluations to judge context recall and answer quality.

The findings show that Prospection-Guided Retrieval substantially outperforms standard similarity and graph-based retrieval methods. On the MemoryQuest benchmark, the prospection framework achieved a memory recall score of approximately 0.72, nearly tripling the recall of the strongest baseline at roughly 0.26 and outperforming graph-based retrieval at 0.16. For recovering the full set of required memories for a query, the prospection method achieved an exact recall score of approximately 0.33 compared to under 0.03 for all baselines. Across response quality evaluations, judges preferred answers generated with prospection over baseline outputs in roughly 89% to 98% of automated comparisons, a trend supported by human judges who favored prospection in 60% to 95% of cases.

These results indicate that adopting an active, simulation-driven retrieval policy resolves the structural bottleneck of passive memory lookup in personal assistants. By anticipating informational needs before generating answers, conversational agents can deliver safer, more proactive, and contextually grounded guidance. The primary operational trade-off is computational cost: executing iterative prospection requires an average of four to five model calls per query compared to one or two for standard methods, which may increase latency and computing expenses in high-throughput environments.

Organizations developing personalized conversational agents should consider piloting iterative prospection layers on top of their existing memory databases rather than re-architecting their underlying storage systems. Future implementation efforts should focus on training lightweight, specialized retrieval models directly on generated prospection traces to reduce computational overhead for real-time applications. While the results demonstrate clear performance gains across multiple benchmarks, organizations should note that performance depends on the quality of the underlying reasoning model, and further piloting is advised before deploying prospection in mission-critical or low-latency production workflows.

Cover for Thinking Ahead: Prospection-Guided Retrieval of Memory with Language Models

Abstract

Long-horizon personalization requires dialogue assistants to retrieve user-specific facts from extended interaction histories. In practice, many relevant facts often have low semanticsimilarity to the query under dense retrieval. Standard Retrieval-Augmented Generation (RAG) and GraphRAG systems are still largely retrospective: they rely on embedding similarity to the query or on fixed graph traversals, so they often miss facts that matter for the user's needs but lie far from the query in embedding space. Inspired by prospection, the human ability to use imagined futures as cues for recall, we introduce Prospection-Guided Retrieval (PGR), which decouples retrieval from how memories are stored. Given a user query, PGR first expands the goal into a short Tree-of-Thought (ToT) or linear chain of plausible next steps, and uses these steps as retrieval probes rather than relying on the original query alone. The facts retrieved by these probes are then used to personalize the next round of prospection, enabling PGR to uncover additional memories that become relevant only after the simulation is grounded in the user's history. We also introduce MemoryQuest, a challenging multi-session benchmark in which each query is annotated with 3--5 dated reference facts subject to a low query-reference similarity constraint. Across 1,625 queries spanning 185 user profiles from 3 publicly available datasets, PGR-TOT substantially improves retrieval, including nearly 3x recall on MemoryQuest over the strongest baseline. In pairwise LLM-as-judge comparisons against baselines, PGR-generated responses are preferred on 89--98% of queries, with blinded human annotations on held-out subsets showing the same trend. Overall, the results demonstrate that explicit prospection yields large gains in long-horizon retrieval and response quality relative to similarity-only baselines.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Prospection-Guided Retrieval
  • 3.1 System Architecture Overview
  • 3.2 Long Term Memory
  • 3.3 Prospection Guided Retrieval
  • 4 Benchmark Dataset: MemoryQuest
  • 4.1 Gap Analysis and Comparison
  • 4.2 Dataset Construction Pipeline
  • 5 Experiments
  • 5.1 Baselines
  • 5.2 Evaluation Metrics
  • 6 Results and Discussion
  • 7 Conclusion
  • References
  • A Appendix
  • A.1 PGR with Alternative Generation Models
  • A.2 Computational Cost.
  • A.3 End-to-End PGR Trace: MemoryQuest Example

Knowls

  1. Knowl 1 — Prospection-Guided Retrieval (PGR) Framework

    model/method

    Prospection-Guided Retrieval (PGR) is a retrieval framework for personalized conversational agents that decouples the retrieval policy from the underlying memory substrate. Unlike conventional Retrieval-Augmented Generation (RAG) and graph traversal methods that match queries retrospectively against stored documents or static graph edges, PGR operationalizes mental simulation through an iterative loop:

    Simulate Future⟶Retrieve Relevant Memories⟶Refine Trajectory\text{Simulate Future} \longrightarrow \text{Retrieve Relevant Memories} \longrightarrow \text{Refine Trajectory}

    Given a user query qq, the framework operates across two main prospection phases:

    1. Query Augmentation: Underspecified user queries are disambiguated with initial context to produce an augmented query qAq_A: qA=LLM(PA(q,ϕ(q;M,Θo)))q_A = \text{LLM}(P_A(q, \phi(q; \mathcal{M}, \Theta_o))) where M\mathcal{M} is the memory store, ϕ(x;M,Θ)\phi(x; \mathcal{M}, \Theta) is an abstract retrieval function returning facts relevant to a sub-query xx under retrieval parameters Θ\Theta, Θo\Theta_o denotes initial augmentation parameters, and PAP_A is an augmentation prompt.

    2. Phase 1 (Initial Prospection): A language model simulates a broad horizon of potential future subgoals or contextual actions, structured as a linear Chain-of-Thought (CoT) or branching Tree-of-Thought (ToT), denoted S0={s10,…,sL0}=LLM(Psim(qA))S^0 = \{s^0_1, \dots, s^0_L\} = \text{LLM}(P_{\text{sim}}(q_A)). Each step sl0s^0_l acts as an independent sub-query probe against M\mathcal{M}. The accumulated fact set RR is initialized as: R←ϕ(q;M,Θ)∪⋃l=1Lϕ(sl0;M,Θ)R \leftarrow \phi(q; \mathcal{M}, \Theta) \cup \bigcup_{l=1}^L \phi(s^0_l; \mathcal{M}, \Theta)

    3. Phase 2 (Personalized Prospection): Using the accumulated memories RR, the prospection sequence is iteratively refined to incorporate user-specific constraints: Si=LLM(Pref(qA,R,Si−1))S^i = \text{LLM}(P_{\text{ref}}(q_A, R, S^{i-1})) For each refined sub-query sli∈Sis^i_l \in S^i, new facts are retrieved: Δi=⋃lϕ(sli;M,Θ)∖R\Delta^i = \bigcup_l \phi(s^i_l; \mathcal{M}, \Theta) \setminus R. The fact set is updated via R←R∪ΔiR \leftarrow R \cup \Delta^i. The process repeats until ∣Δi∣|\Delta^i| drops below a convergence threshold δ\delta or an iteration limit ImaxI_{\text{max}} is reached, returning the final fact set R∗R^* and a prospection reasoning summary.

  2. Knowl 2 — Prospection-Guided Retrieval Algorithm

    algorithm

    The complete execution flow of Prospection-Guided Retrieval (PGR) is defined as follows:

    Input: User query qq, memory store M\mathcal{M}, abstract retrieval function ϕ\phi, augmentation parameters Θo\Theta_o, retrieval parameters Θ\Theta, convergence threshold δ\delta, max iterations ImaxI_{\text{max}}
    Output: Final retrieved fact set R∗R^*
    R←∅R \leftarrow \emptyset
    qA←LLM(PA(q,ϕ(q;M,Θo)))q_A \leftarrow \text{LLM}(P_A(q, \phi(q; \mathcal{M}, \Theta_o)))
    S0←LLM(Psim(qA))S^0 \leftarrow \text{LLM}(P_{\text{sim}}(q_A))
    R←ϕ(q;M,Θ)∪⋃l=1∣S0∣ϕ(sl0;M,Θ)R \leftarrow \phi(q; \mathcal{M}, \Theta) \cup \bigcup_{l=1}^{|S^0|} \phi(s^0_l; \mathcal{M}, \Theta)
    for i=1,2,…,Imaxi = 1, 2, \dots, I_{\text{max}} do
        Si←LLM(Pref(qA,R,Si−1))S^i \leftarrow \text{LLM}(P_{\text{ref}}(q_A, R, S^{i-1}))
        Δi←⋃l=1∣Si∣ϕ(sli;M,Θ)∖R\Delta^i \leftarrow \bigcup_{l=1}^{|S^i|} \phi(s^i_l; \mathcal{M}, \Theta) \setminus R
        R←R∪ΔiR \leftarrow R \cup \Delta^i
        if ∣Δi∣<δ|\Delta^i| < \delta then
            break
        end if
    end for
    return R∗←RR^* \leftarrow R

    In standard implementations, Θ={K,τ}\Theta = \{K, \tau\} where KK is the number of facts retrieved per sub-query and τ\tau is the cosine similarity threshold (e.g., K=5K=5, τ=0.3\tau=0.3). For the initial query augmentation step, a stricter threshold τo=0.6\tau_o = 0.6 is used with Θo={Ko,τo}\Theta_o = \{K_o, \tau_o\}. The convergence threshold is set to δ=5\delta = 5 newly retrieved facts, and Imax=10I_{\text{max}} = 10 iterations.

  3. Knowl 3 — MemoryQuest Benchmark for Long-Horizon Personalization

    definition

    MemoryQuest is a multi-session conversational benchmark designed to evaluate long-term memory retrieval under severe semantic divergence between queries and required evidence. Constructed on top of PersonaLens demographic profiles across 20 domains, it contains 50 user profiles and 535 distinct queries spanning an average of 77.6 dialogue sessions (6.2 exchanges per session) per user.

    The benchmark construction follows a three-step synthetic pipeline:

    1. Query Generation: Candidate queries requiring multi-step planning are generated with an explicit query date and 3–5 ground-truth reference facts. Candidates are filtered to enforce low query–reference cosine similarity (average similarity γ≤0.3\gamma \le 0.3), specifically testing the agent's ability to bypass the semantic similarity bottleneck.
    2. Timeline Synthesis: An LLM constructs a chronological timeline of 5–10 events starting before the earliest required-reference date. Each required reference is placed in exactly one event, with remaining events populated by persona-consistent filler conversations.
    3. Dialogue Expansion: Timeline events are expanded into 5–30 turn conversations conditioned on persona demographics. These dialogues are then converted into atomic facts stored in the user's long-term memory.
  4. Knowl 4 — Comparative Retrieval and Response Win Rates of PGR Across Benchmarks

    data/table

    PGR-TOT (PGR with Tree-of-Thought prospection) was evaluated against three state-of-the-art memory baselines—GraphRAG, TaciTree, and Mem0g—across three benchmarks: MemoryQuest (535 queries, multi-reference), ImplexConv (196 queries, opposed setting with semantically distant context), and PersonaMem (894 queries, 32k history length). Evaluated using GPT-4o for generation, text-embedding-3-small for retrieval, and GPT-5.2 alongside human annotators (Amazon MTurk) for pairwise response judging:

    Dataset MemoryQuest ImplexConv PersonaMem
    Method Recall RecallExact\text{Recall}_{\text{Exact}} Win RatePGR_{\text{PGR}} (%) Recall Win RatePGR_{\text{PGR}} (%) Recall Win RatePGR_{\text{PGR}} (%)
    LLM Humans* LLM Humans* LLM Humans*
    GraphRAG 0.160 0.003 96.6 90.0 0.072 96.2 70.0 0.011 98.4 80.0
    TaciTree 0.235 0.024 94.3 75.0 0.122 89.9 60.0 0.073 88.7 65.0
    Mem0g 0.256 0.015 94.7 95.0 0.050 93.5 70.0 0.121 92.9 65.0
    PGR-TOT 0.723 0.326 – – 0.286 – – 0.229 – –

    Recall is the average fraction of ground-truth reference facts present in the retrieved context. RecallExact\text{Recall}_{\text{Exact}} is the percentage of queries where all required reference facts (3–5 facts) were retrieved. Win RatePGR_{\text{PGR}} denotes the percentage of pairwise comparisons where PGR-TOT responses were preferred over the baseline. PGR-TOT nearly triples the recall of the best baseline on MemoryQuest (0.723 vs. 0.256) and achieves a 20-fold increase in RecallExact\text{Recall}_{\text{Exact}} (0.326 vs. 0.015), while securing 88.7%–98.4% win rates under LLM judges and 60.0%–95.0% under human judges.

  5. Knowl 5 — Structured Long-Term Memory Representation, Upsert, and Periodic Consolidation

    model/method

    PGR operates over an underlying long-term memory store structured around atomic statements called memory facts. Each fact is represented as a structured dictionary containing:

    • fid: A unique sequential fact identifier.
    • info: Concise, atomic text extracted by an LLM from conversation logs.
    • type: One of six taxonomy categories: Identity (demographics, profession), Preference (likes, habits), Goal (intentions, plans), Interest (hobbies, topics), Activity (recurring/past actions), or Event (milestones, situations).
    • frequency: Occurrence count across conversations.
    • related_entities: List of extracted topic/entity keywords.
    • provenance: IDs of source conversation sessions.

    Memory management uses two operational policies:

    1. Upsert: When new dialogue occurs, the LLM determines whether extracted facts update existing items (incrementing frequency and editing info) or are added as new entries.
    2. Consolidation: When 50 new facts accumulate since the last merge, facts are embedded and clustered by cosine similarity. Within each cluster, facts whose most recent timestamps fall within a 7-day time window are merged by an LLM into a single comprehensive fact, aggregating their provenance and identifiers.
  6. Knowl 6 — Ablations on Iterative Refinement, Prospection Strategy, and Summary Conditioning

    empirical result

    Ablation studies on the MemoryQuest benchmark demonstrate the individual contributions of iterative prospection, branching structures, and prospection-augmented generation (PAG):

    1. Iterative Prospection vs. Single-Phase Base:
    • PGR-TOT Recall improves from 0.677 (Base) to 0.723 (Iterative), yielding a 70.13% win rate in generation over the single-phase base.
    • PGR-COT Recall improves from 0.645 (Base) to 0.718 (Iterative), yielding a 75.41% win rate over the base.
    1. Prospection Strategy (ToT vs. CoT):
    • Branching Tree-of-Thought (PGR-TOT) achieves a Recall of 0.723 compared to 0.718 for linear Chain-of-Thought (PGR-COT), winning 55.67% of direct head-to-head pairwise comparisons.
    1. Prospection-Augmented Generation (PAG):
    • Conditioning final answer generation on both retrieved facts and the synthesized prospection summary improves PGR-TOT recall from 0.703 to 0.722 (75.73% win rate against unconditioned generation).
    • For PGR-COT, summary conditioning increases recall from 0.722 to 0.735 (74.37% win rate).
  7. Knowl 7 — Human Annotator and LLM-as-a-Judge Decision Agreement Matrix

    data/table

    Pairwise response quality between PGR-TOT and baseline models was evaluated using human annotators on Amazon MTurk and Azure OpenAI GPT-5.2 on held-out subsets of MemoryQuest (4%), ImplexConv (5%), and PersonaMem (2%). The aggregated agreement matrix across all dataset-baseline pairs is:

    Human \ LM PGR Tie Baseline
    PGR 70.0% 4.0% 2.0%
    Tie 8.0% 0.0% 0.7%
    Baseline 12.0% 2.0% 1.3%

    Humans and the LLM judge concurrently agree on preferring PGR in 70.0% of all evaluated cases. In total, human annotators selected PGR in 76.0% (70.0%+4.0%+2.0%70.0\% + 4.0\% + 2.0\%) of evaluations, validating that LLM judge preference closely aligns with human assessments of anticipatory and proactive reasoning.

  8. Knowl 8 — Cross-LLM Generalization with DeepSeek-V3.2

    empirical result

    To establish that PGR's retrieval gains are model-agnostic, experiments replacing the GPT-4o generation backbone with DeepSeek-V3.2 while maintaining identical embeddings (text-embedding-3-small) and retrieval parameters yielded the following recall results:

    Dataset Base Recall Iterative Recall Base RecallExact\text{Recall}_{\text{Exact}} Iterative RecallExact\text{Recall}_{\text{Exact}}
    MemoryQuest 0.671 0.748 0.245 0.348
    ImplexConv 0.308 0.374 – –
    PersonaMem 0.210 0.229 – –

    Using DeepSeek-V3.2, iterative prospection matches or exceeds the performance observed with GPT-4o on MemoryQuest (0.748 Recall / 0.348 RecallExact\text{Recall}_{\text{Exact}} vs. GPT-4o's 0.723 / 0.326), confirming that the prospection-guided retrieval mechanism generalizes across distinct LLM architectures.

  9. Knowl 9 — Query-Time LLM Call Overhead and Fact Retrieval Volume

    data/table

    The computational overhead of PGR-TOT at inference time was evaluated in terms of the average number of query-time LLM calls and unique facts retrieved across datasets:

    Dataset GraphRAG Mem0g TaciTree PGR-TOT PGR Ph-1 Facts PGR Ph-2 New Facts
    MemoryQuest 1.0 2.0 ∼\sim13.2 4.79 26.78 8.41
    ImplexConv 1.0 2.0 ∼\sim15.9 4.32 22.41 3.85
    PersonaMem 1.0 2.0 ∼\sim9.7 4.41 13.04 4.83

    PGR-TOT executes on average 4.32–4.79 LLM calls per query (composed of 1 Phase-1 initial prospection call, average Phase-2 refinement calls, 1 prospection summary call, and 1 final answer generation call). While higher than single-pass retrieval methods (GraphRAG: 1.0, Mem0g: ∼\sim2.0), PGR-TOT uses fewer than half the inference LLM calls required by tree-pruning methods like TaciTree (∼\sim9.7–15.9).

  10. Knowl 10 — Computational Overhead and Real-Time Latency Trade-off in PGR

    limitation

    PGR achieves higher recall and proactive personalization by planning and exploring multiple hypothesized future paths, but this multi-phase simulation introduces additional query-time LLM invocations (averaging 4–5 calls per query compared to 1–2 for standard dense RAG or GraphRAG). While the embedding-based vector lookups for each sub-step can be executed in parallel with negligible latency, the sequential LLM generation calls for prospection and iterative refinement increase inference latency, creating computational overhead that may constrain real-time deployment without specialized lightweight retriever distillation.

Coverage note — None was omitted; all contributed methods (PGR, memory taxonomy/consolidation), benchmarks (MemoryQuest), algorithms, empirical tables (Table 1-7), ablations, and stated limitations were fully converted into self-contained knowls.

References

  1. 1.Daniel L. Schacter, Roland G. Benoit, and Karl K. Szpunar. Episodic future thinking: Mechanisms and functions. Current Opinion in Behavioral Sciences, 17:41–50, 2017.
  2. 2.Donna Rose Addis, Alana T. Wong, and Daniel L. Schacter. Remembering the past and imagining the future: Common and distinct neural substrates during event construction and elaboration. Neuropsychologia, 45(7):1363–1377, 2007.
  3. 3.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. Retrieval augmented language model pre-training. In International conference on machine learning, pages 3929–3938. PMLR, 2020.
  4. 4.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K%C3%BCttler, Mike Lewis, Wen-tau Yih, Tim Rockt%C3%A4schel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33:9459–9474, 2020.
  5. 5.Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. From local to global: A graph rag approach to query-focused summarization. arXiv preprint arXiv:2404.16130, 2024.
  6. 6.Jing Xu, Arthur Szlam, and Jason Weston. Beyond goldfish memory: Long-term open-domain conversation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5180–5197, Dublin, Ireland, May 2022. Association for Computational Linguistics.
  7. 7.Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, and Dong Yu. Longmemeval: Benchmarking chat assistants on long-term interactive memory, 2024.
  8. 8.Gautier Izacard and Edouard Grave. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics (EACL), pages 874–880. Association for Computational Linguistics, 2021. (FiD: Fusion-in-Decoder approach).
  9. 9.Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology, pages 1–22, 2023.
  10. 10.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. Advances in neural information processing systems, 36:11809–11822, 2023.
  11. 11.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR), 2023.
  12. 12.Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. Mem0: Building production-ready ai agents with scalable long-term memory. arXiv preprint arXiv:2504.19413, 2025.
  13. 13.Demis Hassabis, Dharshan Kumaran, Seralynne D. Vann, and Eleanor A. Maguire. Patients with hippocampal amnesia cannot imagine new experiences. Proceedings of the National Academy of Sciences, 104(5):1726–1731, 2007.
  14. 14.Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, and Yuwei Fang. Evaluating very long-term conversational memory of llm agents, 2024.
  15. 15.Bowen Jiang, Yuan Yuan, Maohao Shen, Zhuoqun Hao, Zhangchen Xu, Zichen Chen, Ziyi Liu, Anvesh Rao Vijjini, Jiashu He, Hanchao Yu, et al. Personamem-v2: Towards personalized intelligence via learning implicit user personas and agentic memory. arXiv preprint arXiv:2512.06688, 2025.
  16. 16.Xintong Li, Jalend Bantupalli, Ria Dharmani, Yuwei Zhang, and Jingbo Shang. Toward multi-session personalized conversation: A large-scale dataset and hierarchical tree framework for implicit reasoning. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 11493–11506. Association for Computational Linguistics, 2025.
  17. 17.Zheng Zhao, Clara Vania, Subhradeep Kayal, Naila Khan, Shay B Cohen, and Emine Yilmaz. PersonaLens: A benchmark for personalization evaluation in conversational AI assistants. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Findings of the Association for Computational Linguistics: ACL 2025, pages 18023–18055, Vienna, Austria, July 2025. Association for Computational Linguistics.
  18. 18.OpenAI. New embedding models and api updates. https://openai.com/index/new-embedding-models-and-api-updates/, 2024. Refers to text-embedding-3-small.

Citation

MLA
Chopra, H., et al. “Thinking Ahead: Prospection-Guided Retrieval of Memory with Language Models”. arXiv, 2026, http://arxiv.org/abs/2605.14177v1.
APA
Chopra, H., Chintalapudi, K. K., Nath, S., White, R. W., & Shah, C. (2026). Thinking Ahead: Prospection-Guided Retrieval of Memory with Language Models. arXiv. http://arxiv.org/abs/2605.14177v1
Chicago
Chopra, H., K. K. Chintalapudi, S. Nath, R. W. White, and C. Shah. 2026. “Thinking Ahead: Prospection-Guided Retrieval of Memory with Language Models”. arXiv. http://arxiv.org/abs/2605.14177v1.
Harvard
Chopra, H. et al. (2026) “Thinking Ahead: Prospection-Guided Retrieval of Memory with Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2605.14177v1.
Vancouver
1. Chopra H, Chintalapudi KK, Nath S, White RW, Shah C (2026) Thinking Ahead: Prospection-Guided Retrieval of Memory with Language Models. arXiv

BibTeX

@article{chopra2026thinking,
  title = {Thinking Ahead: Prospection-Guided Retrieval of Memory with Language Models},
  author = {Chopra, Harshita and Chintalapudi, Krishna Kant and Nath, Suman and White, Ryen W. and Shah, Chirag},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2605.14177v1},
  eprint = {2605.14177}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/