keyword
memory retrieval
Memory retrieval in computational and artificial intelligence systems is the process of searching, accessing, and extracting relevant stored information from an external or persistent memory store to inform downstream processing or response generation. In machine learning and natural language processing architectures, it allows models to dynamically query historical context, factual knowledge, or past representations, such as vector embeddings and key-value pairs, without requiring updates to the underlying model parameters. This mechanism enables systems to overcome fixed context-length constraints, utilize long-term interaction histories, and ground generative outputs in relevant data retrieved through techniques such as semantic similarity search, structured graph traversal, or nearest-neighbor lookups.
3 items

A Multi-Task Embedder For Retrieval Augmented LLMs
Peitian Zhang, Zheng Liu, Shitao Xiao, Zhicheng Dou, Jian-Yun Nie
Why you should read this
Presents LLM-Embedder, a unified multi-task embedding model trained via rank-aware rewards and graded distillation to optimize retrieval across diverse language model augmentation scenarios including knowledge, memory, examples, and tools.
LLMs confront inherent limitations in terms of its knowledge, memory, and action. The retrieval augmentation stands as a vital mechanism to address these limitations, which brings in useful information from external sources to augment the LLM. However, existing retrieval methods encounter two pressing issues. On one hand, the general retrievers are not properly optimized for retrieval augmentation hence exhibit limited effectiveness; on the other hand, the task-specific retrievers excel in the targeted retrieval augmentation scenario, while lack the versatility to handle diverse scenarios. In this work, we propose LLM-Embedder for the unified support of diverse retrieval augmentation scenarios. Our method presents three technical contributions. Firstly, we introduce a new re-ward formulation, namely rank-aware reward. It exploits the ranking position of the desired output among N sampled outputs from the LLM, which leads to fine-grained and robust computation of reward from the LLM's feedback. Secondly, we design a novel distillation objective, called graded distillation. It incorporates both the absolute value and the relative order of the reward for more sufficient utilization of the LLM's feedback. Thirdly, we systematically optimize the multi-task learning, which effectively unifies the multiple retrieval functionalities into one model. In our experiment, LLM-Embedder notably improves the LLM's performances in various downstream tasks, and outperforms both general and task-specific retrievers with a substantial advantage. Our data, code, and model have been released at https://github.com/FlagOpen/FlagEmbedding.
Added
2026-10-02

Thinking Ahead: Prospection-Guided Retrieval of Memory with Language Models
Harshita Chopra, Krishna Chintalapudi, Suman Nath, Ryen W. White, Chirag Shah
Why you should read this
Proposes Prospection-Guided Retrieval, a framework that expands user queries into simulated future actions to retrieve semantically distant interaction memories, nearly tripling recall compared to standard retrieval-augmented generation on long-horizon dialogue tasks.
Long-horizon personalization requires dialogue assistants to retrieve user-specific facts from extended interaction histories. In practice, many relevant facts often have low semanticsimilarity to the query under dense retrieval. Standard Retrieval-Augmented Generation (RAG) and GraphRAG systems are still largely retrospective: they rely on embedding similarity to the query or on fixed graph traversals, so they often miss facts that matter for the user's needs but lie far from the query in embedding space. Inspired by prospection, the human ability to use imagined futures as cues for recall, we introduce Prospection-Guided Retrieval (PGR), which decouples retrieval from how memories are stored. Given a user query, PGR first expands the goal into a short Tree-of-Thought (ToT) or linear chain of plausible next steps, and uses these steps as retrieval probes rather than relying on the original query alone. The facts retrieved by these probes are then used to personalize the next round of prospection, enabling PGR to uncover additional memories that become relevant only after the simulation is grounded in the user's history. We also introduce MemoryQuest, a challenging multi-session benchmark in which each query is annotated with 3--5 dated reference facts subject to a low query-reference similarity constraint. Across 1,625 queries spanning 185 user profiles from 3 publicly available datasets, PGR-TOT substantially improves retrieval, including nearly 3x recall on MemoryQuest over the strongest baseline. In pairwise LLM-as-judge comparisons against baselines, PGR-generated responses are preferred on 89--98% of queries, with blinded human annotations on held-out subsets showing the same trend. Overall, the results demonstrate that explicit prospection yields large gains in long-horizon retrieval and response quality relative to similarity-only baselines.
Added
2026-09-29

Memorizing Transformers
Yuhuai Wu, Markus Norman Rabe, DeLesley Hutchins, Christian Szegedy
Why you should read this
Extends Transformer language models with an approximate k-nearest-neighbors memory over past key-value representations, enabling models to scale test-time context up to 262,000 tokens and immediately apply new knowledge without updating weights.
Language models typically need to be trained or finetuned in order to acquire new knowledge, which involves updating their weights. We instead envision language models that can simply read and memorize new data at inference time, thus acquiring new knowledge immediately. In this work, we extend language models with the ability to memorize the internal representations of past inputs. We demonstrate that an approximate kNN lookup into a non-differentiable memory of recent (key, value) pairs improves language modeling across various benchmarks and tasks, including generic webtext (C4), math papers (arXiv), books (PG-19), code (Github), as well as formal theorems (Isabelle). We show that the performance steadily improves when we increase the size of memory up to 262K tokens. On benchmarks including code and mathematics, we find that the model is capable of making use of newly defined functions and theorems during test time.
Added
2026-09-26
