Built independently by an author, for readers. Read the story and support ChapterPal

keyword

query generation

Query generation is the automated process of creating search queries, questions, or search strings using computational algorithms or language models to facilitate information retrieval. Unlike manual query formulation performed directly by human users, this technique programmatically produces text inputs designed to identify and extract relevant information from a target corpus. In retrieval-augmented generation and multi-step reasoning architectures, language models generate targeted sub-queries to break down complex tasks and fetch the precise external evidence needed to formulate an answer. In document indexing and machine learning workflows, query generation is frequently used for document expansion and synthetic dataset creation, where models predict the plausible questions a given passage could satisfy in order to resolve vocabulary mismatch between users and documents, enhance index representations, and train dense or generative retrieval models.

3 items

Multiview Identifiers Enhanced Generative Retrieval

Multiview Identifiers Enhanced Generative Retrieval

Yongqi Li, Nan Yang, Liang Wang, Furu Wei, Wenjie Li

OrganizationsHong Kong Polytechnic UniversityMicrosoft

Why you should read this

Proposes MINDER, a generative retrieval framework that combines titles, substrings, and synthetic pseudo-queries into multiview passage identifiers to achieve state-of-the-art document ranking performance across multiple benchmark datasets.

Instead of simply matching a query to pre-existing passages, generative retrieval generates identifier strings of passages as the retrieval target. At a cost, the identifier must be distinctive enough to represent a passage. Current approaches use either a numeric ID or a text piece (such as a title or substrings) as the identifier. However, these identifiers cannot cover a passage's content well. As such, we are motivated to propose a new type of identifier, synthetic identifiers, that are generated based on the content of a passage and could integrate contextualized information that text pieces lack. Furthermore, we simultaneously consider multiview identifiers, including synthetic identifiers, titles, and substrings. These views of identifiers complement each other and facilitate the holistic ranking of passages from multiple perspectives. We conduct a series of experiments on three public datasets, and the results indicate that our proposed approach performs the best in generative retrieval, demonstrating its effectiveness and robustness. The code is released at https://github.com/liyongqi67/MINDER.

Added

2026-10-05

Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation

Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation

Satyapriya Krishna, Kalpesh Krishna, Anhad Mohananey, Steven Schwarcz, Adam Stambler, Shyam Upadhyay, Manaal Faruqui

OrganizationsGoogleHarvard UniversityMeta

Why you should read this

Introduces FRAMES, a benchmark of multi-hop questions requiring information synthesis across multiple documents to evaluate retrieval-augmented generation systems simultaneously on factuality, retrieval, and complex reasoning.

Large Language Models (LLMs) have shown significant improvements across cognitive tasks, with an emerging application in enhancing retrieval-augmented generation (RAG) capabilities. These systems require LLMs to understand queries, retrieve relevant information, and synthesize accurate responses. Given their increasing real-world deployment, comprehensive evaluation is crucial. We propose FRAMES (Factuality, Retrieval, And reasoning MEasurement Set), a high-quality dataset designed to test LLMs’ factual responses, retrieval capabilities, and reasoning in generating final answers. Unlike previous work evaluating these abilities in isolation, FRAMES offers a unified framework for assessing LLM performance in end-to-end RAG scenarios. Our dataset comprises challenging multi-hop questions requiring integration of information from multiple sources. Baseline results show that even state-of-the-art LLMs struggle, achieving 0.408 accuracy without retrieval. However, our proposed multi-step retrieval pipeline significantly improves accuracy to 0.66 (>50% improvement). We aim to bridge evaluation gaps and assist in developing more robust RAG systems.

Added

2026-10-01