Retrieval with Learned Similarities
Bailu DingJiaqi Zhai
Develops Mixture-of-Logits and an approximate top-k search algorithm to enable efficient retrieval with complex learned similarities, cutting latency by up to 66x while maintaining over 99% recall across recommendation and question answering tasks.
Large-scale internet applications, including recommendation engines, search systems, and natural language question answering, rely on fast initial retrieval to identify relevant candidates from catalogs containing millions or billions of items. Most standard systems, such as vector databases, employ simple dot-product similarity search to match queries with items. While efficient, dot products impose an expressive bottleneck that degrades retrieval accuracy. Recent advanced methods adopt learned similarity functions or generative indexing to boost relevance, but these models are computationally expensive, complex to index, and difficult to deploy under strict latency limits.
The article evaluates Mixture-of-Logits as a universal similarity framework and demonstrates how to achieve highly accurate, low-latency top-item retrieval across diverse applications. To improve practical training, the authors introduce a mutual information-based load balancing loss for conditional computations in Mixture-of-Logits. They also propose and benchmark exact two-pass and approximate two-stage search algorithms designed to operate efficiently on modern graphics processing units.
The research evaluates these methods across recommendation benchmarks containing up to 674,000 items and the Natural Questions benchmark of roughly 110,000 documents. The authors test sequential recommendation architectures, such as SASRec and HSTU, alongside language models finetuned from T5. They measure retrieval accuracy using Hit Rate and Mean Reciprocal Rank, and evaluate query latency on graphics processors across exact brute-force and approximate retrieval techniques.
The evaluation yielded several key findings. First, the article establishes mathematically that Mixture-of-Logits is a universal approximator capable of representing any complex similarity function. Second, adding Mixture-of-Logits to recommendation systems improved top-1 Hit Rate by an average of 29.1%, top-10 Hit Rate by 16.3%, and Mean Reciprocal Rank by 18.1% compared to standard dot-product baselines. Third, on the Natural Questions dataset, the approach achieved a top-100 Hit Rate of 97.0%, outperforming state-of-the-art generative and dense retrieval models. Fourth, the proposed approximate search methods achieved up to a 66-fold reduction in latency compared to exact brute-force search while preserving over 99% of the retrieval accuracy.
These results demonstrate that organizations do not need to choose between the high accuracy of complex learned similarity models and the fast execution of dot-product vector search. Mixture-of-Logits leverages the parallel compute power of modern accelerators to provide high throughput at latency budgets comparable to standard vector search. Furthermore, because the retrieval algorithms interface naturally with existing vector search operations, organizations can upgrade retrieval quality without completely redesigning their underlying database infrastructure.
Engineering and infrastructure teams should consider adopting the proposed Retrieval with Learned Similarities framework for accelerator-based search pipelines. Teams can select approximate search heuristics based on their specific workload needs: the average dot-product heuristic offers the lowest latency and is agnostic to the number of embedding components, while combined candidate generation provides the highest accuracy when recall requirements are exceptionally stringent. Before enterprise-scale rollout, organizations should conduct pilot testing on internal billion-scale datasets to validate memory overheads, indexing costs, and custom hardware kernel optimizations.
The findings are supported by consistent results across both structured recommendation datasets and unstructured language tasks. However, the study focuses on datasets of up to hundreds of thousands of items; exact performance at extreme multi-billion-item scales will depend on network bandwidth, distributed infrastructure, and hardware configurations.
- Paper: Accelerating Large-Scale Inference with Anisotropic Vector Quantization, Ruiqi Guo et al. (2020). Provides the foundational vector quantization and fast Maximum Inner Product Search (MIPS) approximation techniques that the source generalizes to non-linear learned similarity functions.
- Paper: ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction, Keshav Santhanam et al. (2022). Introduces late-interaction and multi-vector retrieval mechanisms, establishing the expressiveness and indexing trade-offs that motivate Mixture-of-Logits similarities.
- Paper: ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT, Omar Khattab et al. (2020). Establishes contextualized late-interaction scoring across token embeddings, an expressive multi-vector retrieval paradigm that the source seeks to generalize and accelerate.
- Paper: Transformer Memory as a Differentiable Search Index, Yi Tay et al. (2022). Introduces generative retrieval via direct item identifier decoding, providing one of the key advanced learned similarity settings analyzed in the source.
- Paper: DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning, Hussein Hazimeh et al. (2021). Covers differentiable gating and expert selection in mixture-of-experts architectures, underpinning the design and load-balancing objectives in Mixture-of-Logits scoring.
- Paper: On the Representation Collapse of Sparse Mixture of Experts, Zewen Chi et al. (2022). Examines representation collapse and load balancing in mixture-of-experts models, which directly motivates the mutual-information load-balancing loss utilized in the source.
- Paper: Sampling-bias-corrected neural modeling for large corpus item recommendations, Xinyang Yi et al. (2019). Analyzes large-scale neural retrieval and logit adjustments in two-tower recommendation architectures, which the source builds upon for sequential retrieval.
- Paper: On the Theoretical Limitations of Embedding-Based Retrieval, Orion Weller et al. (2026). Investigates the theoretical expressiveness and capacity bounds of embedding-based retrieval, providing formal limits for the learned similarity representations developed in the source.
- Paper: Retrieval Needs Multivectors: An Exponential Separation, Mihir Agarwal et al. (2026). Proves an exponential representational separation between single-vector and multi-vector retrieval architectures, theoretically reinforcing the expressive similarity modeling demonstrated in the source.
- Paper: Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale, Siddharth Gollapudi et al. (2026). Examines whether large language models can perform in-context retrieval across million-token scales, exploring an alternative learned retrieval paradigm beyond indexed similarity approximations.
- Paper: MeMo: Memory as a Model, Ryan Wei Heng Quek et al. (2026). Extends learned retrieval-augmented reasoning by introducing parametric memory models that replace standard similarity indexing pipelines entirely.
