RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder
Shitao XiaoZheng LiuYingxia ShaoZhao Cao
Proposes an asymmetric masked auto-encoder pre-training framework that forces language models to generate superior sentence embeddings for dense retrieval, establishing state-of-the-art performance on BEIR and MS MARCO benchmarks.
Dense text retrieval systems are critical components of modern search engines, recommendation systems, and web applications. While large-scale language models have advanced information retrieval, most standard models rely on token-level pre-training tasks that fail to build strong sentence-level semantic representations. Existing solutions often use self-contrastive learning, which requires computationally expensive negative sampling and artificial data generation, or standard auto-encoding methods that do not extract sufficient training signals from input text.
The article demonstrates and evaluates RetroMAE, a novel pre-training framework based on masked auto-encoding specifically designed to produce superior sentence embeddings for dense retrieval. The central objective is to force the model to capture deep semantic meaning by making the text reconstruction task significantly more demanding while maintaining high computational efficiency.
The proposed framework employs an asymmetric architecture and asymmetric masking strategy. A standard 12-layer encoder processes input text with a moderate masking ratio of 15% to 30% to generate a comprehensive sentence embedding. An extremely simplified single-layer decoder then reconstructs the original sentence using an aggressive masking ratio of 50% to 70% alongside the generated sentence embedding. The approach also incorporates an enhanced decoding mechanism that enables full token reconstruction across diversified contexts without requiring negative samples or complex data augmentation. The authors evaluated the model against numerous generic and retrieval-oriented baseline models across the 18-dataset BEIR zero-shot benchmark, as well as supervised benchmarks including MS MARCO and Natural Questions using standard hardware.
The evaluation produced several key findings. First, in zero-shot retrieval across diverse domains, the framework achieved an average score of 45.2 on the BEIR benchmark, outperforming the strongest baseline model by 4.5 percentage points. Second, under supervised fine-tuning on MS MARCO and Natural Questions, it consistently outperformed existing models across standard ranking metrics. Third, when paired with standard knowledge distillation on the MS MARCO benchmark, the model achieved a top ranking score of 41.6, surpassing sophisticated dense retrieval systems such as ColBERTv2 and ERNIE-Search. Finally, ablation analyses confirmed that a single-layer decoder paired with high decoder masking and enhanced decoding delivers superior results compared to deeper decoder architectures.
These findings indicate that significant performance gains in dense retrieval can be achieved through better pre-training task design rather than simply increasing model size or pre-training data volume. The asymmetric design ensures computational efficiency during training while providing superior transferability across specialized domains like biomedical search, fact checking, and question answering without requiring extensive fine-tuning.
Organizations developing search and retrieval applications should consider adopting masked auto-encoder architectures with asymmetric masking for sentence embedding pipelines. For implementation, teams should maintain a lightweight, single-layer decoder and moderate-to-aggressive masking parameters to maximize representation quality while keeping computational costs manageable.
Confidence in these findings is supported by rigorous evaluations across established industry benchmarks and multiple baseline models. However, the study evaluated only standard base-sized models trained on moderate corpus volumes. Stakeholders should note that performance characteristics on very large model scales and massive multi-modal datasets remain areas for further empirical validation.
- Paper: Unsupervised Corpus Aware Language Model Pre-training for Dense Passage Retrieval, Luyu Gao et al. (2022). Read coCondenser first to see an earlier corpus-aware pretraining approach for dense retrieval, which provides a direct point of comparison for RetroMAE’s retrieval-oriented pretraining design.
- Paper: Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval, Lee Xiong et al. (2021). ANCE establishes the dense-retrieval training setting and hard-negative challenges that help make RetroMAE’s pretraining objective and retrieval gains meaningful.
- Paper: Unsupervised Dense Information Retrieval with Contrastive Learning, Gautier Izacard et al. (2021). Contriever shows how contrastive pretraining can produce dense retrievers, clarifying the alternative training paradigm against which RetroMAE’s masked-autoencoder approach can be understood.
No sufficiently relevant recommendations were found.
