RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder

Shitao XiaoZheng LiuYingxia ShaoZhao Cao

article2022EMNLP202 citations

Proposes an asymmetric masked auto-encoder pre-training framework that forces language models to generate superior sentence embeddings for dense retrieval, establishing state-of-the-art performance on BEIR and MS MARCO benchmarks.

Listen

Dense text retrieval systems are critical components of modern search engines, recommendation systems, and web applications. While large-scale language models have advanced information retrieval, most standard models rely on token-level pre-training tasks that fail to build strong sentence-level semantic representations. Existing solutions often use self-contrastive learning, which requires computationally expensive negative sampling and artificial data generation, or standard auto-encoding methods that do not extract sufficient training signals from input text.

The article demonstrates and evaluates RetroMAE, a novel pre-training framework based on masked auto-encoding specifically designed to produce superior sentence embeddings for dense retrieval. The central objective is to force the model to capture deep semantic meaning by making the text reconstruction task significantly more demanding while maintaining high computational efficiency.

The proposed framework employs an asymmetric architecture and asymmetric masking strategy. A standard 12-layer encoder processes input text with a moderate masking ratio of 15% to 30% to generate a comprehensive sentence embedding. An extremely simplified single-layer decoder then reconstructs the original sentence using an aggressive masking ratio of 50% to 70% alongside the generated sentence embedding. The approach also incorporates an enhanced decoding mechanism that enables full token reconstruction across diversified contexts without requiring negative samples or complex data augmentation. The authors evaluated the model against numerous generic and retrieval-oriented baseline models across the 18-dataset BEIR zero-shot benchmark, as well as supervised benchmarks including MS MARCO and Natural Questions using standard hardware.

The evaluation produced several key findings. First, in zero-shot retrieval across diverse domains, the framework achieved an average score of 45.2 on the BEIR benchmark, outperforming the strongest baseline model by 4.5 percentage points. Second, under supervised fine-tuning on MS MARCO and Natural Questions, it consistently outperformed existing models across standard ranking metrics. Third, when paired with standard knowledge distillation on the MS MARCO benchmark, the model achieved a top ranking score of 41.6, surpassing sophisticated dense retrieval systems such as ColBERTv2 and ERNIE-Search. Finally, ablation analyses confirmed that a single-layer decoder paired with high decoder masking and enhanced decoding delivers superior results compared to deeper decoder architectures.

These findings indicate that significant performance gains in dense retrieval can be achieved through better pre-training task design rather than simply increasing model size or pre-training data volume. The asymmetric design ensures computational efficiency during training while providing superior transferability across specialized domains like biomedical search, fact checking, and question answering without requiring extensive fine-tuning.

Organizations developing search and retrieval applications should consider adopting masked auto-encoder architectures with asymmetric masking for sentence embedding pipelines. For implementation, teams should maintain a lightweight, single-layer decoder and moderate-to-aggressive masking parameters to maximize representation quality while keeping computational costs manageable.

Confidence in these findings is supported by rigorous evaluations across established industry benchmarks and multiple baseline models. However, the study evaluated only standard base-sized models trained on moderate corpus volumes. Stakeholders should note that performance characteristics on very large model scales and massive multi-modal datasets remain areas for further empirical validation.

arXiv: 2205.12035staoxiao/RetroMAE

No sufficiently relevant recommendations were found.

Cover for RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder

Abstract

Despite pre-training’s progress in many important NLP tasks, it remains to explore effective pre-training strategies for dense retrieval. In this paper, we propose RetroMAE, a new retrieval oriented pre-training paradigm based on Masked Auto-Encoder (MAE). RetroMAE is highlighted by three critical designs. 1) A novel MAE workflow, where the input sentence is polluted for encoder and decoder with different masks. The sentence embedding is generated from the encoder’s masked input; then, the original sentence is recovered based on the sentence embedding and the decoder’s masked input via masked language modeling. 2) Asymmetric model structure, with a full-scale BERT like transformer as encoder, and a one-layer transformer as decoder. 3) Asymmetric masking ratios, with a moderate ratio for encoder: 1530%, and an aggressive ratio for decoder: 5070%. Our framework is simple to realize and empirically competitive: the pre-trained models dramatically improve the SOTA performances on a wide range of dense retrieval benchmarks, like BEIR and MS MARCO. The source code and pre-trained models are made publicly available at https://github.com/staoxiao/RetroMAE so as to inspire more interesting research.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Methodology
  • 3.1 Encoding
  • 3.2 Decoding
  • 3.3 Enhanced Decoding
  • 4 Experimental Studies
  • 4.1 Experiment Settings
  • 4.2 Main Results
  • 4.3 Ablation Studies
  • 5 Conclusion
  • 6 Limitations
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — RetroMAE uses a sentence-embedding bottleneck with asymmetric masking and model capacity

    model/method

    RetroMAE pre-trains a dense-retrieval encoder by reconstructing an input sentence from a sentence embedding and a separately masked version of that sentence. Let X=(x_1,{x}N)beasequenceofbe a sequence ofNtokens.Twoindependentlymaskedinputsaremade:theencoderinputmasksamoderatefractionoftokens,whilethedecoderinputmasksalargerfraction.TheencoderisaBERT−base−stylebidirectionaltransformerwith12layersandhiddensize768;itsfinal‘[CLS]‘representationisthesentenceembeddingtokens. Two independently masked inputs are made: the encoder input masks a moderate fraction of tokens, while the decoder input masks a larger fraction. The encoder is a BERT-base-style bidirectional transformer with 12 layers and hidden size 768; its final `[CLS]` representation is the sentence embeddingh=\Phi{\mathrm{enc}}(\widetilde X_{\mathrm{enc}})_{[\mathrm{CLS}]}$. The encoder masking ratio is 15–30% (0.3 in the default implementation).

    For basic decoding, the decoder receives hh followed by the decoder-input token embeddings with their position embeddings. It is a single-layer transformer, and its masked-token predictions are trained with cross-entropy. The decoder masking ratio is 50–70% (0.5 by default). The total pre-training loss is the sum of the decoder reconstruction loss and the BERT-style encoder masked-language-model loss, L=Ldec+LencL=L_{\mathrm{dec}}+L_{\mathrm{enc}}. The paper also proposes enhanced decoding, which changes the decoder’s attention and reconstructs every input token rather than only decoder-masked tokens.

  2. Knowl 2 — Enhanced decoding reconstructs every token from a position-specific context

    model/method

    RetroMAE’s enhanced decoder uses two streams for a sentence of NN tokens. Let hh be the encoder’s sentence embedding, e(xi)e(x_i) the embedding of token xix_i, and pip_i the position embedding at index ii. The query stream contains the sentence embedding at each position with a position embedding added; the context stream contains the sentence embedding followed by all original token embeddings with position embeddings:

    H1=[h+p0,…,h+pN],H2=[h,e(x1)+p1,…,e(xN)+pN].H_1=[h+p_0,\ldots,h+p_N],\qquad H_2=[h,e(x_1)+p_1,\ldots,e(x_N)+p_N].

    A single transformer layer computes each token’s reconstruction using a position-specific attention mask. For token position i∈{1,…,N}i\in\{1,\ldots,N\}, let SiS_i be a sampled set of visible token positions that excludes ii. Among context indices j∈{0,…,N}j\in\{0,\ldots,N\}, attention is allowed to the sentence-embedding position j=0j=0 and to positions in SiS_i; all other positions, including the token’s own position ii, are masked. Equivalently, the additive attention mask has value 00 at allowed entries and −∞-\infty elsewhere. The attention output, together with the query stream’s residual representation, is used to predict xix_i. Cross-entropy is applied to all NN tokens. Thus each token is reconstructed without attending to itself and with a sampled context, making all input tokens training targets and providing position-dependent reconstruction contexts.

  3. Knowl 3 — RetroMAE leads the reported BEIR transfer results

    data/table

    The paper evaluates retrieval using NDCG@10 on the 18 BEIR datasets, reporting transfer after MS MARCO fine-tuning. The table compares RetroMAE with the strongest non-RetroMAE baseline on each dataset; the final row reports the mean across datasets. RetroMAE has the highest score on 12 datasets, ties for the highest score on SCIDOCS, and reaches a mean of 0.452 versus 0.407 for the strongest baseline overall, Condenser.

    Dataset Best competing baseline Baseline RetroMAE
    TREC-COVID Condenser 0.750 0.772
    BioASQ Condenser 0.322 0.421
    NFCorpus LaPraDoR 0.310 0.308
    NQ Condenser 0.486 0.518
    HotpotQA SEED 0.541 0.635
    FiQA-2018 DeBERTa 0.299 0.316
    Signal-1M(RT) SimCSE 0.262 0.265
    TREC-NEWS RoBERTa 0.385 0.428
    Robust04 RoBERTa 0.384 0.447
    ArguAna LaPraDoR 0.499 0.433
    Touche-2020 RoBERTa 0.299 0.237
    CQADupStack Condenser 0.347 0.317
    Quora Condenser 0.853 0.847
    DBPedia Condenser 0.339 0.390
    SCIDOCS LaPraDoR 0.150 0.150
    FEVER Condenser 0.691 0.774
    Climate-FEVER RoBERTa 0.222 0.232
    SciFact Condenser 0.593 0.653
    Average Condenser 0.407 0.452
  4. Knowl 4 — RetroMAE improves supervised retrieval after DPR or ANCE fine-tuning

    data/table

    The paper evaluates the pre-trained encoders after supervised fine-tuning with DPR and ANCE on MS MARCO passage retrieval and Natural Questions. The following values compare RetroMAE with Condenser, the strongest competing pre-trained model for the listed metrics in these results. For MS MARCO, the reported metrics are MRR@10 and Recall@10; for Natural Questions, the comparison below uses Recall@10. RetroMAE is higher in every listed comparison.

    Fine-tuning Model MS MARCO MRR@10 MS MARCO R@10 Natural Questions R@10
    DPR Condenser 0.3357 0.6082 0.7562
    DPR RetroMAE 0.3553 0.6356 0.7704
    ANCE Condenser 0.3635 0.6388 0.7903
    ANCE RetroMAE 0.3822 0.6677 0.8044

    The gains over Condenser include 0.0196 in MS MARCO MRR@10 and 0.0142 in Natural Questions Recall@10 with DPR, and 0.0187 and 0.0141, respectively, with ANCE.

  5. Knowl 5 — Knowledge-distilled RetroMAE reaches 0.416 MRR@10 on MS MARCO

    empirical result

    In the paper’s MS MARCO passage-retrieval evaluation with knowledge distillation, RetroMAE reaches MRR@10 of 0.416, Recall@10 of 0.709, Recall@100 of 0.927, and Recall@1000 of 0.988. The reported MRR@10 scores for comparison systems are AR2 0.395, ColBERTv2 0.397, RocketQAv2 0.388, and ERNIE-Search 0.401. RetroMAE therefore exceeds the strongest listed MRR@10 score, ERNIE-Search’s 0.401, by 0.015. For RetroMAE, distillation trains a BERT-base cross-encoder on hard negatives returned by an ANCE-fine-tuned bi-encoder, then fine-tunes the bi-encoder to minimize KL divergence from the cross-encoder.

  6. Knowl 6 — Pre-training data and implementation settings

    experimental setup

    RetroMAE’s pre-training uses English Wikipedia and BookCorpus, the corpora used for BERT, and also examines MS MARCO as in-domain pre-training data. The paper reports that MS MARCO pre-training is important for MS MARCO retrieval but unnecessary for results on other evaluation datasets. The encoder uses 12 bidirectional transformer layers, hidden size 768, and a 30,522-token vocabulary; the decoder has one transformer layer. Default encoder and decoder masking ratios are 0.3 and 0.5. Training runs for 8 epochs with AdamW, learning rate 10−410^{-4}, and batch size 32 per device, on 8 NVIDIA A100 40GB GPUs.

    Evaluation covers MS MARCO passage retrieval, Natural Questions, and BEIR. The reported MS MARCO passage task has 502,939 training queries, 6,980 development queries, and 8.8 million candidate passages. Natural Questions has 79,168 training queries, 8,757 development queries, 3,610 test queries, and 21,015,324 candidate Wikipedia passages. BEIR transfer is evaluated on 18 datasets after MS MARCO fine-tuning. DPR and ANCE are used for supervised fine-tuning.

  7. Knowl 7 — Enhanced decoding improves DPR-fine-tuned retrieval metrics

    data/table

    An ablation compares the enhanced decoder with basic decoding under DPR fine-tuning, using the same RetroMAE setup. Enhanced decoding improves every listed metric on both MS MARCO and Natural Questions. The comparison is especially relevant because basic decoding supervises only decoder-masked tokens, whereas enhanced decoding predicts every input token from position-specific contexts.

    Dataset Decoding MRR@10 MRR@100 R@100 R@1000
    MS MARCO Basic 0.3462 0.6218 0.8813 0.9725
    MS MARCO Enhanced 0.3553 0.6356 0.8922 0.9763
    Natural Questions Basic 0.7562 0.8291 0.8540 0.8759
    Natural Questions Enhanced 0.7704 0.8399 0.8604 0.8812
  8. Knowl 8 — Decoder and encoder masking ratios have different useful ranges

    data/table

    DPR-fine-tuning ablations vary decoder masking with and without enhanced decoding, and encoder masking with enhanced decoding. Values below are MRR@10 on MS MARCO and Natural Questions. Aggressive decoder masking outperforms the 0.15 setting: the best tested ratio is 0.5 with enhanced decoding and 0.7 without it. For the encoder, increasing masking from 0.15 to 0.3 improves MS MARCO MRR@10 slightly and leaves Natural Questions nearly unchanged, whereas a 0.9 encoder ratio reduces performance on both datasets.

    Masking factor Setting MS MARCO MRR@10 Natural Questions MRR@10
    Decoder, enhanced 0.15 0.3496 0.7608
    Decoder, enhanced 0.50 0.3553 0.7704
    Decoder, enhanced 0.90 0.3514 0.7609
    Decoder, basic 0.15 0.3440 0.7519
    Decoder, basic 0.70 0.3508 0.7593
    Decoder, basic 0.90 0.3441 0.7576
    Encoder, enhanced 0.15 0.3501 0.7703
    Encoder, enhanced 0.30 0.3553 0.7704
    Encoder, enhanced 0.90 0.3365 0.7599
  9. Knowl 9 — Increasing decoder depth does not improve retrieval in the tested ablation

    empirical result

    The paper compares one-, two-, and three-layer decoders without enhanced decoding, because enhanced decoding is only applicable to a single-layer transformer in this setup. Under DPR fine-tuning, increasing decoder depth produces no empirical gain in the reported MRR@10 values, despite the greater decoder size; the authors identify the one-layer decoder as the preferred design.

    Decoder transformer layers MS MARCO MRR@10 Natural Questions MRR@10
    1 0.3462 0.7562
    2 0.3446 0.7561
    3 0.3439 0.7563
  10. Knowl 10 — Evaluation does not establish scaling behavior beyond BERT-base and moderate data

    limitation

    The paper’s empirical studies use BERT-base-scale transformers and a moderate amount of pre-training data, partly because of computational-resource constraints. The experiments therefore do not establish how RetroMAE behaves with larger networks or substantially more pre-training data; the authors identify both scaling directions as requiring further study.

Coverage note — The per-method baseline cells in the BEIR and supervised evaluation tables are omitted; the knowls retain BEIR’s per-dataset strongest-baseline comparisons and the principal supervised comparisons needed to represent the reported findings.

References

  1. 1.Hangbo Bao, Li Dong, and Furu Wei. 2021. Beit: Bert pre-training of image transformers. arXiv preprint arXiv:2106.08254.
  2. 2.Wei-Cheng Chang, Felix X Yu, Yin-Wen Chang, Yiming Yang, and Sanjiv Kumar. 2020. Pre-training tasks for embedding-based large-scale retrieval. arXiv preprint arXiv:2002.03932.
  3. 3.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR.
  4. 4.Yung-Sung Chuang, Rumen Dangovski, Hongyin Luo, Yang Zhang, Shiyu Chang, Marin Soljacic, Shang-Wen Li, Wen-tau Yih, Yoon Kim, and James R. Glass. 2022. Diffcse: Difference-based contrastive learning for sentence embeddings. CoRR, abs/2204.10298.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,, pages 4171–4186. Association for Computational Linguistics.
  6. 6.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019. Unified language model pre-training for natural language understanding and generation. Advances in Neural Information Processing Systems, 32.
  7. 7.Alaaeldin El-Nouby, Gautier Izacard, Hugo Touvron, Ivan Laptev, Herve Jegou, and Edouard Grave. 2021. Are large-scale datasets necessary for self-supervised pre-training? arXiv preprint arXiv:2112.10740.
  8. 8.Luyu Gao and Jamie Callan. 2021. Condenser: a pre-training architecture for dense retrieval. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 981–993.
  9. 9.Luyu Gao and Jamie Callan. 2022. Unsupervised corpus aware language model pre-training for dense passage retrieval. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2843–2853, Dublin, Ireland.
  10. 10.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. Simcse: Simple contrastive learning of sentence embeddings. arXiv preprint arXiv:2104.08821.
  11. 11.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. Realm: Retrieval-augmented language model pre-training. arXiv preprint arXiv:2002.08909.
  12. 12.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020a. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9729–9738.
  13. 13.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020b. Deberta: Decoding-enhanced bert with disentangled attention. arXiv preprint arXiv:2006.03654.
  14. 14.Herve Jegou, Matthijs Douze, and Cordelia Schmid. 2010. Product quantization for nearest neighbor search. IEEE transactions on pattern analysis and machine intelligence, 33(1):117–128.
  15. 15.Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy. 2020. Spanbert: Improving pre-training by representing and predicting spans. Transactions of the Association for Computational Linguistics, 8:64–77.
  16. 16.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, pages 6769–6781.
  17. 17.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.
  18. 18.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461.
  19. 19.Chunyuan Li, Xiang Gao, Yuan Li, Baolin Peng, Xiujun Li, Yizhe Zhang, and Jianfeng Gao. 2020. Optimus: Organizing sentences via pre-trained modeling of a latent space. arXiv preprint arXiv:2004.04092.
  20. 20.Jimmy Lin, Rodrigo Nogueira, and Andrew Yates. 2021. Pretrained transformers for text ranking: Bert and beyond. Synthesis Lectures on Human Language Technologies, 14(4):1–325.
  21. 21.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  22. 22.Shuqi Lu, Di He, Chenyan Xiong, Guolin Ke, Waleed Malik, Zhicheng Dou, Paul Bennett, Tie-Yan Liu, and Arnold Overwijk. 2021. Less is more: Pretrain a strong Siamese encoder for dense text retrieval using a weak decoder. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2780–2791.
  23. 23.Wenhao Lu, Jian Jiao, and Ruofei Zhang. 2020. Twinbert: Distilling knowledge to twin-structured bert models for efficient retrieval. arXiv preprint arXiv:2002.06275.
  24. 24.Yuxiang Lu, Yiding Liu, Jiaxiang Liu, Yunsheng Shi, Zhengjie Huang, Shikun Feng Yu Sun, Hao Tian, Hua Wu, Shuaiqiang Wang, Dawei Yin, et al. 2022. Ernie-search: Bridging cross-encoder with dual-encoder via self on-the-fly distillation for dense passage retrieval. arXiv preprint arXiv:2205.09153.
  25. 25.Yi Luan, Jacob Eisenstein, Kristina Toutanova, and Michael Collins. 2021. Sparse, dense, and attentional representations for text retrieval. Transactions of the Association for Computational Linguistics, 9:329–345.
  26. 26.Yu A Malkov and Dmitry A Yashunin. 2018. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE transactions on pattern analysis and machine intelligence, 42(4):824–836.
  27. 27.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. Ms marco: A human generated machine reading comprehension dataset. In CoCo@ NIPS.
  28. 28.Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Y Zhao, Yi Luan, Keith B Hall, Ming-Wei Chang, et al. 2021. Large dual encoders are generalizable retrievers. arXiv preprint arXiv:2112.07899.
  29. 29.Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2020. Rocketqa: An optimized training approach to dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2010.08191.
  30. 30.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training.
  31. 31.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683.
  32. 32.Ruiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao, Qiaoqiao She, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2021. Rocketqav2: A joint training method for dense passage retrieval and passage re-ranking. arXiv preprint arXiv:2110.07367.
  33. 33.Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2021. Colbertv2: Effective and efficient retrieval via lightweight late interaction. arXiv preprint arXiv:2112.01488.
  34. 34.Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Xuyi Chen, Han Zhang, Xin Tian, Danxiang Zhu, Hao Tian, and Hua Wu. 2019. Ernie: Enhanced representation through knowledge integration. arXiv preprint arXiv:1904.09223.
  35. 35.Nandan Thakur, Nils Reimers, Andreas Rucklé, Abhishek Srivastava, and Iryna Gurevych. 2021. Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models. arXiv preprint arXiv:2104.08663.
  36. 36.Kexin Wang, Nils Reimers, and Iryna Gurevych. 2021. Tsdae: Using transformer-based sequential denoising auto-encoder for unsupervised sentence embedding learning. arXiv preprint arXiv:2104.06979.
  37. 37.Shitao Xiao, Zheng Liu, Weihao Han, Jianjin Zhang, Defu Lian, Yeyun Gong, Qi Chen, Fan Yang, Hao Sun, Yingxia Shao, et al. 2022a. Distill-vq: Learning retrieval oriented vector quantization by distilling knowledge from dense embeddings. arXiv preprint arXiv:2204.00185.
  38. 38.Shitao Xiao, Zheng Liu, Yingxia Shao, Tao Di, Bhuvan Middha, Fangzhao Wu, and Xing Xie. 2022b. Training large-scale news recommenders with pretrained language models in the loop. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 4215–4225.
  39. 39.Shitao Xiao, Zheng Liu, Yingxia Shao, Defu Lian, and Xing Xie. 2021. Matching-oriented product quantization for ad-hoc retrieval. arXiv preprint arXiv:2104.07858.
  40. 40.Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul Bennett, Junaid Ahmed, and Arnold Overwijk. 2020. Approximate nearest neighbor negative contrastive learning for dense text retrieval. arXiv preprint arXiv:2007.00808.
  41. 41.Canwen Xu, Daya Guo, Nan Duan, and Julian McAuley. 2022. Laprador: Unsupervised pretrained dense retriever for zero-shot text retrieval. arXiv preprint arXiv:2203.06169.
  42. 42.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32.
  43. 43.Hang Zhang, Yeyun Gong, Yelong Shen, Jiancheng Lv, Nan Duan, and Weizhu Chen. 2021. Adversarial retriever-ranker for dense text retrieval. arXiv preprint arXiv:2110.03611.
  44. 44.Jianjin Zhang, Zheng Liu, Weihao Han, Shitao Xiao, Ruicheng Zheng, Yingxia Shao, Hao Sun, Hanqing Zhu, Premkumar Srinivasan, Weiwei Deng, et al. 2022. Uni-retriever: Towards learning the unified embedding based retriever in bing sponsored search. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 4493–4501.

Citation

MLA
Xiao, S., et al. “RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 538–48, https://doi.org/10.18653/v1/2022.emnlp-main.35.
APA
Xiao, S., Liu, Z., Shao, Y., & Cao, Z. (2022). RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 538–548. https://doi.org/10.18653/v1/2022.emnlp-main.35
Chicago
Xiao, S., Z. Liu, Y. Shao, and Z. Cao. 2022. “RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 538–48. https://doi.org/10.18653/v1/2022.emnlp-main.35.
Harvard
Xiao, S. et al. (2022) “RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 538–548. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.35.
Vancouver
1. Xiao S, Liu Z, Shao Y, Cao Z (2022) RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 538–548

BibTeX

@inproceedings{xiao-etal-2022-retromae,
    title = "{R}etro{MAE}: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder",
    author = "Xiao, Shitao  and
      Liu, Zheng  and
      Shao, Yingxia  and
      Cao, Zhao",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.35/",
    doi = "10.18653/v1/2022.emnlp-main.35",
    pages = "538--548"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/