Built independently by an author, for readers. Read the story and support ChapterPal

keyword

self-contrastive learning

Self-contrastive learning is a machine learning representation technique in which a model learns meaningful feature representations by contrasting different internal views, states, or representations generated from the same input instance. Rather than relying entirely on heavy external data augmentations or separately collected positive pairs, this approach creates positive associations from internal variations of a single sample, such as different network layers, masking patterns, or stochastic perturbations like dropout. By pulling these self-generated positive representations closer while pushing representations of distinct instances apart, self-contrastive learning promotes a well-distributed latent space, reduces computational overhead, and improves downstream performance across tasks such as classification, sentence embedding, and information retrieval.

1 item

RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder

RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder

Shitao Xiao, Zheng Liu, Yingxia Shao, Zhao Cao

OrganizationsBeijing University of Posts and TelecommunicationsHuawei

Why you should read this

Proposes an asymmetric masked auto-encoder pre-training framework that forces language models to generate superior sentence embeddings for dense retrieval, establishing state-of-the-art performance on BEIR and MS MARCO benchmarks.

Despite pre-training’s progress in many important NLP tasks, it remains to explore effective pre-training strategies for dense retrieval. In this paper, we propose RetroMAE, a new retrieval oriented pre-training paradigm based on Masked Auto-Encoder (MAE). RetroMAE is highlighted by three critical designs. 1) A novel MAE workflow, where the input sentence is polluted for encoder and decoder with different masks. The sentence embedding is generated from the encoder’s masked input; then, the original sentence is recovered based on the sentence embedding and the decoder’s masked input via masked language modeling. 2) Asymmetric model structure, with a full-scale BERT like transformer as encoder, and a one-layer transformer as decoder. 3) Asymmetric masking ratios, with a moderate ratio for encoder: 15~30%, and an aggressive ratio for decoder: 50~70%. Our framework is simple to realize and empirically competitive: the pre-trained models dramatically improve the SOTA performances on a wide range of dense retrieval benchmarks, like BEIR and MS MARCO. The source code and pre-trained models are made publicly available at https://github.com/staoxiao/RetroMAE so as to inspire more interesting research.

Added

2026-09-30