Built independently by an author, for readers. Read the story and support ChapterPal

keyword

unsupervised adaptation

Unsupervised adaptation is a machine learning process that modifies a pre-trained model or its internal representations to perform effectively on a new target domain, data distribution, or task without using labeled target data. By relying entirely on unlabeled inputs or self-supervised learning objectives, this approach enables a model to overcome performance degradation caused by domain shift or dataset bias. It typically functions by aligning the statistical properties of feature representations between source and target distributions, learning domain-invariant representations through adversarial training, or optimizing pretext tasks that restructure embeddings for specific downstream applications. As a result, unsupervised adaptation allows models to generalize and transfer knowledge to novel environments while avoiding the high costs and logistical constraints associated with collecting ground-truth annotations.

5 items

Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval

Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval

Chaofan Li, Zheng Liu, Shitao Xiao, Yingxia Shao, Defu Lian

OrganizationsBeijing Academy of Artificial IntelligenceBeijing University of Posts and TelecommunicationsHong Kong Polytechnic UniversityUniversity of Science and Technology of China

Why you should read this

Develops an unsupervised adaptation method using embedding-based auto-encoding and auto-regression tasks to transform autoregressive language models into effective dense retrieval encoders, achieving state-of-the-art performance on MSMARCO and BEIR benchmarks.

Dense retrieval calls for discriminative embeddings to represent the semantic relationship between query and document. It may benefit from the using of large language models (LLMs), given LLMs’ strong capability on semantic understanding. However, the LLMs are learned by auto-regression, whose working mechanism is completely different from representing whole text as one discriminative embedding. Thus, it is imperative to study how to adapt LLMs properly so that they can be effectively initialized as the backbone encoder for dense retrieval. In this paper, we propose a novel approach, called Llama2Vec, which performs unsupervised adaptation of LLM for its dense retrieval application. Llama2Vec consists of two pretext tasks: EBAE (Embedding-Based Auto-Encoding) and EBAR (Embedding-Based Auto-Regression), where the LLM is prompted to reconstruct the input sentence and predict the next sentence based on its text embeddings. Llama2Vec is simple, lightweight, but highly effective. It is used to adapt LLaMA-2-7B on the Wikipedia corpus. With a moderate steps of adaptation, it substantially improves the model’s fine-tuned performances on a variety of dense retrieval benchmarks. Notably, it results in the new state-of-the-art performances on popular benchmarks, such as passage and document retrieval on MSMARCO, and zero-shot retrieval on BEIR. The model and source code will be made publicly available to facilitate the future research. Our model is available at https://github.com/FlagOpen/FlagEmbedding.

Added

2026-10-01

ReCo: Retrieve and Co-segment for Zero-shot Transfer

ReCo: Retrieve and Co-segment for Zero-shot Transfer

Gyungin Shin, Weidi Xie, Samuel Albanie

OrganizationsDepartment of EngineeringShanghai Jiao Tong UniversityUniversity of CambridgeUniversity of Oxford

Why you should read this

Proposes a zero-shot semantic segmentation framework that combines vision-language image retrieval with cross-image co-segmentation to build open-vocabulary segmenters from unlabeled data without requiring any manual pixel annotations.

Semantic segmentation has a broad range of applications, but its real-world impact has been significantly limited by the prohibitive annotation costs necessary to enable deployment. Segmentation methods that forgo supervision can side-step these costs, but exhibit the inconvenient requirement to provide labelled examples from the target distribution to assign concept names to predictions. An alternative line of work in language-image pre-training has recently demonstrated the potential to produce models that can both assign names across large vocabularies of concepts and enable zero-shot transfer for classification, but do not demonstrate commensurate segmentation abilities. We leverage the retrieval abilities of one such language-image pre-trained model, CLIP, to dynamically curate training sets from unlabelled images for arbitrary collections of concept names, and leverage the robust correspondences offered by modern image representations to co-segment entities among the resulting collections. The synthetic segment collections are then employed to construct a segmentation model (without requiring pixel labels) whose knowledge of concepts is inherited from the scalable pre-training process of CLIP. We demonstrate that our approach, termed Retrieve and Co-segment (ReCo) performs favourably to conventional unsupervised segmentation approaches while inheriting the convenience of nameable predictions and zero-shot transfer. We also demonstrate ReCo’s ability to generate specialist segmenters for extremely rare objects.

Added

2026-09-26

Learning Transferable Features with Deep Adaptation Networks

Learning Transferable Features with Deep Adaptation Networks

Mingsheng Long, Yue Cao, Jianmin Wang, Michael I. Jordan

OrganizationsTsinghua UniversityUniversity of California Berkeley

Why you should read this

Introduces Deep Adaptation Networks, an architecture that aligns task-specific feature distributions across domains using multi-kernel mean embedding matching to reduce dataset bias with linear computational scaling.

Recent studies reveal that a deep neural network can learn transferable features which generalize well to novel tasks for domain adaptation. However, as deep features eventually transition from general to specific along the network, the feature transferability drops significantly in higher layers with increasing domain discrepancy. Hence, it is important to formally reduce the dataset bias and enhance the transferability in task-specific layers. In this paper, we propose a new Deep Adaptation Network (DAN) architecture, which generalizes deep convolutional neural network to the domain adaptation scenario. In DAN, hidden representations of all task-specific layers are embedded in a reproducing kernel Hilbert space where the mean embeddings of different domain distributions can be explicitly matched. The domain discrepancy is further reduced using an optimal multi-kernel selection method for mean embedding matching. DAN can learn transferable features with statistical guarantees, and can scale linearly by unbiased estimate of kernel embedding. Extensive empirical evidence shows that the proposed architecture yields state-of-the-art image classification error rates on standard domain adaptation benchmarks.

Added

2026-09-09

Domain-Adversarial Training of Neural Networks

Domain-Adversarial Training of Neural Networks

Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, Victor Lempitsky

OrganizationsSkolkovo Institute of Science and TechnologyUniversité de SherbrookeUniversité Laval

Why you should read this

Introduces a novel and easily implementable domain-adversarial training method that enables neural networks to achieve state-of-the-art domain adaptation performance without requiring labeled target domain data.

We introduce a new representation learning approach for domain adaptation, in which data at training and test time come from similar but different distributions. Our approach is directly inspired by the theory on domain adaptation suggesting that, for effective domain transfer to be achieved, predictions must be made based on features that cannot discriminate between the training (source) and test (target) domains. The approach implements this idea in the context of neural network architectures that are trained on labeled data from the source domain and unlabeled data from the target domain (no labeled target-domain data is necessary). As the training progresses, the approach promotes the emergence of features that are (i) discriminative for the main learning task on the source domain and (ii) indiscriminate with respect to the shift between the domains. We show that this adaptation behaviour can be achieved in almost any feed-forward model by augmenting it with few standard layers and a new gradient reversal layer. The resulting augmented architecture can be trained using standard backpropagation and stochastic gradient descent, and can thus be implemented with little effort using any of the deep learning packages. We demonstrate the success of our approach for two distinct classification problems (document sentiment analysis and image classification), where state-of-the-art domain adaptation performance on standard benchmarks is achieved. We also validate the approach for descriptor learning task in the context of person re-identification application.

Added

2025-11-08

Creative Commons License