keyword
negative sampling strategies
Negative sampling strategies are methods used in machine learning, recommender systems, and representation learning to select unobserved, non-relevant, or dissimilar items to serve as negative training examples alongside positive target instances. Because computing loss functions over an entire vocabulary, catalog, or dataset is often computationally intractable, these strategies approximate the overall data distribution by drawing a representative subset of negative samples. Approaches range from simple uniform random selection and in-batch sampling to frequency-based weighting, dynamic synthetic generation, and hard negative mining that identifies challenging contrastive pairs. By balancing training efficiency against the risk of sampling bias and false negatives, these strategies ensure that models receive informative gradient signals to accurately separate true positive associations from irrelevant or non-preferred data.
2 items

Plug-In Diffusion Model for Sequential Recommendation
Haokai Ma, Ruobing Xie, Lei Meng, Xin Chen, Xu Zhang, Leyu Lin, Zhanhui Kang
Why you should read this
Proposes a model-agnostic plug-in framework that leverages time-interval diffusion models to generate preference distributions across all items, mitigating data sparsity and noisy interactions in sequential recommenders through behavior reweighting, positive augmentation, and noise-free negative sampling.
Pioneering efforts have verified the effectiveness of the diffusion models in exploring the informative uncertainty for recommendation. Considering the difference between recommendation and image synthesis tasks, existing methods have undertaken tailored refinements to the diffusion and reverse process. However, these approaches typically use the highest-score item in corpus for user interest prediction, leading to the ignorance of the user's generalized preference contained within other items, thereby remaining constrained by the data sparsity issue. To address this issue, this paper presents a novel Plug-In Diffusion Model for Recommendation (PDRec) framework, which employs the diffusion model as a flexible plugin to jointly take full advantage of the diffusion-generating user preferences on all items. Specifically, PDRec first infers the users' dynamic preferences on all items via a time-interval diffusion model and proposes a Historical Behavior Reweighting (HBR) mechanism to identify the high-quality behaviors and suppress noisy behaviors. In addition to the observed items, PDRec proposes a Diffusion-based Positive Augmentation (DPA) strategy to leverage the top-ranked unobserved items as the potential positive samples, bringing in informative and diverse soft signals to alleviate data sparsity. To alleviate the false negative sampling issue, PDRec employs Noise-free Negative Sampling (NNS) to select stable negative samples for ensuring effective model optimization. Extensive experiments and analyses on four datasets have verified the superiority of the proposed PDRec over the state-of-the-art baselines and showcased the universality of PDRec as a flexible plugin for commonly-used sequential encoders in different recommendation scenarios. The code is available in https://github.com/hulkima/PDRec.
Added
2026-09-26

Debiased Contrastive Learning of Unsupervised Sentence Representations
Kun Zhou, Beichen Zhang, Wayne Xin Zhao, Ji-Rong Wen
Why you should read this
Proposes a debiased contrastive learning framework that improves unsupervised sentence embeddings by downweighting false negatives and generating optimized noise-based negative samples to overcome representation anisotropy.
Recently, contrastive learning has been shown to be effective in improving pre-trained language models (PLM) to derive high-quality sentence representations. It aims to pull close positive examples to enhance the alignment while push apart irrelevant negatives for the uniformity of the whole representation space. However, previous works mostly adopt in-batch negatives or sample from training data at random. Such a way may cause the sampling bias that improper negatives (e.g., false negatives and anisotropy representations) are used to learn sentence representations, which will hurt the uniformity of the representation space. To address it, we present a new framework DCLR (Debiased Contrastive Learning of unsupervised sentence Representations) to alleviate the influence of these improper negatives. In DCLR, we design an instance weighting method to punish false negatives and generate noise-based negatives to guarantee the uniformity of the representation space. Experiments on seven semantic textual similarity tasks show that our approach is more effective than competitive baselines. Our code and data are publicly available at the link: https://github.com/RUCAIBox/DCLR.
Added
2026-09-26
