Plug-In Diffusion Model for Sequential Recommendation

Haokai MaRuobing XieLei MengXin ChenXu ZhangLeyu LinZhanhui Kang

article2024AAAI81 citations

Proposes a model-agnostic plug-in framework that leverages time-interval diffusion models to generate preference distributions across all items, mitigating data sparsity and noisy interactions in sequential recommenders through behavior reweighting, positive augmentation, and noise-free negative sampling.

Listen

Modern online platforms rely heavily on sequential recommendation systems to anticipate user needs by analyzing past behavior sequences. However, because most users interact with only a tiny fraction of available items, these systems suffer from severe data sparsity. While generative diffusion models—widely recognized for their success in image generation—have recently been applied to recommendations to model uncertainty, existing approaches only use the single highest-scoring generated item. This narrow focus discards valuable preference information across the rest of the catalog and fails to address temporal dynamics.

The article develops and evaluates the Plug-In Diffusion Model for Recommendation (PDRec), a flexible framework designed to improve sequential recommenders by leveraging diffusion-generated user preferences across all catalog items.

The authors designed a time-interval diffusion model that accounts for the timing between user actions and integrates three modular mechanisms into existing recommendation architectures. First, it reweights historical user actions to filter out noise and emphasize critical interactions. Second, it identifies high-scoring unobserved items to serve as positive training signals, expanding user interests and addressing data sparsity. Third, it selects reliable negative samples from low-scoring unobserved items to prevent incorrect negative labeling. The evaluation tested PDRec across four real-world datasets spanning e-commerce, gaming, books, and music, pairing it with three standard sequential recommendation backbones and assessing performance in single-domain and cross-domain settings.

The evaluation yielded three key findings. First, integrating PDRec produced statistically significant accuracy improvements across all baseline models and datasets, showing gains of up to 13.88% in ranking quality metrics such as Normalized Discounted Cumulative Gain. Second, the performance benefits were especially pronounced on sparser datasets, where data limitations are most severe. Third, in cross-domain recommendation tasks, PDRec achieved improvements of up to 38.3% over standalone diffusion models, demonstrating that its modular components effectively curb negative transfer when merging behavioral streams across domains.

These findings indicate that diffusion models are most valuable when utilized as flexible, pre-trained plugins rather than standalone recommenders. By mining preference signals across the entire item catalog to denoise history and augment training data, organizations can significantly enhance recommendation precision without replacing their existing underlying recommendation architectures. Furthermore, using a pre-trained diffusion model for inference rather than end-to-end retraining helps control computational overhead.

Engineering and data teams should consider adopting this plug-in approach to enhance current sequential recommendation pipelines, particularly for product categories with sparse user interaction. For immediate exploration, organizations can pilot the framework in cross-domain or cold-start scenarios using the authors' publicly available codebase, evaluating the trade-off between offline pre-computation costs and online accuracy gains. Future work should focus on developing advanced hard-negative sampling strategies and testing the framework across broader industry operational environments.

Confidence in these findings is supported by rigorous multi-dataset benchmarking and consistent improvements across different model architectures. However, decision-makers should note that optimal performance requires tuning dataset-specific thresholds (such as truncation bounds and negative sampling proportions), and production deployment will depend on managing the computational resources required for diffusion model inference across very large product catalogs.

Cover for Plug-In Diffusion Model for Sequential Recommendation

Abstract

Pioneering efforts have verified the effectiveness of the diffusion models in exploring the informative uncertainty for recommendation. Considering the difference between recommendation and image synthesis tasks, existing methods have undertaken tailored refinements to the diffusion and reverse process. However, these approaches typically use the highest-score item in corpus for user interest prediction, leading to the ignorance of the user's generalized preference contained within other items, thereby remaining constrained by the data sparsity issue. To address this issue, this paper presents a novel Plug-In Diffusion Model for Recommendation (PDRec) framework, which employs the diffusion model as a flexible plugin to jointly take full advantage of the diffusion-generating user preferences on all items. Specifically, PDRec first infers the users' dynamic preferences on all items via a time-interval diffusion model and proposes a Historical Behavior Reweighting (HBR) mechanism to identify the high-quality behaviors and suppress noisy behaviors. In addition to the observed items, PDRec proposes a Diffusion-based Positive Augmentation (DPA) strategy to leverage the top-ranked unobserved items as the potential positive samples, bringing in informative and diverse soft signals to alleviate data sparsity. To alleviate the false negative sampling issue, PDRec employs Noise-free Negative Sampling (NNS) to select stable negative samples for ensuring effective model optimization. Extensive experiments and analyses on four datasets have verified the superiority of the proposed PDRec over the state-of-the-art baselines and showcased the universality of PDRec as a flexible plugin for commonly-used sequential encoders in different recommendation scenarios. The code is available in https://github.com/hulkima/PDRec.

Table of Contents

  • Introduction
  • Related Work
  • Time-Interval Diffusion Model
  • Plug-In Diffusion Model Task Formulation and Overall Framework
  • Historical Behavior Reweighting
  • Diffusion-based Positive Augmentation
  • Noise-free Negative Sampling
  • Optimization Objectives
  • Experiments
  • Experimental Settings
  • Performance Comparison on SR (RQ1)
  • Ablation Study (RQ2)
  • Universality Analysis of PDRec (RQ3)
  • Results of Cross-domain SR (RQ4)
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Plug-In Diffusion Model for Recommendation (PDRec) Framework

    model/method

    The Plug-In Diffusion Model for Recommendation (PDRec) is a model-agnostic and task-agnostic framework for sequential recommendation (SR) that utilizes a pre-trained diffusion model as an external plugin to guide the training of base sequential recommenders (such as GRU4Rec, SASRec, or CL4SRec).

    Rather than solely using the highest-scoring generated item as the final recommendation, PDRec extracts generated preference scores across the entire item corpus I\mathcal{I}. These scores are deployed across three decoupled plug-in modules during sequence encoder training:

    1. Historical Behavior Reweighting (HBR): Denoises observed user interactions by weighting the supervised loss using the diffusion model's assigned scores and ranks for observed items.
    2. Diffusion-based Positive Augmentation (DPA): Selects top-ranked unobserved items predicted by diffusion, verifies them via user-sequence matching, and introduces them as soft positive training targets through self-distillation.
    3. Noise-free Negative Sampling (NNS): Samples safe negative items exclusively from the lowest-scored tail of unobserved items according to the diffusion model, preventing false negative sampling.

    Because the diffusion model is queried offline or prior to training, PDRec avoids heavy inference computation at recommendation test time while transferring global, generative collaborative knowledge to the sequential recommender.

  2. Knowl 2 — Time-Interval Diffusion Recommendation Model (TI-DiffRec)

    model/method

    The Time-Interval Diffusion Recommendation model (TI-DiffRec) generates corpus-wide user preferences by incorporating behavioral time intervals into the diffusion process.

    Given a user uu's interaction sequence Su={i1u,i2u,…,ipu}\mathcal{S}_u = \{i_1^u, i_2^u, \dots, i_p^u\} with corresponding timestamps Tu={t1u,t2u,…,tpu}\mathcal{T}_u = \{t_1^u, t_2^u, \dots, t_p^u\}, the time-interval weight wujw_u^j for interaction ijui_j^u is computed as:

    wuj=wmin⁡+tuj−tu1tup−tu1(wmax⁡−wmin⁡)w_u^j = w_{\min} + \frac{t_u^j - t_u^1}{t_u^p - t_u^1} (w_{\max} - w_{\min})

    where wmin⁡w_{\min} and wmax⁡w_{\max} are predefined lower and upper bounds. The initial state vector x0′=[x1,x2,…,x∣I∣]∈R∣I∣\mathbf{x}_0' = [x_1, x_2, \dots, x_{|\mathcal{I}|}] \in \mathbb{R}^{|\mathcal{I}|} is constructed such that xiju=wujx_{i_j^u} = w_u^j if user uu interacted with item ijui_j^u, and 00 otherwise.

    In the forward diffusion process, noise is added across steps t∈{1,…,T}t \in \{1, \dots, T\} via transition kernel:

    q(xt′∣xt−1′):=N(xt′;1−βtxt−1′,βtI)q(\mathbf{x}_t' \mid \mathbf{x}_{t-1}') := \mathcal{N}(\mathbf{x}_t'; \sqrt{1 - \beta_t}\mathbf{x}_{t-1}', \beta_t \mathbf{I})

    allowing direct sampling via xt′=αˉtx0′+1−αˉtϵ\mathbf{x}_t' = \sqrt{\bar{\alpha}_t}\mathbf{x}_0' + \sqrt{1 - \bar{\alpha}_t}\boldsymbol{\epsilon} with ϵ∼N(0,I)\boldsymbol{\epsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I}), αt=1−βt\alpha_t = 1 - \beta_t, and αˉt=∏s=1tαs\bar{\alpha}_t = \prod_{s=1}^t \alpha_s.

    In the reverse denoising process, the original interaction profile is recovered through learnable reverse transitions:

    pθ(xt−1′∣xt′):=N(xt−1′;μθ(xt′,t),Σθ(xt′,t))p_\theta(\mathbf{x}_{t-1}' \mid \mathbf{x}_t') := \mathcal{N}(\mathbf{x}_{t-1}'; \boldsymbol{\mu}_\theta(\mathbf{x}_t', t), \boldsymbol{\Sigma}_\theta(\mathbf{x}_t', t))

    parameterized by deep neural networks θ\theta. Reverse inference yields user uu's predicted preference vector x^0′∈R∣I∣\hat{\mathbf{x}}_0' \in \mathbb{R}^{|\mathcal{I}|} across all items.

  3. Knowl 3 — Historical Behavior Reweighting (HBR)

    model/method

    Historical Behavior Reweighting (HBR) uses the diffusion-generated preference scores for observed items to denoise historical sequences during sequential recommender training.

    Let ou∈R∣Iu+∣\mathbf{o}_u \in \mathbb{R}^{|\mathcal{I}_u^+|} and ru∈R∣Iu+∣\mathbf{r}_u \in \mathbb{R}^{|\mathcal{I}_u^+|} denote the diffusion-predicted scores and their corresponding ascending rank indices (where rank 1 is highest) for user uu's observed items Iu+\mathcal{I}_u^+, extracted from x^0′\hat{\mathbf{x}}_0'. The unclipped reweighting vector w^u\hat{\mathbf{w}}_u is formulated as:

    w^u=(1−ωr)⋅ωs⋅ou−min⁡(ou)max⁡(ou)−min⁡(ou)+ωr⋅1+max⁡(ru)−rumax⁡(ru)\hat{\mathbf{w}}_u = (1 - \omega_r) \cdot \omega_s \cdot \frac{\mathbf{o}_u - \min(\mathbf{o}_u)}{\max(\mathbf{o}_u) - \min(\mathbf{o}_u)} + \omega_r \cdot \frac{1 + \max(\mathbf{r}_u) - \mathbf{r}_u}{\max(\mathbf{r}_u)}

    where ωr∈[0,1]\omega_r \in [0, 1] balances preference score and ranking priorities, and ωs\omega_s is a sequence normalization factor defined as:

    ωs=len⁡(Su)∑ou−min⁡(ou)max⁡(ou)−min⁡(ou)\omega_s = \frac{\operatorname{len}(\mathcal{S}_u)}{\sum \frac{\mathbf{o}_u - \min(\mathbf{o}_u)}{\max(\mathbf{o}_u) - \min(\mathbf{o}_u)}}

    To prevent extreme values from dominating training gradients, the final reweight vector wu\mathbf{w}_u is truncated and scaled:

    wu=ωf⋅min⁡(max⁡(cw,min⁡(w^u)),max⁡(w^u))\mathbf{w}_u = \omega_f \cdot \min\left(\max\left(c_w, \min(\hat{\mathbf{w}}_u)\right), \max(\hat{\mathbf{w}}_u)\right)

    where cwc_w is the lower truncation threshold and ωf\omega_f is a rescaling weight.

  4. Knowl 4 — Diffusion-Based Positive Augmentation (DPA)

    model/method

    Diffusion-based Positive Augmentation (DPA) transfers unobserved item preferences generated by TI-DiffRec into the sequential recommender as self-distilled soft positive samples.

    Let uu=x^0′∖ou\mathbf{u}_u = \hat{\mathbf{x}}_0' \setminus \mathbf{o}_u denote the diffusion preference scores over the unobserved items Iu−=I∖Iu+\mathcal{I}_u^- = \mathcal{I} \setminus \mathcal{I}_u^+. DPA performs positive candidate selection and distillation via two steps:

    1. Candidate Retrieval and Matching: The top-mm highest-scoring unobserved items from uu\mathbf{u}_u are selected to form candidate set tu={t1,t2,…,tm}\mathbf{t}_u = \{t_1, t_2, \dots, t_m\} with item embeddings Tu=[t1,t2,…,tm]\mathbf{T}_u = [\mathbf{t}_1, \mathbf{t}_2, \dots, \mathbf{t}_m]. Matching scores mu∈Rm\mathbf{m}_u \in \mathbb{R}^m are computed against user uu's sequence representation hu\mathbf{h}_u (produced by the sequential encoder from the sequence's final item):

    mu=[hu⊤t1,hu⊤t2,…,hu⊤tm]\mathbf{m}_u = [\mathbf{h}_u^\top \mathbf{t}_1, \mathbf{h}_u^\top \mathbf{t}_2, \dots, \mathbf{h}_u^\top \mathbf{t}_m]

    1. Augmentation Selection and Distillation: The candidate items are re-ranked by matching score mu\mathbf{m}_u, and the top-nn (n≤mn \le m) items are retained as soft positive augmentations su\mathbf{s}_u. These soft positives are supervised via a distillation loss LD\mathcal{L}_D:

    LD=−∑(u,i)∈R+[yu,ilog⁡y^u,i]\mathcal{L}_D = - \sum_{(u, i) \in \mathcal{R}^+} \left[ y_{u, i} \log \hat{y}_{u, i} \right]

    where R+\mathcal{R}^+ denotes the set of user-augmented pairs (u,i)(u, i) with i∈sui \in \mathbf{s}_u, target label yu,i=1y_{u, i} = 1, and y^u,i=σ(hu⊤vi)\hat{y}_{u, i} = \sigma(\mathbf{h}_u^\top \mathbf{v}_i) is the predicted probability.

  5. Knowl 5 — Noise-Free Negative Sampling (NNS)

    model/method

    Noise-free Negative Sampling (NNS) mitigates false-negative errors during sequential recommender training by sampling negative items exclusively from the low-confidence tail of diffusion-predicted preferences.

    Given unobserved item corpus Iu−=I∖Iu+\mathcal{I}_u^- = \mathcal{I} \setminus \mathcal{I}_u^+, items are sorted in descending order of their diffusion preferences uu\mathbf{u}_u to yield ranked list Ku\mathcal{K}_u of length lu=∣Ku∣l_u = |\mathcal{K}_u|. NNS assigns non-zero sampling probability only to items ranked after cutoff index ⌊ωmlu⌋\lfloor \omega_m l_u \rfloor:

    PNNS(j∣Iu−)={1(1−ωm)lu,j∈Ku[⌊ωmlu⌋:lu]0,otherwiseP^{\text{NNS}}(j \mid \mathcal{I}_u^-) = \begin{cases} \frac{1}{(1 - \omega_m) l_u}, & j \in \mathcal{K}_u[\lfloor \omega_m l_u \rfloor : l_u] \\ 0, & \text{otherwise} \end{cases}

    where ωm∈(0,1)\omega_m \in (0, 1) is the initial proportion parameter determining the safe slice threshold. A higher ωm\omega_m restricts negative candidate selection to a strictly lower-preference tail, avoiding unobserved items that the user may actually prefer.

  6. Knowl 6 — PDRec Joint Optimization Objective

    equation

    The overall training objective L\mathcal{L} of the sequential recommender within the PDRec framework combines a reweighted recommendation loss LR\mathcal{L}_R and a self-distillation loss LD\mathcal{L}_D:

    L=LR+ωdLD\mathcal{L} = \mathcal{L}_R + \omega_d \mathcal{L}_D

    where ωd\omega_d is the weight of the distillation loss. The components are defined as:

    LR=−∑(u,i)∈R[wu,i⋅yu,ilog⁡y^u,i+(1−yu,i)log⁡(1−y^u,i)]\mathcal{L}_R = - \sum_{(u, i) \in \mathcal{R}} \left[ w_{u, i} \cdot y_{u, i} \log \hat{y}_{u, i} + (1 - y_{u, i}) \log (1 - \hat{y}_{u, i}) \right]

    LD=−∑(u,i)∈R+[yu,ilog⁡y^u,i]\mathcal{L}_D = - \sum_{(u, i) \in \mathcal{R}^+} \left[ y_{u, i} \log \hat{y}_{u, i} \right]

    where:

    • R\mathcal{R} denotes the primary training set comprising observed positive interactions (labeled yu,i=1y_{u, i} = 1), random negative samples, and safe negative items sampled via Noise-free Negative Sampling (NNS) (labeled yu,i=0y_{u, i} = 0).
    • R+\mathcal{R}^+ denotes the set of soft positive augmentation items su\mathbf{s}_u selected via Diffusion-based Positive Augmentation (DPA), with target yu,i=1y_{u, i} = 1.
    • wu,iw_{u, i} is the sample weight from the Historical Behavior Reweighting (HBR) vector wu\mathbf{w}_u for observed positives, and 11 for negative samples.
    • y^u,i=hu⊤vi\hat{y}_{u, i} = \mathbf{h}_u^\top \mathbf{v}_i is the predicted interaction probability between user sequence representation hu\mathbf{h}_u and item embedding vi\mathbf{v}_i.
  7. Knowl 7 — Sequential Recommendation Performance Comparison Across Four Datasets

    data/table

    The performance of PDRec plugged into three sequential recommendation backbones (GRU4Rec, SASRec, CL4SRec) alongside diffusion baselines (T-DiffRec, TI-DiffRec) across Amazon Toy, Amazon Game, Douban Book, and Douban Music datasets under 99 negative test items per positive instance:

    Datasets Metrics T-DiffRec TI-DiffRec GRU4Rec +PDRec (Improv.) SASRec +PDRec (Improv.) CL4SRec +PDRec (Improv.)
    Toy N@1 0.1033 0.1058 0.0878 0.0899 (2.39%) 0.1095 0.1247 (13.88%) 0.1125 0.1254 (11.47%)
    N@5 0.1564 0.1618 0.1515 0.1617 (6.73%) 0.1779 0.2023 (13.72%) 0.1802 0.2041 (13.26%)
    N@10 0.1758 0.1823 0.1755 0.1879 (7.07%) 0.2020 0.2286 (13.17%) 0.2046 0.2305 (12.66%)
    HR@5 0.2055 0.2151 0.2128 0.2300 (8.08%) 0.2423 0.2752 (13.58%) 0.2438 0.2776 (13.86%)
    HR@10 0.2657 0.2787 0.2874 0.3112 (8.28%) 0.3169 0.3568 (12.59%) 0.3195 0.3595 (12.52%)
    AUC 0.5911 0.5968 0.5670 0.5909 (4.22%) 0.5771 0.6060 (5.01%) 0.5805 0.6068 (4.53%)
    Game N@1 0.1611 0.1746 0.1667 0.1808 (8.46%) 0.2111 0.2191 (3.79%) 0.2106 0.2180 (3.51%)
    N@5 0.2567 0.2723 0.2818 0.2996 (6.32%) 0.3310 0.3382 (2.18%) 0.3294 0.3368 (2.25%)
    N@10 0.2895 0.3040 0.3199 0.3380 (5.66%) 0.3682 0.3753 (1.93%) 0.3682 0.3750 (1.85%)
    HR@5 0.3451 0.3618 0.3893 0.4091 (5.09%) 0.4409 0.4475 (1.50%) 0.4385 0.4456 (1.62%)
    HR@10 0.4469 0.4600 0.5071 0.5282 (4.16%) 0.5559 0.5626 (1.21%) 0.5584 0.5638 (0.97%)
    AUC 0.7217 0.7234 0.7601 0.7786 (2.43%) 0.7865 0.7908 (0.61%) 0.7857 0.7905 (0.61%)
    Book N@1 0.3194 0.3275 0.3072 0.3359 (9.34%) 0.3594 0.3656 (1.73%) 0.3554 0.3621 (1.89%)
    N@5 0.4398 0.4491 0.4433 0.4757 (7.31%) 0.4948 0.5063 (2.32%) 0.4942 0.5047 (2.12%)
    N@10 0.4671 0.4776 0.4765 0.5091 (6.84%) 0.5272 0.5393 (2.30%) 0.5276 0.5376 (1.90%)
    HR@5 0.5459 0.5557 0.5643 0.6004 (6.40%) 0.6148 0.6306 (2.57%) 0.6166 0.6304 (2.24%)
    HR@10 0.6300 0.6435 0.6667 0.7033 (5.49%) 0.7150 0.7323 (2.42%) 0.7197 0.7317 (1.67%)
    AUC 0.8160 0.8202 0.8541 0.8728 (2.19%) 0.8790 0.8898 (1.23%) 0.8820 0.8895 (0.85%)
    Music N@1 0.3401 0.3494 0.3299 0.3540 (7.31%) 0.3753 0.3826 (1.95%) 0.3689 0.3755 (1.79%)
    N@5 0.4709 0.4773 0.4725 0.5000 (5.82%) 0.5170 0.5283 (2.19%) 0.5096 0.5211 (2.26%)
    N@10 0.4987 0.5049 0.5069 0.5348 (5.50%) 0.5503 0.5620 (2.13%) 0.5435 0.5558 (2.26%)
    HR@5 0.5852 0.5886 0.5987 0.6287 (5.01%) 0.6421 0.6573 (2.37%) 0.6353 0.6504 (2.38%)
    HR@10 0.6706 0.6738 0.7048 0.7361 (4.44%) 0.7447 0.7612 (2.22%) 0.7400 0.7573 (2.34%)
    AUC 0.8329 0.8318 0.8768 0.8908 (1.60%) 0.8962 0.9040 (0.87%) 0.8939 0.9026 (0.97%)

    Key observations:

    1. TI-DiffRec consistently outperforms T-DiffRec across all metrics and datasets, verifying the benefit of time-interval reweighting in diffusion.
    2. PDRec consistently and significantly improves all base sequential models (p<0.05p < 0.05 with paired t-tests). Performance gains are highest on the sparsest dataset (Amazon Toy with density 0.0378%, yielding up to 13.88% relative improvement on N@1 for SASRec).
  8. Knowl 8 — Cross-Domain Sequential Recommendation Performance and Ablations

    data/table

    Cross-domain sequential recommendation (CDSR) performance when evaluating PDRec and its incremental ablation versions on chronological multi-domain interaction sequences between Amazon Game and Toy domains:

    Setting Algorithms N@1 N@5 N@10 N@20 N@50 HR@5 HR@10 HR@20 HR@50 AUC
    Game →\to Toy T-DiffRec (M) 0.0981 0.1520 0.1727 0.1934 0.2375 0.2029 0.2673 0.3494 0.5780 0.5924
    TI-DiffRec (M) 0.1053 0.1598 0.1806 0.2008 0.2407 0.2111 0.2759 0.3562 0.5623 0.5932
    SASRec (M) 0.1267 0.2019 0.2261 0.2490 0.2785 0.2722 0.3472 0.4380 0.5873 0.5951
    +HBR 0.1283 0.2061 0.2311 0.2533 0.2835 0.2785 0.3558 0.4438 0.5972 0.6092
    +HBR+NNS 0.1264 0.2068 0.2323 0.2542 0.2844 0.2815 0.3606 0.4480 0.6013 0.6123
    +HBR+NNS+DPA 0.1302 0.2093 0.2348 0.2574 0.2873 0.2826 0.3616 0.4515 0.6026 0.6106
    Toy →\to Game T-DiffRec (M) 0.1674 0.2643 0.2977 0.3247 0.3597 0.3548 0.4584 0.5655 0.7428 0.7232
    TI-DiffRec (M) 0.1709 0.2757 0.3096 0.3378 0.3721 0.3723 0.4773 0.5887 0.7622 0.7407
    SASRec (M) 0.2273 0.3532 0.3905 0.4190 0.4467 0.4674 0.5826 0.6955 0.8342 0.8007
    +HBR 0.2332 0.3597 0.3963 0.4250 0.4547 0.4741 0.5872 0.7006 0.8501 0.8145
    +HBR+NNS 0.2352 0.3601 0.3975 0.4257 0.4557 0.4733 0.5890 0.7002 0.8517 0.8138
    +HBR+NNS+DPA 0.2363 0.3623 0.3992 0.4275 0.4572 0.4761 0.5904 0.7022 0.8520 0.8153

    Key takeaways:

    1. Direct sequence mixing in CDSR benefits from HBR by dynamically downweighting cross-domain noise and negative transfer.
    2. Successive integration of HBR, NNS, and DPA delivers monotonic improvements on almost all metrics, outperforming the standalone diffusion recommenders by up to 38.3%.
  9. Knowl 9 — Ablation and Replacement Analysis of PDRec Modules

    empirical result

    Systematic ablation on SASRec, GRU4Rec, and CL4SRec confirms the specific utility of each component and the indispensability of the underlying diffusion model:

    1. Effect of HBR: Adding HBR to SASRec consistently improves N@10 and HR@10 across all datasets by suppressing noisy interactions in user history via diffusion-derived weights.
    2. Effect of NNS: Adding NNS to SASRec+HBR further enhances performance by eliminating false-negative samples from training, drawing negative instances strictly from the low-confidence tail.
    3. Effect of DPA: Full PDRec (SASRec+HBR+NNS+DPA) outperforms SASRec+HBR+NNS, showing that top-ranked unobserved items provide diverse and valid positive signals that alleviate data sparsity.
    4. Necessity of Diffusion Generative Prior: Replacing TI-DiffRec in PDRec with a second SASRec model (evaluating PDRec w/o TI-DiffRec) leads to a substantial performance drop across all four datasets (e.g., N@10 on Amazon Toy drops from ~0.23 to ~0.20, and HR@10 drops from ~0.36 to ~0.31). This verifies that the full-corpus visibility and probabilistic generative modeling of diffusion models provide generalization signals that standard sequential encoders cannot replicate.

Coverage note — None was omitted; all contributed models (PDRec, TI-DiffRec, HBR, DPA, NNS), equations, primary benchmark experiments, CDSR experiments, and ablation findings are fully covered.

References

  1. 1.Brempong, E. A.; Kornblith, S.; Chen, T.; Parmar, N.; Minderer, M.; and Norouzi, M. 2022. Denoising pretraining for semantic segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR).
  2. 2.Chen, G.; Zhang, X.; Su, Y.; Lai, Y.; Xiang, J.; Zhang, J.; and Zheng, Y. 2023a. Win-Win: A Privacy-Preserving Federated Framework for Dual-Target Cross-Domain Recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI).
  3. 3.Chen, H.; He, J.; Xu, W.; Feng, T.; Liu, M.; Song, T.; Yao, R.; and Qiao, Y. 2023b. Enhanced Multi-Relationships Integration Graph Convolutional Network for Inferring Substitutable and Complementary Items. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI).
  4. 4.Chen, Y.; Liu, Z.; Li, J.; McAuley, J.; and Xiong, C. 2022. Intent contrastive learning for sequential recommendation. In Proceedings of the ACM Web Conference (WWW)).
  5. 5.Du, H.; Yuan, H.; Huang, Z.; Zhao, P.; and Zhou, X. 2023. Sequential Recommendation with Diffusion Models.
  6. 6.Hidasi, B.; Karatzoglou, A.; Baltrunas, L.; and Tikk, D. 2016. Session-based recommendations with recurrent neural networks. In Proceedings of International Conference on Learning Representations (ICLR).
  7. 7.Ho, J.; Jain, A.; and Abbeel, P. 2020. Denoising diffusion probabilistic models. Proceedings of Advances in neural information processing systems ((NeurIPS).
  8. 8.Ho, J.; Saharia, C.; Chan, W.; Fleet, D. J.; Norouzi, M.; and Salimans, T. 2022. Cascaded diffusion models for high fidelity image generation. The Journal of Machine Learning Research (JMLR).
  9. 9.Kang, W.-C.; and McAuley, J. 2018. Self-attentive sequential recommendation. In Proceedings of International Conference on Data Mining (ICDM).
  10. 10.Kingma, D. P.; and Welling, M. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114.
  11. 11.Li, J.; Ren, P.; Chen, Z.; Ren, Z.; Lian, T.; and Ma, J. 2017. Neural attentive session-based recommendation. In Proceedings of ACM International Conference on Information and Knowledge Management (CIKM).
  12. 12.Li, M.; Zhao, X.; Lyu, C.; Zhao, M.; Wu, R.; and Guo, R. 2022. MLP4Rec: A Pure MLP Architecture for Sequential Recommendations.
  13. 13.Li, Z.; Sun, A.; and Li, C. 2023. DiffuRec: A Diffusion Model for Sequential Recommendation. arXiv preprint arXiv:2304.00686.
  14. 14.Lin, G.; Gao, C.; Li, Y.; Zheng, Y.; Li, Z.; Jin, D.; and Li, Y. 2022. Dual Contrastive Network for Sequential Recommendation with User and Item-Centric Perspectives. arXiv preprint arXiv:2209.08446.
  15. 15.Lopez Alcaraz, J. M.; and Strodthoff, N. 2023. Diffusion-based time series imputation and forecasting with structured atate apace models. Transactions on machine learning research (TMLR).
  16. 16.Ma, H.; Li, X.; Meng, L.; and Meng, X. 2021. Comparative study of adversarial training methods for cold-start recommendation. In Proceedings of ADVM.
  17. 17.Ma, H.; Qi, Z.; Dong, X.; Li, X.; Zheng, Y.; and Meng, X. M. L. 2023a. Cross-Modal Content Inference and Feature Enrichment for Cold-Start Recommendation. Proceedings of IJCNN.
  18. 18.Ma, H.; Xie, R.; Meng, L.; Chen, X.; Zhang, X.; Lin, L.; and Zhou, J. 2023b. Exploring False Hard Negative Sample in Cross-Domain Recommendation. In Proceedings of the ACM Conference on Recommender Systems (RecSys).
  19. 19.Ma, H.; Xie, R.; Meng, L.; Chen, X.; Zhang, X.; Lin, L.; and Zhou, J. 2023c. Triple Sequence Learning for Cross-domain Recommendation. ACM Trans. Inf. Syst. (TOIS).
  20. 20.Meng, L.; Feng, F.; He, X.; Gao, X.; and Chua, T.-S. 2020. Heterogeneous fusion of semantic and collaborative information for visually-aware food recommendation. In Proceedings of MM.
  21. 21.Moon, J.; Jeong, Y.; Chae, D.-K.; Choi, J.; Shim, H.; and Lee, J. 2023. CoMix: Collaborative filtering with mixup for implicit datasets. Information Sciences.
  22. 22.Nichol, A. Q.; and Dhariwal, P. 2021. Improved denoising diffusion probabilistic models. In Proceedings of International Conference on Machine Learning (ICML). PMLR.
  23. 23.Shi, W.; Chen, J.; Feng, F.; Zhang, J.; Wu, J.; Gao, C.; and He, X. 2023. On the Theories Behind Hard Negative Sampling for Recommendation. In Proceedings of the ACM Web Conference (WWW).
  24. 24.Shi, Y.; De Bortoli, V.; Deligiannidis, G.; and Doucet, A. 2022. Conditional simulation using diffusion Schrödinger bridges. In Proceedings of Uncertainty in Artificial Intelligence (UAI).
  25. 25.Sun, F.; Liu, J.; Wu, J.; Pei, C.; Lin, X.; Ou, W.; and Jiang, P. 2019. BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer. In Proceedings of ACM International Conference on Information and Knowledge Management (CIKM).
  26. 26.Tang, J.; and Wang, K. 2018. Personalized top-n sequential recommendation via convolutional sequence embedding. In Proceedings of ACM International Conference on Web Search and Data Mining (WSDM).
  27. 27.Tashiro, Y.; Song, J.; Song, Y.; and Ermon, S. 2021. Csdi: Conditional score-based diffusion models for probabilistic time series imputation. Proceedings of Advances in Neural Information Processing Systems (NeurIPs).
  28. 28.Walker, J.; Zhong, T.; Zhang, F.; Gao, Q.; and Zhou, F. 2022. Recommendation via collaborative diffusion generative model. In Proceedings of International Conference on Knowledge Science, Engineering and Management, 593–605. Springer.
  29. 29.Wang, W.; Xu, Y.; Feng, F.; Lin, X.; He, X.; and Chua, T.-S. 2023. Diffusion Recommender Model. Proceedings of International Conference on Research on Development in Information Retrieval (SIGIR).
  30. 30.Wu, B.; He, X.; Wu, L.; Zhang, X.; and Ye, Y. 2023. Graph-augmented co-attention model for socio-sequential recommendation. IEEE Transactions on Systems, Man, and Cybernetics: Systems (SMC).
  31. 31.Xia, L.; Huang, C.; Xu, Y.; Dai, P.; Zhang, X.; Yang, H.; Pei, J.; and Bo, L. 2021. Knowledge-enhanced hierarchical graph transformer network for multi-behavior recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI).
  32. 32.Xie, X.; Sun, F.; Liu, Z.; Wu, S.; Gao, J.; Zhang, J.; Ding, B.; and Cui, B. 2022. Contrastive learning for sequential recommendation. In Proceedings of IEEE International Conference on Data Engineering (ICDE).
  33. 33.Xu, C.; Zhao, P.; Liu, Y.; Xu, J.; S. Sheng, V. S. S.; Cui, Z.; Zhou, X.; and Xiong, H. 2019. Recurrent convolutional neural network for sequential recommendation. In Proceedings of International World Wide Web Conferences (WWW).
  34. 34.Zhang, M.; Wu, S.; Yu, X.; Liu, Q.; and Wang, L. 2022. Dynamic graph neural networks for sequential recommendation. IEEE Transactions on Knowledge and Data Engineering (TKDE).
  35. 35.Zheng, X.; Su, J.; Liu, W.; and Chen, C. 2022. DDGHM: Dual Dynamic Graph with Hybrid Metric Training for Cross-Domain Sequential Recommendation. In Proceedings of ACM International Conference on Multimedia (ACM MM).

Citation

MLA
Ma, H., et al. “Plug-in Diffusion Model for Sequential Recommendation”. arXiv, 2024, http://arxiv.org/abs/2401.02913v1.
APA
Ma, H., Xie, R., Meng, L., Chen, X., Zhang, X., Lin, L., & Kang, Z. (2024). Plug-in Diffusion Model for Sequential Recommendation. arXiv. http://arxiv.org/abs/2401.02913v1
Chicago
Ma, H., R. Xie, L. Meng, et al. 2024. “Plug-in Diffusion Model for Sequential Recommendation”. arXiv. http://arxiv.org/abs/2401.02913v1.
Harvard
Ma, H. et al. (2024) “Plug-in Diffusion Model for Sequential Recommendation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.02913v1.
Vancouver
1. Ma H, Xie R, Meng L, Chen X, Zhang X, Lin L, Kang Z (2024) Plug-in Diffusion Model for Sequential Recommendation. arXiv

BibTeX

@article{ma2024plug,
  title = {Plug-in Diffusion Model for Sequential Recommendation},
  author = {Ma, Haokai and Xie, Ruobing and Meng, Lei and Chen, Xin and Zhang, Xu and Lin, Leyu and Kang, Zhanhui},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.02913v1},
  eprint = {2401.02913}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF