An Attentive Inductive Bias for Sequential Recommendation beyond the Self-Attention

Yehjin ShinJeongwhan ChoiHyowon WiNoseong Park

article2024AAAI120 citations

Reveals that self-attention in sequential recommendation behaves as a low-pass filter causing representation oversmoothing, and introduces BSARec, a Fourier transform-based architecture that integrates high-frequency signals to capture abrupt short-term user preferences alongside long-term interests.

Listen

Modern online platforms rely heavily on sequential recommendation systems to predict what item a user will interact with next based on their recent activity. While transformer-based models using self-attention mechanisms have become the industry standard for this task, they face fundamental design limitations. Specifically, standard self-attention inherently behaves like a low-pass filter, smoothing out detailed variations across sequence layers and leading to representation loss known as oversmoothing. Consequently, existing models excel at tracking long-term, stable user preferences but struggle to capture rapid, short-term changes in interest, such as sudden shifts in browsing trends.

The article introduces and evaluates a novel recommendation architecture called Beyond Self-Attention for Sequential Recommendation (BSARec). The objective is to demonstrate that integrating frequency-domain processing directly into sequential transformers resolves oversmoothing and significantly enhances recommendation accuracy.

The authors designed an architecture that combines standard self-attention with a structured frequency filter using the discrete Fourier transform. This filter injects an explicit structural assumption—that successive item interactions are directly linked—while dynamically rebalancing low-frequency signals (representing persistent interests) and high-frequency signals (representing abrupt, short-term trends). The approach was evaluated across six standard benchmark datasets spanning e-commerce, user reviews, music, and movies, and compared against seven established baseline models without using negative-sampling shortcuts.

The findings confirm that BSARec consistently outperforms all baseline models across all datasets and evaluation metrics. Most notably, the model achieved a 27.49% improvement in top-10 hit rate on the LastFM music dataset over the strongest baseline. It also outperformed complex frameworks that rely on computationally intensive contrastive learning techniques. Furthermore, computational analysis showed that BSARec adds negligible parameter overhead (less than 0.1% increase) and trains significantly faster than contrastive learning methods, requiring only a minor training time increase of about 7% relative to standard self-attentive models.

These results indicate that sequential recommendation engines do not require overly complex training objectives or massive parameter scaling to achieve high performance. Instead, resolving the low-pass filtering flaw of self-attention enables systems to detect sudden user interest shifts without sacrificing long-term profile modeling. For organizations operating digital platforms, adopting this approach can improve recommendation relevance and user engagement with minimal compute and latency overhead.

Engineering and product teams should consider piloting frequency-rescaled attention layers in existing transformer-based recommendation pipelines as a lightweight upgrade over pure self-attention. Because optimal frequency-weighting parameters vary across domains, teams should calibrate the balance between self-attention and inductive bias based on dataset characteristics. While the findings provide high confidence across standard public benchmarks, real-world deployment should be preceded by online A/B testing to validate performance under live traffic dynamics and latency constraints.

arXiv: 2312.10325
Cover for An Attentive Inductive Bias for Sequential Recommendation beyond the Self-Attention

Abstract

Sequential recommendation (SR) models based on Transformers have achieved remarkable successes. The self-attention mechanism of Transformers for computer vision and natural language processing suffers from the oversmoothing problem, i.e., hidden representations becoming similar to tokens. In the SR domain, we, for the first time, show that the same problem occurs. We present pioneering investigations that reveal the low-pass filtering nature of self-attention in the SR, which causes oversmoothing. To this end, we propose a novel method called Beyond Self-Attention for Sequential Recommendation (BSARec), which leverages the Fourier transform to i) inject an inductive bias by considering fine-grained sequential patterns and ii) integrate low and high-frequency information to mitigate oversmoothing. Our discovery shows significant advancements in the SR domain and is expected to bridge the gap for existing Transformer-based SR models. We test our proposed approach through extensive experiments on 6 benchmark datasets. The experimental results demonstrate that our model outperforms 7 baseline methods in terms of recommendation performance. Our code is available at https://github.com/yehjin-shin/BSARec.

Table of Contents

  • Introduction
  • Preliminaries
  • Problem Formulation
  • Self-Attention for Sequential Recommendation
  • Discrete vs. Graph Fourier Transform
  • Motivation
  • Proposed Method
  • Embedding Layer
  • Beyond Self-Attention Encoder
  • Prediction Layer and Training
  • Relation to Previous Models
  • Experiments
  • Experimental Setup
  • Experimental Results
  • Ablation, Sensitivity, and Additional Studies
  • Model Complexity and Runtime Analyses
  • Related Work
  • Sequential Recommendation
  • Oversmoothing and Transformers
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Self-attention is provably a low-pass filter in sequential recommendation

    theoretical result

    For a self-attention matrix A=softmax⁡(QKT/d)A=\operatorname{softmax}(QK^{\mathsf T}/\sqrt d), where Q,K∈RN×dQ,K\in\mathbb{R}^{N\times d} are the query and key matrices, NN is the sequence length, and dd is the key dimension, the paper proves that repeated self-attention removes high-frequency sequence information. Let x∈RNx\in\mathbb{R}^{N}, and let LFC⁡[⋅]\operatorname{LFC}[\cdot] and HFC⁡[⋅]\operatorname{HFC}[\cdot] denote the low- and high-frequency components under the discrete Fourier transform. Then the paper's theorem states

    lim⁡t→∞∥HFC⁡[Atx]∥2∥LFC⁡[Atx]∥2=0.\lim_{t\to\infty}\frac{\left\|\operatorname{HFC}[A^t x]\right\|_2}{\left\|\operatorname{LFC}[A^t x]\right\|_2}=0.

    Thus, regardless of the particular query and key values, repeatedly applying the self-attention matrix asymptotically produces a low-frequency, oversmoothed representation. The associated empirical analysis shows increasing cosine similarity between sequence representations and rapid singular-value decay as Transformer depth increases, indicating loss of feature diversity and rank.

  2. Knowl 2 — Fourier decomposition supplies the attentive inductive bias

    definition

    BSARec treats the NN positions of an item-embedding sequence as a ring graph and applies a discrete Fourier transform along the sequence-position axis. For a scalar sequence x∈RNx\in\mathbb{R}^{N}, let x^=Fx\widehat{x}=Fx be its Fourier spectrum and let cc be the cutoff retaining the cc lowest-frequency components. The low- and high-frequency projections are defined by

    LFC⁡c[x]=F−1(x^1,…,x^c,0,…,0),\operatorname{LFC}_c[x]=F^{-1}(\widehat{x}_1,\ldots,\widehat{x}_c,0,\ldots,0), HFC⁡c[x]=F−1(0,…,0,x^c+1,…,x^N).\operatorname{HFC}_c[x]=F^{-1}(0,\ldots,0,\widehat{x}_{c+1},\ldots,\widehat{x}_{N}).

    The same projections are applied independently to every embedding dimension of a matrix X∈RN×DX\in\mathbb{R}^{N\times D}. Under the ring-graph interpretation, low-frequency components vary smoothly across neighboring positions and represent persistent, long-term interests, whereas high-frequency components vary locally and represent abrupt or short-term preference changes. Using the fixed Fourier structure gives BSARec an inductive bias that does not have to be learned from the training data; the frequency rescaler then learns how strongly to use each component.

  3. Knowl 3 — BSARec mixes Fourier inductive bias with trainable self-attention

    model/method

    BSARec first converts a user's chronologically ordered interaction sequence into a fixed-length embedding matrix. If s=(s1,…,sN)s=(s_1,\ldots,s_N) is the sequence after truncating the oldest items or padding with zeros, M∈R∣V∣×DM\in\mathbb{R}^{|V|\times D} is the item-embedding matrix, and P∈RN×DP\in\mathbb{R}^{N\times D} is a trainable positional-embedding matrix, the input is

    X0=Dropout⁡(LayerNorm⁡(E+P)),Ei=Msi.X^0=\operatorname{Dropout}\bigl(\operatorname{LayerNorm}(E+P)\bigr),\qquad E_i=M_{s_i}.

    At layer ℓ\ell, the beyond-self-attention operation combines a trainable self-attention matrix AℓA^\ell with a Fourier-based attentive inductive-bias operator AIBℓA^\ell_{\mathrm{IB}}:

    Sℓ=A~ℓXℓ=αAIBℓXℓ+(1−α)AℓXℓ,S^\ell=\widetilde{A}^{\ell}X^\ell =\alpha A^\ell_{\mathrm{IB}}X^\ell+(1-\alpha)A^\ell X^\ell,

    where α\alpha controls the trade-off between the fixed Fourier bias and self-attention, and Aℓ=softmax⁡(Qℓ(Kℓ)T/d)A^\ell=\operatorname{softmax}(Q^\ell(K^\ell)^{\mathsf T}/\sqrt d). The Fourier operator is

    AIBℓXℓ=LFC⁡c[Xℓ]+β HFC⁡c[Xℓ],A^\ell_{\mathrm{IB}}X^\ell=\operatorname{LFC}_c[X^\ell]+\beta\,\operatorname{HFC}_c[X^\ell],

    where β\beta is learned either as a scalar or as a DD-dimensional vector, allowing the model to rescale high-frequency information. BSARec uses multi-head versions of this operation, concatenates the head outputs, applies an output projection, and then applies a pointwise feed-forward network, dropout, a residual connection, and layer normalization:

    X~ℓ=GELU⁡(X^ℓW1ℓ+b1ℓ)W2ℓ+b2ℓ,\widetilde{X}^{\ell}=\operatorname{GELU}(\widehat{X}^{\ell}W_1^\ell+b_1^\ell)W_2^\ell+b_2^\ell, Xℓ+1=LayerNorm⁡(Xℓ+X^ℓ+Dropout⁡(X~ℓ)).X^{\ell+1}=\operatorname{LayerNorm}\bigl(X^\ell+\widehat{X}^{\ell}+\operatorname{Dropout}(\widetilde{X}^{\ell})\bigr).

    Here DD is the embedding dimension, W1ℓ,W2ℓ∈RD×DW_1^\ell,W_2^\ell\in\mathbb{R}^{D\times D} and b1ℓ,b2ℓ∈RDb_1^\ell,b_2^\ell\in\mathbb{R}^{D} are learned feed-forward parameters, and X^ℓ\widehat{X}^{\ell} is the multi-head output. The high-frequency term is the mechanism intended to counteract the low-pass behavior of self-attention while the self-attention term preserves the ability to learn non-obvious item dependencies.

  4. Knowl 4 — Next-item prediction uses the final BSARec representation and full-item cross-entropy

    equation

    Let XtuLX^L_{t_u} be the output at the final nonpadding position tut_u of the LL-layer BSARec encoder for user uu, and let ei∈RDe_i\in\mathbb{R}^{D} be the embedding of candidate item ii from the item matrix MM. BSARec assigns item ii the score

    y^i=eiTXtuL.\widehat{y}_i=e_i^{\mathsf T}X^L_{t_u}.

    If g∈Vg\in V is the ground-truth next item and VV is the complete item set, training minimizes the categorical cross-entropy loss

    L=−log⁡exp⁡(y^g)∑i∈Vexp⁡(y^i).\mathcal{L}=-\log\frac{\exp(\widehat{y}_g)}{\sum_{i\in V}\exp(\widehat{y}_i)}.

    The prediction score is therefore a dot-product similarity between the user's final sequential representation and every item embedding, and the loss treats next-item recommendation as classification over the entire item set.

  5. Knowl 5 — Benchmark evaluation covers six datasets and seven sequential-recommendation baselines

    experimental setup

    BSARec was evaluated on Amazon Beauty, Amazon Sports, Amazon Toys, Yelp, LastFM, and MovieLens-1M. Reviews and ratings were converted to implicit feedback using the preprocessing protocol adopted by the paper. The seven comparison methods were Caser and GRU4Rec, SASRec, BERT4Rec, FMLPRec, DuoRec, and FEARec, covering CNN/RNN models, ordinary Transformer-based models, and Transformer models with contrastive learning.

    The implementation used PyTorch on an NVIDIA RTX 3090 with 16 GB memory. The main settings were embedding dimension D=64D=64, maximum sequence length N=50N=50, L=2L=2 BSA blocks, batch size 256256, and Adam optimization with learning rate selected from {5×10−4,10−3}\{5\times10^{-4},10^{-3}\}. The Fourier cutoff cc was selected from {1,3,5,7,9}\{1,3,5,7,9\}, α\alpha from {0.1,0.3,0.5,0.7,0.9}\{0.1,0.3,0.5,0.7,0.9\}, and the number of attention heads from {1,2,4}\{1,2,4\}. Recommendation quality was measured with HR@5, HR@10, HR@20, NDCG@5, NDCG@10, and NDCG@20, evaluating rankings over the full item set without negative sampling.

  6. Knowl 6 — BSARec achieves the best recommendation accuracy on every reported dataset and metric

    data/table

    The benchmark comparison evaluates BSARec against Caser, GRU4Rec, SASRec, BERT4Rec, FMLPRec, DuoRec, and FEARec. The values below report BSARec's score and its relative improvement over the strongest baseline for each dataset and metric. HR is Hit Rate and NDCG is Normalized Discounted Cumulative Gain.

    Could not parse LaTeX table

    BSARec is the best method for all 36 dataset-metric combinations. It also exceeds DuoRec and FEARec, despite not using contrastive learning; the largest reported gain is 27.49% in LastFM HR@10.

  7. Knowl 7 — Both trainable self-attention and Fourier inductive bias are necessary

    data/table

    The ablation study compares the complete BSARec layer with a self-attention-only variant, an attentive-inductive-bias-only variant, and a variant that replaces the frequency-specific vector β\beta with a single scalar. HR@20 and NDCG@20 are reported for Amazon Beauty and Amazon Toys.

    Could not parse LaTeX table

    The Fourier-only variant outperforms the self-attention-only variant on both datasets, showing the value of the fixed sequential-frequency structure. However, the complete mixture of AA and AIBA_{\mathrm{IB}} performs best overall, demonstrating that the inductive bias does not replace learned attention. The scalar-β\beta variant is also generally weaker than the default frequency-rescaling design, supporting the use of a frequency-dimension-specific rescaler.

  8. Knowl 8 — The frequency mixture adapts across datasets and captures abrupt preference changes

    empirical result

    The coefficient α\alpha exhibits dataset-dependent behavior rather than a universally optimal balance: Beauty favors larger α\alpha values, whereas ML-1M achieves its best reported accuracy at α=0.3\alpha=0.3. The cutoff cc is likewise dataset-dependent: Beauty performs best at c=5c=5, while ML-1M improves as cc increases over the tested values.

    The learned high-frequency scale β\beta is larger in the first BSA block than in the second across the evaluated datasets, indicating that emphasizing high-frequency signals early is useful. LastFM and Beauty learn especially large β\beta values. In a LastFM case study, a heavy user whose history is dominated by rock music abruptly changes preference toward pop; only BSARec among the compared models recommends the subsequently consumed pop artist. The authors interpret this example as evidence that the high-frequency pathway can preserve sudden short-term changes that low-pass attention tends to suppress.

  9. Knowl 9 — BSARec adds little parameter overhead and avoids the cost of contrastive-learning baselines

    data/table

    The parameter and training-time comparison below was measured on Amazon Beauty and MovieLens-1M. The number of parameters is reported together with runtime per training epoch in seconds.

    Could not parse LaTeX table

    BSARec adds only 384 parameters relative to SASRec on each dataset. It trains faster per epoch than the contrastive-learning methods DuoRec and FEARec. On ML-1M, BSARec is reported as 7.02% slower than SASRec, but the additional runtime is accompanied by the benchmark accuracy improvement.

  10. Knowl 10 — Several existing sequential-recommendation models arise as restricted BSARec designs

    model/method

    BSARec provides a unifying view of several frequency- and attention-based sequential recommenders. Setting α=0\alpha=0 removes the Fourier inductive-bias term and leaves pure self-attention, making BSARec equivalent in architecture to SASRec; BSARec still uses categorical cross-entropy whereas SASRec uses binary cross-entropy. DuoRec has the same α=0\alpha=0 architectural special case when its additional contrastive-learning objective is ignored.

    FMLPRec uses a discrete Fourier filter without self-attention, but its filter is learned directly and is reported to favor low-pass behavior. BSARec instead combines self-attention with a non-learned Fourier structure and learns the high-frequency rescaling through β\beta. FEARec also separates low- and high-frequency information, but learns frequency processing before its Transformer encoder and adds contrastive learning and frequency normalization. BSARec performs better in the reported experiments with a simpler encoder that injects and rescales both frequency ranges inside every BSA block.

Coverage note — The theorem's proof, appendix-only dataset statistics, and supplementary per-dataset sensitivity and ablation values were deliberately omitted because the task excludes derivations and the main contributed conclusions are already represented by the stated theorem, method, main benchmark, ablation, and efficiency results.

References

  1. 1.Chen, Q.; Zhao, H.; Li, W.; Huang, P.; and Ou, W. 2019. Behavior sequence transformer for e-commerce recommendation in alibaba. In Proceedings of the 1st International Workshop on Deep Learning Practice for High-dimensional Sparse Data, 1–4.
  2. 2.Choi, J.; Hong, S.; Park, N.; and Cho, S.-B. 2023a. Blurring-Sharpening Process Models for Collaborative Filtering. In SIGIR.
  3. 3.Choi, J.; Hong, S.; Park, N.; and Cho, S.-B. 2023b. GREAD: Graph Neural Reaction-Diffusion Networks. In ICML.
  4. 4.Choi, J.; Jeon, J.; and Park, N. 2021. LT-OCF: Learnable-Time ODE-based Collaborative Filtering. In CIKM.
  5. 5.Choi, J.; Wi, H.; Kim, J.; Shin, Y.; Lee, K.; Trask, N.; and Park, N. 2023c. Graph Convolutions Enrich the Self-Attention in Transformers! arXiv preprint arXiv:2312.04234.
  6. 6.Choi, J.; Wi, H.; Lee, C.; Cho, S.-B.; Lee, D.; and Park, N. 2023d. RDGCL: Reaction-Diffusion Graph Contrastive Learning for Recommendation. arXiv preprint arXiv:2312.16563.
  7. 7.Dong, Y.; Cordonnier, J.-B.; and Loukas, A. 2021. Attention is not all you need: Pure attention loses rank doubly exponentially with depth. In ICML, 2793–2803. PMLR.
  8. 8.Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. ICLR.
  9. 9.Du, X.; Yuan, H.; Zhao, P.; Qu, J.; Zhuang, F.; Liu, G.; Liu, Y.; and Sheng, V. S. 2023. Frequency Enhanced Hybrid Attention Network for Sequential Recommendation. In SIGIR, 78–88.
  10. 10.Fan, Z.; Liu, Z.; Peng, H.; and Yu, P. S. 2023. Addressing the Rank Degeneration in Sequential Recommendation via Singular Spectrum Smoothing. arXiv preprint arXiv:2306.11986.
  11. 11.Gao, C.; Zheng, Y.; Li, N.; Li, Y.; Qin, Y.; Piao, J.; Quan, Y.; Chang, J.; Jin, D.; He, X.; et al. 2023. A survey of graph neural networks for recommender systems: Challenges, methods, and directions. ACM Transactions on Recommender Systems, 1(1): 1–51.
  12. 12.Gong, C.; Wang, D.; Li, M.; Chandra, V.; and Liu, Q. 2021. Vision transformers with patch diversification. arXiv preprint arXiv:2104.12753.
  13. 13.Guo, X.; Wang, Y.; Du, T.; and Wang, Y. 2023. Contranorm: A contrastive learning perspective on oversmoothing and beyond. In ICLR.
  14. 14.Hansen, C.; Hansen, C.; Maystre, L.; Mehrotra, R.; Brost, B.; Tomasi, F.; and Lalmas, M. 2020. Contextual and sequential user embeddings for large-scale music recommendation. In RecSys, 53–62.
  15. 15.Harper, F. M.; and Konstan, J. A. 2015. The movielens datasets: History and context. Acm transactions on interactive intelligent systems (tiis), 5(4): 1–19.
  16. 16.He, R.; and McAuley, J. 2016. Fusing similarity models with markov chains for sparse sequential recommendation. In ICDM, 191–200. IEEE.
  17. 17.He, X.; Deng, K.; Wang, X.; Li, Y.; Zhang, Y.; and Wang, M. 2020. LightGCN: Simplifying and Powering Graph Convolution Network for Recommendation. In SIGIR.
  18. 18.He, Y.; and Wai, H.-T. 2021. Identifying first-order lowpass graph signals using perron frobenius theorem. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 5285–5289. IEEE.
  19. 19.Hidasi, B.; Karatzoglou, A.; Baltrunas, L.; and Tikk, D. 2016. Session-based recommendations with recurrent neural networks. In ICLR.
  20. 20.Hong, S.; Jo, M.; Kook, S.; Jung, J.; Wi, H.; Park, N.; and Cho, S.-B. 2022. TimeKit: A Time-series Forecasting-based Upgrade Kit for Collaborative Filtering. In 2022 IEEE International Conference on Big Data (Big Data), 565–574. IEEE.
  21. 21.Huang, X.; Qian, S.; Fang, Q.; Sang, J.; and Xu, C. 2018. CSAN: Contextual self-attention network for user sequential recommendation. In ACM MM, 447–455.
  22. 22.Jiang, J.; Zhang, P.; Luo, Y.; Li, C.; Kim, J. B.; Zhang, K.; Wang, S.; Xie, X.; and Kim, S. 2023. AdaMCT: adaptive mixture of CNN-transformer for sequential recommendation. In CIKM.
  23. 23.Jiang, S.; Qian, X.; Mei, T.; and Fu, Y. 2016. Personalized travel sequence recommendation on multi-source big social media. IEEE Transactions on Big Data, 2(1): 43–56.
  24. 24.Kang, W.-C.; and McAuley, J. 2018. Self-attentive sequential recommendation. In ICDM, 197–206. IEEE.
  25. 25.Kong, T.; Kim, T.; Jeon, J.; Choi, J.; Lee, Y.-C.; Park, N.; and Kim, S.-W. 2022. Linear, or Non-Linear, That is the Question! In WSDM, 517–525.
  26. 26.Krichene, W.; and Rendle, S. 2020. On sampled metrics for item recommendation. In Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining, 1748–1757.
  27. 27.Lee, Y.-C.; Kim, S.-W.; and Lee, D. 2018. gOCCF: Graph-theoretic one-class collaborative filtering based on uninteresting items. In AAAI, volume 32.
  28. 28.Li, J.; Wang, Y.; and McAuley, J. 2020. Time interval aware self-attention for sequential recommendation. In WSDM, 322–330.
  29. 29.Li, Q.; Han, Z.; and Wu, X.-M. 2018. Deeper insights into graph convolutional networks for semi-supervised learning. In AAAI.
  30. 30.Lin, G.; Gao, C.; Zheng, Y.; Chang, J.; Niu, Y.; Song, Y.; Gai, K.; Li, Z.; Jin, D.; Li, Y.; et al. 2023. Mixed Attention Network for Cross-domain Sequential Recommendation. arXiv preprint arXiv:2311.08272.
  31. 31.Liu, Q.; Yan, F.; Zhao, X.; Du, Z.; Guo, H.; Tang, R.; and Tian, F. 2023. Diffusion Augmentation for Sequential Recommendation. In CIKM, 1576–1586.
  32. 32.McAuley, J.; Targett, C.; Shi, Q.; and Van Den Hengel, A. 2015. Image-based recommendations on styles and substitutes. In Proceedings of the 38th international ACM SIGIR conference on research and development in information retrieval, 43–52.
  33. 33.Meyer, C. D.; and Stewart, I. 2023. Matrix analysis and applied linear algebra. SIAM.
  34. 34.Qiu, R.; Huang, Z.; Yin, H.; and Wang, Z. 2022. Contrastive learning for representation degeneration problem in sequential recommendation. In WSDM, 813–823.
  35. 35.Rendle, S.; Freudenthaler, C.; and Schmidt-Thieme, L. 2010. Factorizing personalized markov chains for next-basket recommendation. In TheWebConf (former WWW), 811–820.
  36. 36.Rusch, T. K.; Chamberlain, B.; Rowbottom, J.; Mishra, S.; and Bronstein, M. 2022. Graph-Coupled Oscillator Networks. In ICML, volume 162, 18888–18909.
  37. 37.Sandryhaila, A.; and Moura, J. M. 2014. Discrete signal processing on graphs: Frequency analysis. IEEE Transactions on Signal Processing, 62(12): 3042–3054.
  38. 38.Schedl, M.; Zamani, H.; Chen, C.-W.; Deldjoo, Y.; and Elahi, M. 2018. Current challenges and visions in music recommender systems research. International Journal of Multimedia Information Retrieval, 7: 95–116.
  39. 39.Shin, Y.; Choi, J.; Wi, H.; and Park, N. 2023. An Attentive Inductive Bias for Sequential Recommendation Beyond the Self-Attention. arXiv preprint arXiv:2312.10325.
  40. 40.Sun, F.; Liu, J.; Wu, J.; Pei, C.; Lin, X.; Ou, W.; and Jiang, P. 2019. BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer. In CIKM, 1441–1450.
  41. 41.Tang, J.; and Wang, K. 2018. Personalized top-n sequential recommendation via convolutional sequence embedding. In WSDM, 565–573.
  42. 42.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017. Attention is all you need. In NeurIPS.
  43. 43.Wang, P.; Zheng, W.; Chen, T.; and Wang, Z. 2022. Anti-Oversmoothing in Deep Vision Transformers via the Fourier Domain Analysis: From Theory to Practice. In ICLR.
  44. 44.Wu, J.; Cai, R.; and Wang, H. 2020. Dej´ a vu: A contextualized temporal attention mechanism for sequential recommendation. In TheWebConf (former WWW), 2199–2209.
  45. 45.Wu, L.; Li, S.; Hsieh, C.-J.; and Sharpnack, J. 2020. SSE-PT: Sequential recommendation via personalized transformer. In RecSys, 328–337.
  46. 46.Wu, S.; Sun, F.; Zhang, W.; Xie, X.; and Cui, B. 2022. Graph neural networks in recommender systems: a survey. ACM Computing Surveys, 55(5): 1–37.
  47. 47.Ying, R.; He, R.; Chen, K.; Eksombatchai, P.; Hamilton, W. L.; and Leskovec, J. 2018. Graph Convolutional Neural Networks for Web-Scale Recommender Systems. In KDD.
  48. 48.Yue, Z.; Wang, Y.; He, Z.; Zeng, H.; McAuley, J.; and Wang, D. 2023. Linear Recurrent Units for Sequential Recommendation. arXiv preprint arXiv:2310.02367.
  49. 49.Zhang, T.; Zhao, P.; Liu, Y.; Sheng, V. S.; Xu, J.; Wang, D.; Liu, G.; Zhou, X.; et al. 2019. Feature-level Deeper Self-Attention Network for Sequential Recommendation. In IJCAI, 4320–4326.
  50. 50.Zhou, D.; Kang, B.; Jin, X.; Yang, L.; Lian, X.; Jiang, Z.; Hou, Q.; and Feng, J. 2021. Deepvit: Towards deeper vision transformer. arXiv preprint arXiv:2103.11886.
  51. 51.Zhou, K.; Wang, H.; Zhao, W. X.; Zhu, Y.; Wang, S.; Zhang, F.; Wang, Z.; and Wen, J.-R. 2020. S3-rec: Self-supervised learning for sequential recommendation with mutual information maximization. In CIKM, 1893–1902.
  52. 52.Zhou, K.; Yu, H.; Zhao, W. X.; and Wen, J.-R. 2022. Filter-enhanced MLP is all you need for sequential recommendation. In TheWebConf (former WWW), 2388–2399.
  53. 53.Zhou, P.; Ye, Q.; Xie, Y.; Gao, J.; Wang, S.; Kim, J. B.; You, C.; and Kim, S. 2023. Attention Calibration for Transformer-based Sequential Recommendation. In CIKM, 3595–3605.

Citation

MLA
Shin, Y., et al. “An Attentive Inductive Bias for Sequential Recommendation Beyond the Self-Attention”. arXiv, 2023, http://arxiv.org/abs/2312.10325v2.
APA
Shin, Y., Choi, J., Wi, H., & Park, N. (2023). An Attentive Inductive Bias for Sequential Recommendation beyond the Self-Attention. arXiv. http://arxiv.org/abs/2312.10325v2
Chicago
Shin, Y., J. Choi, H. Wi, and N. Park. 2023. “An Attentive Inductive Bias for Sequential Recommendation Beyond the Self-Attention”. arXiv. http://arxiv.org/abs/2312.10325v2.
Harvard
Shin, Y. et al. (2023) “An Attentive Inductive Bias for Sequential Recommendation beyond the Self-Attention”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.10325v2.
Vancouver
1. Shin Y, Choi J, Wi H, Park N (2023) An Attentive Inductive Bias for Sequential Recommendation beyond the Self-Attention. arXiv

BibTeX

@article{shin2023attentive,
  title = {An Attentive Inductive Bias for Sequential Recommendation beyond the Self-Attention},
  author = {Shin, Yehjin and Choi, Jeongwhan and Wi, Hyowon and Park, Noseong},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.10325v2},
  eprint = {2312.10325}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF