An Attentive Inductive Bias for Sequential Recommendation beyond the Self-Attention
Yehjin ShinJeongwhan ChoiHyowon WiNoseong Park
Reveals that self-attention in sequential recommendation behaves as a low-pass filter causing representation oversmoothing, and introduces BSARec, a Fourier transform-based architecture that integrates high-frequency signals to capture abrupt short-term user preferences alongside long-term interests.
Modern online platforms rely heavily on sequential recommendation systems to predict what item a user will interact with next based on their recent activity. While transformer-based models using self-attention mechanisms have become the industry standard for this task, they face fundamental design limitations. Specifically, standard self-attention inherently behaves like a low-pass filter, smoothing out detailed variations across sequence layers and leading to representation loss known as oversmoothing. Consequently, existing models excel at tracking long-term, stable user preferences but struggle to capture rapid, short-term changes in interest, such as sudden shifts in browsing trends.
The article introduces and evaluates a novel recommendation architecture called Beyond Self-Attention for Sequential Recommendation (BSARec). The objective is to demonstrate that integrating frequency-domain processing directly into sequential transformers resolves oversmoothing and significantly enhances recommendation accuracy.
The authors designed an architecture that combines standard self-attention with a structured frequency filter using the discrete Fourier transform. This filter injects an explicit structural assumption—that successive item interactions are directly linked—while dynamically rebalancing low-frequency signals (representing persistent interests) and high-frequency signals (representing abrupt, short-term trends). The approach was evaluated across six standard benchmark datasets spanning e-commerce, user reviews, music, and movies, and compared against seven established baseline models without using negative-sampling shortcuts.
The findings confirm that BSARec consistently outperforms all baseline models across all datasets and evaluation metrics. Most notably, the model achieved a 27.49% improvement in top-10 hit rate on the LastFM music dataset over the strongest baseline. It also outperformed complex frameworks that rely on computationally intensive contrastive learning techniques. Furthermore, computational analysis showed that BSARec adds negligible parameter overhead (less than 0.1% increase) and trains significantly faster than contrastive learning methods, requiring only a minor training time increase of about 7% relative to standard self-attentive models.
These results indicate that sequential recommendation engines do not require overly complex training objectives or massive parameter scaling to achieve high performance. Instead, resolving the low-pass filtering flaw of self-attention enables systems to detect sudden user interest shifts without sacrificing long-term profile modeling. For organizations operating digital platforms, adopting this approach can improve recommendation relevance and user engagement with minimal compute and latency overhead.
Engineering and product teams should consider piloting frequency-rescaled attention layers in existing transformer-based recommendation pipelines as a lightweight upgrade over pure self-attention. Because optimal frequency-weighting parameters vary across domains, teams should calibrate the balance between self-attention and inductive bias based on dataset characteristics. While the findings provide high confidence across standard public benchmarks, real-world deployment should be preceded by online A/B testing to validate performance under live traffic dynamics and latency constraints.
- Paper: Self-Attentive Sequential Recommendation, Wang-Cheng Kang et al. (2018). Introduces SASRec, the foundational self-attentive sequential recommendation architecture whose oversmoothing and low-pass filtering limitations are directly analyzed and resolved by BSARec.
- Paper: BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer, Fei Sun et al. (2019). Establishes bidirectional Transformer encoding for sequential recommendation, serving as a primary baseline and benchmark for self-attention modeling in sequential recommendation.
- Paper: Behavior sequence transformer for e-commerce recommendation in Alibaba, Qiwei Chen et al. (2019). Provides early architectural groundwork for integrating Transformer self-attention mechanisms to model behavior sequence dynamics in recommender systems.
- Paper: Personalized Top-N Sequential Recommendation via Convolutional Sequence Embedding, Jiaxi Tang et al. (2018). Demonstrates the necessity of capturing local and skip-behavior inductive biases in sequential recommendation prior to pure attention-based paradigms.
- Paper: Session-based Recommendations with Recurrent Neural Networks, Balázs Hidasi et al. (2016). Pioneers the core formulation of neural sequential recommendation, providing essential context on how item transition modeling evolved before Transformer architectures.
- Paper: An Embarrassingly Simple Graph Heuristic Reveals Shortcut-Solvable Benchmarks for Sequential Recommendation, Haoyu Han et al. (2026). Critically re-examines standard sequential recommendation benchmarks by evaluating whether the sophisticated inductive biases developed by advanced sequence models truly outperform simple transition heuristics.
