Performance of recommender algorithms on top-n recommendation tasks

P. CremonesiY. KorenRoberto Turrin

article2010RecSys1,551 citations

Demonstrates that minimizing rating prediction error like RMSE fails to optimize top-N recommendation quality, exposing how popularity bias skews standard evaluations and offering simpler collaborative filtering variants that achieve superior ranking accuracy.

Listen

Commercial recommender systems typically present users with a short list of top recommended items rather than numerical rating predictions. Despite this real-world application, the industry and research community widely evaluate and optimize algorithms using rating error metrics, such as root mean squared error. This divergence creates a significant operational challenge: optimizing models to predict exact numerical ratings may not deliver the most relevant recommendations to users, risking misallocated development resources and poor user engagement.

The main objective of the article is to systematically evaluate how collaborative filtering algorithms perform on the top-N recommendation task using accuracy metrics—specifically precision and recall—and to demonstrate whether improvements in rating error metrics translate into better top-ranked recommendations.

To evaluate performance, the analysis assessed multiple personalized and non-personalized recommendation algorithms on two benchmark datasets: the MovieLens dataset (one million ratings across 6,040 users and 3,883 movies) and the Netflix dataset (100 million ratings across 480,189 users and 17,770 movies). The testing protocol measured each model's ability to rank a user's known five-star rating above a random sample of 1,000 unrated items. To account for popularity skew, the evaluation tested models on the full item catalog as well as on a segmented long-tail subset representing 94% to 98% of less-popular items.

The investigation produced four central findings. First, rating error metrics do not reliably reflect recommendation accuracy, and reducing error does not guarantee better ranking performance. Second, in full-catalog evaluations, simple non-personalized popularity baselines matched or exceeded the performance of sophisticated, personalized models optimized for rating error. Third, a streamlined matrix factorization variant called PureSVD consistently delivered the highest precision and recall across both datasets; on the MovieLens dataset, for example, PureSVD achieved a top-10 recall of approximately 52%, outperforming the sophisticated SVD++ model (around 43%) and Asymmetric-SVD (around 28%). Finally, the evaluation showed that including the top 2% to 6% most popular items heavily distorts accuracy measurements, masking the true personalized performance of models on niche catalog items.

These findings indicate that businesses building recommender systems may face substantial performance risks and inflated compute costs by focusing solely on rating prediction metrics. High-complexity models engineered to minimize rating error do not necessarily enhance user-facing top-item lists. Furthermore, measuring accuracy across an entire catalog without isolating popular items creates a misleading signal that favors trivial baseline recommendations rather than novel discovery.

Organizations should re-evaluate their recommendation objectives by designing systems directly for top-N ranking accuracy rather than rating error. Practitioners should consider adopting PureSVD as an efficient, highly scalable baseline, leveraging off-the-shelf sparse matrix decomposition libraries without extensive parameter tuning. When tuning PureSVD, teams should increase the number of latent factors (such as moving from 50 to 150 or 300 factors) to significantly enhance accuracy on long-tail items. To prevent evaluation bias, testing sets should explicitly separate the most popular items from the long tail.

These conclusions are supported by robust, consistent empirical patterns observed across two large-scale movie rating datasets. However, confidence should be tempered by certain boundary conditions: the study evaluated offline movie consumption benchmarks, and accuracy metrics conservatively assumed that randomly sampled unrated items were irrelevant. Additional analysis and online pilot testing are warranted to explore advanced missing-value imputation strategies, confidence-weighted matrix factorization, and generalizability to other commercial domains.

Cover for Performance of recommender algorithms on top-n recommendation tasks

Abstract

In many commercial systems, the ‘best bet’ recommendations are shown, but the predicted rating values are not. This is usually referred to as a top-N recommendation task, where the goal of the recommender system is to find a few specific items which are supposed to be most appealing to the user. Common methodologies based on error metrics (such as RMSE) are not a natural fit for evaluating the top-N recommendation task. Rather, top-N performance can be directly measured by alternative methodologies based on accuracy metrics (such as precision/recall).

An extensive evaluation of several state-of-the art recommender algorithms suggests that algorithms optimized for minimizing RMSE do not necessarily perform as expected in terms of top-N recommendation task. Results show that improvements in RMSE often do not translate into accuracy improvements. In particular, a naive non-personalized algorithm can outperform some common recommendation approaches and almost match the accuracy of sophisticated algorithms. Another finding is that the very few top popular items can skew the top-N performance. The analysis points out that when evaluating a recommender algorithm on the top-N recommendation task, the test set should be chosen carefully in order to not bias accuracy metrics towards non-personalized solutions. Finally, we offer practitioners new variants of two collaborative filtering algorithms that, regardless of their RMSE, significantly outperform other recommender algorithms in pursuing the top-N recommendation task, with offering additional practical advantages. This comes at surprise given the simplicity of these two methods.

Table of Contents

  • 1. INTRODUCTION
  • 2. TESTING METHODOLOGY
  • 2.1 Popular items vs. long-tail
  • 3. COLLABORATIVE ALGORITHMS
  • 3.1 Non-personalized models
  • 3.2 Neighborhood models
  • 3.2.1 Non-normalized Cosine Neighborhood
  • 3.3 Latent Factor Models
  • 3.3.1 PureSVD
  • 4. RESULTS
  • 4.1 Movielens dataset
  • 4.2 Netflix dataset
  • 5. DISCUSSION - PureSVD
  • 6. CONCLUSIONS
  • 7. REFERENCES

Knowls

  1. Knowl 1 — Top-N Recommendation Evaluation Protocol via Unrated Item Sampling

    algorithm

    To evaluate top-NN recommendation accuracy directly on ranking quality rather than rating error metrics, a testing protocol evaluates whether a known 5-star test item can be ranked ahead of randomly chosen unrated items.

    The dataset is partitioned into a training set MM and a test set TT, where TT contains only 5-star ratings from a held-out set (e.g., a validation probe set). For each test rating (u,i)eT(u, i) e T, 1000 items unrated by user uu are randomly sampled under the assumption that most are irrelevant. The model predicts scores for the target item ii and the 1000 unrated items, forms a descending ranking, and determines if item ii falls within the top-NN positions.

    Input: Training rating matrix MM, test set of 5-star ratings TT, recommendation list cutoff NN
    Output: recall(N)\text{recall}(N), precision(N)\text{precision}(N)
    hits←0\text{hits} \leftarrow 0
    Train recommendation model on MM
    for each (u,i)∈T(u, i) \in T do
        Randomly select a set SuS_u of 1000 items not rated by user uu in MM or TT
        Predict recommendation scores r^uj\hat{r}_{uj} for all j∈{i}∪Suj \in \{i\} \cup S_u
        Rank the 1001 candidate items in descending order of r^uj\hat{r}_{uj}
        p←rank of item i in the ordered list (1-indexed)p \leftarrow \text{rank of item } i \text{ in the ordered list (1-indexed)}
        if p≤Np \le N then
            hits←hits+1\text{hits} \leftarrow \text{hits} + 1
    recall(N)←hits∣T∣\text{recall}(N) \leftarrow \frac{\text{hits}}{|T|}
    precision(N)←hitsN⋅∣T∣=recall(N)N\text{precision}(N) \leftarrow \frac{\text{hits}}{N \cdot |T|} = \frac{\text{recall}(N)}{N}
    return recall(N)\text{recall}(N), precision(N)\text{precision}(N)

    Because some of the 1000 randomly selected items might in reality be relevant to user uu, this evaluation methodology provides a conservative lower bound on true precision and recall.

  2. Knowl 2 — PureSVD Matrix Factorization Model for Top-N Recommendation

    model/method

    PureSVD adapts truncated Singular Value Decomposition (SVD) for top-NN recommendation tasks by treating all unobserved ratings in the user-item rating matrix R∈Rn×mR \in \mathbb{R}^{n \times m} as explicit zeros (or an imputed constant) rather than missing values. The full rating matrix is approximated as:

    R^=UΣQT\hat{R} = U \Sigma Q^T

    where U∈Rn×fU \in \mathbb{R}^{n \times f} and Q∈Rm×fQ \in \mathbb{R}^{m \times f} are matrices with orthonormal columns, and Σ∈Rf×f\Sigma \in \mathbb{R}^{f \times f} is a diagonal matrix containing the ff largest singular values.

    By defining the user factor matrix as P=UΣP = U \Sigma, the orthogonality of QQ implies:

    P=UΣ=RQP = U \Sigma = R Q

    Thus, the latent factor representation of user uu is pu=ruQp_u = r_u Q, where ru∈Rmr_u \in \mathbb{R}^m is the uu-th row vector of RR. The top-NN recommendation score r^ui\hat{r}_{ui} for user uu and item ii is computed as:

    r^ui=ruQqiT\hat{r}_{ui} = r_u Q q_i^T

    where qi∈Rfq_i \in \mathbb{R}^f is the ii-th row of QQ.

    This formulation yields several practical advantages:

    1. Factorization can be solved directly using standard sparse SVD packages (e.g., Lanczos algorithms) without iterative gradient descent or hyperparameter tuning.
    2. Users are represented purely as combinations of item features (ruQr_u Q), allowing instant fold-in of new users and ratings online without model retraining.
    3. The predicted scores act as association metrics optimized for ranking rather than calibrated star rating values.
  3. Knowl 3 — Non-Normalized Cosine Neighborhood Model (NNCosNgbr)

    model/method

    The Non-Normalized Cosine Neighborhood (NNCosNgbr) algorithm is an item-item collaborative filtering model tailored for top-NN item ranking rather than explicit rating prediction.

    Standard item-item kk-nearest neighbors (kNN) normalizes the predicted rating by the sum of item similarities:

    r^uinorm=bui+∑j∈Dk(u;i)dij(ruj−buj)∑j∈Dk(u;i)dij\hat{r}_{ui}^{\text{norm}} = b_{ui} + \frac{\sum_{j \in D^k(u; i)} d_{ij} (r_{uj} - b_{uj})}{\sum_{j \in D^k(u; i)} d_{ij}}

    where buib_{ui} is the static user-item baseline bias, Dk(u;i)D^k(u; i) is the set of kk items rated by user uu most similar to item ii, and dijd_{ij} is the shrunk item similarity.

    NNCosNgbr removes the normalizing denominator from the prediction rule:

    r^ui=bui+∑j∈Dk(u;i)dij(ruj−buj)\hat{r}_{ui} = b_{ui} + \sum_{j \in D^k(u; i)} d_{ij} (r_{uj} - b_{uj})

    Omitting the denominator rewards candidate items that have a higher number of strong neighbors in the user's profile, providing a confidence-weighted association score suitable for ranking.

    The similarity sijs_{ij} between items ii and jj is computed as the uncentered cosine similarity across all rows of the rating matrix, setting missing ratings to zero:

    cos⁡(i,j)=i⋅j∥i∥2∥j∥2\cos(i, j) = \frac{\mathbf{i} \cdot \mathbf{j}}{\|\mathbf{i}\|_2 \|\mathbf{j}\|_2}

    where i,j∈Rn\mathbf{i}, \mathbf{j} \in \mathbb{R}^n are the rating column vectors for items ii and jj. The shrunk similarity dijd_{ij} is given by:

    dij=nijnij+λ1cos⁡(i,j)d_{ij} = \frac{n_{ij}}{n_{ij} + \lambda_1} \cos(i, j)

    where nijn_{ij} is the number of users who rated both items ii and jj, and λ1\lambda_1 is a shrinkage parameter (set to λ1=100\lambda_1 = 100).

  4. Knowl 4 — Test Set Partitioning into Short-Head and Long-Tail Items

    experimental setup

    Rating distributions in recommendation datasets follow heavy-tailed power-law distributions. For instance, in the Netflix dataset, approximately 33% of all ratings belong to only 1.7% of the most popular items (302 items); in the MovieLens dataset, 33% of ratings belong to 5.5% of items (213 items).

    To prevent top-NN accuracy metrics from being skewed by trivial popularity baselines, the test set TT (consisting of 5-star user-item interactions) is split into two disjoint subsets:

    • TheadT_{\text{head}} (short-head): Contains test ratings for the small fraction of items accounting for the top 33% of all ratings in the dataset.
    • TlongT_{\text{long}} (long-tail): Contains test ratings for the remaining 98% (Netflix) or 94.5% (MovieLens) of less popular items.

    Evaluating recommender models separately on TlongT_{\text{long}} isolates their capability to recommend novel, non-trivial items without distortion from extreme popularity bias.

  5. Knowl 5 — Discrepancy Between RMSE Minimization and Top-N Recommendation Quality

    empirical result

    Experiments on MovieLens (1M ratings) and Netflix (100M ratings) demonstrate that minimizing Root Mean Squared Error (RMSE) does not guarantee or correlate monotonically with superior top-NN recommendation accuracy (precision and recall):

    1. Models with state-of-the-art RMSE performance (e.g., SVD++ with Netflix RMSE 0.8911, Asymmetric SVD with RMSE 0.9000, and Pearson Correlation Neighborhood CorNgbr with RMSE 0.9406) are consistently outperformed on top-NN recall and precision by models not optimized for RMSE, such as PureSVD and NNCosNgbr.
    2. On the full item test set, a non-personalized popularity baseline (TopPop) achieves top-NN recall and precision comparable to or exceeding RMSE-optimized models (matching Asymmetric SVD on MovieLens and outperforming CorNgbr on Netflix).
    3. Standard RMSE evaluation measures prediction error only on items users chose to rate, introducing selection bias and ignoring the vast space of unrated items. In contrast, top-NN evaluation assesses ranking quality across unrated items, aligning better with models (like PureSVD) trained considering all user-item pairs.
  6. Knowl 6 — Top-N Recommendation Performance Across Full and Long-Tail Item Sets

    empirical result

    Top-NN evaluation of collaborative filtering models on MovieLens and Netflix yields distinct behavior when contrasting full item catalogs against long-tail subsets:

    • Full test set: PureSVD achieves the highest recall and precision on both datasets (e.g., recall@10≈0.52\text{recall}@10 \approx 0.52 on MovieLens, ≈0.45\approx 0.45 on Netflix), followed by NNCosNgbr (recall@10≈0.44\text{recall}@10 \approx 0.44 on MovieLens). The non-personalized TopPop algorithm achieves surprisingly high recall (recall@10≈0.29\text{recall}@10 \approx 0.29 on MovieLens), matching Asymmetric SVD and outperforming Pearson correlation neighborhood (CorNgbr).
    • Long-tail test set (TlongT_{\text{long}}): Excluding the top popular items (accounting for 33% of ratings) causes TopPop accuracy to drop to near zero. PureSVD remains the top-performing model (recall@10≈0.40\text{recall}@10 \approx 0.40 on MovieLens). SVD++ achieves the highest accuracy among RMSE-oriented models. On Netflix, the standard neighborhood model CorNgbr exhibits a relative accuracy improvement, becoming one of the strongest performers on long-tail items despite its poor performance on the full catalog.
  7. Knowl 7 — Dimensionality Effects of PureSVD on Long-Tail Recommendation

    empirical result

    The optimal latent factor dimensionality ff for PureSVD depends on whether the evaluation targets the entire item catalog or long-tail items specifically:

    • When evaluated on the full test set (including short-head popular items), PureSVD with a smaller number of latent factors (f=50f = 50) delivers optimal or near-optimal top-NN recall and precision.
    • When evaluated on the long-tail test set TlongT_{\text{long}}, PureSVD accuracy improves with higher dimensionality, reaching peak performance at f=150f = 150 on MovieLens and f=300f = 300 on Netflix.

    This behavior occurs because the dominant singular vectors in truncated SVD capture aggregate variance associated with high-frequency popular items, whereas additional, higher-index singular vectors encode fine-grained features necessary to accurately differentiate and rank niche, long-tail items.

Coverage note — Deliberately omitted prior-work baseline specifications (standard formulas for MovieAvg, SVD++, and Asymmetric SVD) as they are standard collaborative filtering techniques cited by the paper rather than original contributions.

References

  1. 1.C. Anderson. The Long Tail: Why the Future of Business Is Selling Less of More. Hyperion, July 2006.
  2. 2.R. Bambini, P. Cremonesi, and R. Turrin. Recommender Systems Handbook, chapter A Recommender System for an IPTV Service Provider: a Real Large-Scale Production Environment. Springer, 2010.
  3. 3.J. Bennett and S. Lanning. The Netflix Prize. Proceedings of KDD Cup and Workshop, pages 3–6, 2007.
  4. 4.M. W. Berry. Large-scale sparse singular value computations. The International Journal of Supercomputer Applications, 6(1):13–49, Spring 1992.
  5. 5.O. Celma and P. Cano. From hits to niches? or how popular artists can bias music recommendation and discovery. Las Vegas, USA, August 2008.
  6. 6.P. Cremonesi, E. Lentini, M. Matteucci, and R. Turrin. An evaluation methodology for recommender systems. 4th Int. Conf. on Automated Solutions for Cross Media Content and Multi-channel Distribution, pages 224–231, Nov 2008.
  7. 7.M. Deshpande and G. Karypis. Item-based top-n recommendation algorithms. ACM Transactions on Information Systems (TOIS), 22(1):143–177, 2004.
  8. 8.J. Herlocker, J. Konstan, L. Terveen, and J. Riedl. Evaluating collaborative filtering recommender systems. ACM Transactions on Information Systems (TOIS), 22(1):5–53, 2004.
  9. 9.Y. Hu, Y. Koren, and C. Volinsky. Collaborative filtering for implicit feedback datasets. Data Mining, IEEE International Conference on, 0:263–272, 2008.
  10. 10.Y. Koren. Factorization meets the neighborhood: a multifaceted collaborative filtering model. In KDD '08: Proceeding of the 14th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 426–434, New York, NY, USA, 2008. ACM.
  11. 11.Y. Koren. Collaborative filtering with temporal dynamics. In KDD '09: Proceedings of the 15th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 447–456, New York, NY, USA, 2009. ACM.
  12. 12.Y. Koren, R. M. Bell, and C. Volinsky. Matrix factorization techniques for recommender systems. IEEE Computer, 42(8):30–37, 2009.
  13. 13.B. Miller, I. Albert, S. Lam, J. Konstan, and J. Riedl. MovieLens unplugged: experiences with an occasionally connected recommender system. Proceedings of the 8th international conference on Intelligent user interfaces, pages 263–266, 2003.
  14. 14.A. Paterek. Improving regularized singular value decomposition for collaborative filtering. Proceedings of KDD Cup and Workshop, 2007.
  15. 15.B. Sarwar, G. Karypis, J. Konstan, and J. Reidl. Item-based collaborative filtering recommendation algorithms. 10th Int. Conf. on World Wide Web, pages 285–295, 2001.
  16. 16.B. Sarwar, G. Karypis, J. Konstan, and J. Riedl. Application of Dimensionality Reduction in Recommender System-A Case Study. Defense Technical Information Center, 2000.

Citation

MLA
Cremonesi, P., et al. “Performance of Recommender Algorithms on Top-n Recommendation Tasks”. Proceedings of the Fourth ACM Conference on Recommender Systems, 2010, pp. 39–46, https://doi.org/10.1145/1864708.1864721.
APA
Cremonesi, P., Koren, Y., & Turrin, R. (2010). Performance of recommender algorithms on top-n recommendation tasks. Proceedings of the Fourth ACM Conference on Recommender Systems, 39–46. https://doi.org/10.1145/1864708.1864721
Chicago
Cremonesi, P., Y. Koren, and R. Turrin. 2010. “Performance of Recommender Algorithms on Top-n Recommendation Tasks”. Proceedings of the Fourth ACM Conference on Recommender Systems, 39–46. https://doi.org/10.1145/1864708.1864721.
Harvard
Cremonesi, P., Koren, Y. and Turrin, R. (2010) “Performance of recommender algorithms on top-n recommendation tasks”, Proceedings of the fourth ACM conference on Recommender systems. ACM, pp. 39–46. Available at: https://doi.org/10.1145/1864708.1864721.
Vancouver
1. Cremonesi P, Koren Y, Turrin R (2010) Performance of recommender algorithms on top-n recommendation tasks. In: Proceedings of the fourth ACM conference on Recommender systems. ACM, pp 39–46

BibTeX

@inproceedings{Cremonesi_2010, series={RecSys ’10}, title={Performance of recommender algorithms on top-n recommendation tasks}, url={http://dx.doi.org/10.1145/1864708.1864721}, DOI={10.1145/1864708.1864721}, booktitle={Proceedings of the fourth ACM conference on Recommender systems}, publisher={ACM}, author={Cremonesi, Paolo and Koren, Yehuda and Turrin, Roberto}, year={2010}, month=Sept, pages={39–46}, collection={RecSys ’10} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF