Bayesian probabilistic matrix factorization using Markov chain Monte Carlo

R. SalakhutdinovA. Mnih

article2008ICML1,561 citations

Presents a fully Bayesian treatment of probabilistic matrix factorization using Markov chain Monte Carlo methods to automatically control model complexity and significantly improve recommendation accuracy on large-scale collaborative filtering datasets like Netflix.

Listen

Modern recommendation systems rely heavily on collaborative filtering to predict user preferences, but standard factor-based algorithms face severe challenges with large, sparse, and imbalanced datasets. Conventional methods usually rely on point estimation, fitting a single set of parameters by maximizing the posterior distribution. While computationally fast, this approach requires extensive, manual tuning of regularization parameters to avoid overfitting and struggles to accurately predict preferences for users with very few recorded ratings.

The article evaluates a fully Bayesian approach to Probabilistic Matrix Factorization to determine whether integrating over all model parameters and hyperparameters provides automated complexity control and superior recommendation accuracy at scale. The primary objective is to demonstrate that Markov chain Monte Carlo simulation methods can be practically applied to massive datasets despite common assumptions that they are too computationally slow for enterprise-scale systems.

To demonstrate this approach, the authors implemented Gibbs sampling—a statistical simulation technique that cycles through variables to draw representative samples—and tested it on the benchmark Netflix dataset. This dataset comprises over 100 million movie ratings from more than 480,000 users across nearly 17,800 movie titles. The Bayesian framework was benchmarked against standard matrix factorization, logistic matrix factorization, and singular value decomposition across various model dimensions ranging from 30 to 300 features.

The findings show that the Bayesian model consistently outperforms conventional methods, lowering the root mean squared error by roughly 1.7% to 3.4% over standard matrix factorization across different model capacities. Unlike point-estimate models, which quickly overfit as feature dimensions grow, the Bayesian model’s accuracy steadily improves as model capacity expands up to 300 dimensions (yielding a test error of 0.8954). The most pronounced accuracy gains appear among infrequent users with few observed ratings, where conventional models typically fail. Additionally, the empirical posterior distributions of the factors proved to be non-Gaussian, explaining why this simulation approach outperforms variational approximation techniques.

These results demonstrate that organizations do not need to restrict model complexity artificially to prevent overfitting when using a fully Bayesian approach. By accounting for parameter uncertainty, recommendation engines can avoid costly and brittle manual parameter tuning while delivering significantly more dependable predictions. Furthermore, because the Bayesian model outputs a full predictive distribution rather than a single fixed score, organizations can quantify prediction confidence to make safer, risk-informed recommendation decisions.

Organizations seeking to improve recommendation engines should adopt Bayesian matrix factorization, particularly if their platforms suffer from sparse user activity or severe data imbalance. To offset the primary trade-off—computational demand that scales cubically with feature dimensions—engineering teams should exploit the model's structure by parallelizing the sampling process across multi-core systems and initializing chains using standard point estimates to speed up burn-in time.

Confidence in these findings is high due to testing on an industry-scale dataset and validation on hidden benchmark test data. However, practitioners should be aware of key operational limitations: sampling requires noticeable computation time (from roughly 13 minutes per step at 30 dimensions to 220 minutes at 300 dimensions on a single processor), and diagnosing exact mathematical convergence still relies on empirical rules of thumb rather than guaranteed stopping criteria.

Cover for Bayesian probabilistic matrix factorization using Markov chain Monte Carlo

Abstract

Low-rank matrix approximation methods provide one of the simplest and most effective approaches to collaborative filtering. Such models are usually fitted to data by finding a MAP estimate of the model parameters, a procedure that can be performed efficiently even on very large datasets. However, unless the regularization parameters are tuned carefully, this approach is prone to overfitting because it finds a single point estimate of the parameters. In this paper we present a fully Bayesian treatment of the Probabilistic Matrix Factorization (PMF) model in which model capacity is controlled automatically by integrating over all model parameters and hyperparameters. We show that Bayesian PMF models can be efficiently trained using Markov chain Monte Carlo methods by applying them to the Netflix dataset, which consists of over 100 million movie ratings. The resulting models achieve significantly higher prediction accuracy than PMF models trained using MAP estimation.

Table of Contents

  • 1. Introduction
  • 2. Probabilistic Matrix Factorization
  • 3. Bayesian Probabilistic Matrix Factorization
  • 3.1. The Model
  • 3.2. Predictions
  • 3.3. Inference
  • Gibbs sampling for Bayesian PMF
  • 4. Experimental Results
  • 4.1. Description of the dataset
  • 4.2. Training PMF models
  • 4.3. Training Bayesian PMF models
  • 4.4. Results
  • 5. Conclusions
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Bayesian Probabilistic Matrix Factorization Generative Model

    model/method

    Bayesian Probabilistic Matrix Factorization (BPMF) models an N×MN \times M preference matrix of ratings R∈RN×MR \in \mathbb{R}^{N \times M}, where RijR_{ij} represents the rating assigned by user i∈{1,…,N}i \in \{1, \dots, N\} to item j∈{1,…,M}j \in \{1, \dots, M\}, using DD-dimensional latent feature vectors Ui∈RDU_i \in \mathbb{R}^D for user ii and Vj∈RDV_j \in \mathbb{R}^D for item jj.

    The conditional distribution of the observed ratings (likelihood) with observation noise precision α\alpha is given by:

    p(R∣U,V,α)=∏i=1N∏j=1M[N(Rij|UiTVj,α−1)]Iijp(R | U, V, \alpha) = \prod_{i=1}^N \prod_{j=1}^M \left[ \mathcal{N}\left(R_{ij} \middle| U_i^T V_j, \alpha^{-1}\right) \right]^{I_{ij}}

    where Iij∈{0,1}I_{ij} \in \{0, 1\} is an indicator variable equal to 11 if user ii rated movie jj, and 00 otherwise. The user and movie latent factor matrices U∈RD×NU \in \mathbb{R}^{D \times N} and V∈RD×MV \in \mathbb{R}^{D \times M} are placed under multivariate Gaussian priors:

    p(U∣μU,ΛU)=∏i=1NN(Ui|μU,ΛU−1),p(V∣μV,ΛV)=∏j=1MN(Vj|μV,ΛV−1)p(U | \mu_U, \Lambda_U) = \prod_{i=1}^N \mathcal{N}\left(U_i \middle| \mu_U, \Lambda_U^{-1}\right), \quad p(V | \mu_V, \Lambda_V) = \prod_{j=1}^M \mathcal{N}\left(V_j \middle| \mu_V, \Lambda_V^{-1}\right)

    To perform fully Bayesian inference, conjugate Gaussian-Wishart hyperpriors are placed on the user hyperparameters ΘU={μU,ΛU}\Theta_U = \{\mu_U, \Lambda_U\} and movie hyperparameters ΘV={μV,ΛV}\Theta_V = \{\mu_V, \Lambda_V\}:

    p(ΘU∣Θ0)=N(μU|μ0,(β0ΛU)−1)W(ΛU∣W0,ν0)p(\Theta_U | \Theta_0) = \mathcal{N}\left(\mu_U \middle| \mu_0, (\beta_0 \Lambda_U)^{-1}\right) \mathcal{W}(\Lambda_U | W_0, \nu_0)

    p(ΘV∣Θ0)=N(μV|μ0,(β0ΛV)−1)W(ΛV∣W0,ν0)p(\Theta_V | \Theta_0) = \mathcal{N}\left(\mu_V \middle| \mu_0, (\beta_0 \Lambda_V)^{-1}\right) \mathcal{W}(\Lambda_V | W_0, \nu_0)

    where W(Λ∣W0,ν0)=1C∣Λ∣(ν0−D−1)/2exp⁡(−12Tr(W0−1Λ))\mathcal{W}(\Lambda | W_0, \nu_0) = \frac{1}{C} |\Lambda|^{(\nu_0 - D - 1)/2} \exp\left(-\frac{1}{2} \text{Tr}(W_0^{-1} \Lambda)\right) denotes the Wishart distribution with ν0\nu_0 degrees of freedom, D×DD \times D scale matrix W0W_0, and normalizing constant CC. The hyperprior parameter set is Θ0={μ0,ν0,W0,β0}\Theta_0 = \{\mu_0, \nu_0, W_0, \beta_0\}, typically initialized with μ0=0\mu_0 = 0, ν0=D\nu_0 = D, and W0=ID×DW_0 = I_{D \times D} by symmetry.

  2. Knowl 2 — Exact Conditional Distributions for BPMF Gibbs Sampling

    equation

    Due to conjugate prior choices, all full conditional posterior distributions in Bayesian Probabilistic Matrix Factorization have closed-form expressions that can be sampled exactly.

    The conditional distribution of each user latent feature vector Ui∈RDU_i \in \mathbb{R}^D given item features VV, ratings RR, noise precision α\alpha, and user hyperparameters ΘU={μU,ΛU}\Theta_U = \{\mu_U, \Lambda_U\} is multivariate Gaussian:

    p(Ui∣R,V,ΘU,α)=N(Ui|μi∗,[Λi∗]−1)p(U_i | R, V, \Theta_U, \alpha) = \mathcal{N}\left(U_i \middle| \mu_i^*, [\Lambda_i^*]^{-1}\right)

    where

    Λi∗=ΛU+α∑j=1MIijVjVjT\Lambda_i^* = \Lambda_U + \alpha \sum_{j=1}^M I_{ij} V_j V_j^T

    μi∗=[Λi∗]−1(α∑j=1MIijRijVj+ΛUμU)\mu_i^* = [\Lambda_i^*]^{-1} \left( \alpha \sum_{j=1}^M I_{ij} R_{ij} V_j + \Lambda_U \mu_U \right)

    The conditional distribution for each movie feature vector Vj∈RDV_j \in \mathbb{R}^D given UU, RR, α\alpha, and ΘV={μV,ΛV}\Theta_V = \{\mu_V, \Lambda_V\} is symmetrically Gaussian with precision Λj∗=ΛV+α∑i=1NIijUiUiT\Lambda_j^* = \Lambda_V + \alpha \sum_{i=1}^N I_{ij} U_i U_i^T and mean μj∗=[Λj∗]−1(α∑i=1NIijRijUi+ΛVμV)\mu_j^* = [\Lambda_j^*]^{-1} \left( \alpha \sum_{i=1}^N I_{ij} R_{ij} U_i + \Lambda_V \mu_V \right).

    The conditional distribution of the user hyperparameters ΘU={μU,ΛU}\Theta_U = \{\mu_U, \Lambda_U\} conditioned on UU and hyperprior settings Θ0={μ0,ν0,W0,β0}\Theta_0 = \{\mu_0, \nu_0, W_0, \beta_0\} is Gaussian-Wishart:

    p(μU,ΛU∣U,Θ0)=N(μU|μ0∗,(β0∗ΛU)−1)W(ΛU∣W0∗,ν0∗)p(\mu_U, \Lambda_U | U, \Theta_0) = \mathcal{N}\left(\mu_U \middle| \mu_0^*, (\beta_0^* \Lambda_U)^{-1}\right) \mathcal{W}(\Lambda_U | W_0^*, \nu_0^*)

    where

    μ0∗=β0μ0+NUˉβ0+N,β0∗=β0+N,ν0∗=ν0+N\mu_0^* = \frac{\beta_0 \mu_0 + N \bar{U}}{\beta_0 + N}, \quad \beta_0^* = \beta_0 + N, \quad \nu_0^* = \nu_0 + N

    [W0∗]−1=W0−1+NSˉ+β0Nβ0+N(μ0−Uˉ)(μ0−Uˉ)T[W_0^*]^{-1} = W_0^{-1} + N \bar{S} + \frac{\beta_0 N}{\beta_0 + N} (\mu_0 - \bar{U})(\mu_0 - \bar{U})^T

    Uˉ=1N∑i=1NUi,Sˉ=1N∑i=1NUiUiT\bar{U} = \frac{1}{N} \sum_{i=1}^N U_i, \quad \bar{S} = \frac{1}{N} \sum_{i=1}^N U_i U_i^T

    The conditional distribution for movie hyperparameters ΘV={μV,ΛV}\Theta_V = \{\mu_V, \Lambda_V\} is symmetric with MM replacing NN, and Vˉ=1M∑j=1MVj\bar{V} = \frac{1}{M} \sum_{j=1}^M V_j and SˉV=1M∑j=1MVjVjT\bar{S}_V = \frac{1}{M} \sum_{j=1}^M V_j V_j^T replacing Uˉ\bar{U} and Sˉ\bar{S}.

  3. Knowl 3 — Gibbs Sampling for Bayesian Probabilistic Matrix Factorization

    algorithm

    The Gibbs sampling algorithm generates posterior samples by alternately drawing hyperparameter samples and latent feature vectors from their full conditionals. Because user latent vectors are conditionally independent given item factors and user hyperparameters (and vice versa for movie features), feature updates can be computed completely in parallel across users and items.

    Input: Rating matrix R∈RN×MR \in \mathbb{R}^{N \times M}, indicator matrix I∈{0,1}N×MI \in \{0, 1\}^{N \times M}, latent dimension DD, observation precision α\alpha, hyperpriors Θ0={μ0,ν0,W0,β0}\Theta_0 = \{\mu_0, \nu_0, W_0, \beta_0\}, number of iterations TT
    Output: Sequence of parameter samples {(Ut,Vt)}t=1T\{(U^t, V^t)\}_{t=1}^T
    Initialize U1∈RD×NU^1 \in \mathbb{R}^{D \times N} and V1∈RD×MV^1 \in \mathbb{R}^{D \times M} from linear PMF MAP estimates
    for t=1t = 1 to TT do
        Sample user hyperparameters ΘUt={μUt,ΛUt}∼p(ΘU∣Ut,Θ0)\Theta_U^t = \{\mu_U^t, \Lambda_U^t\} \sim p(\Theta_U | U^t, \Theta_0)
        Sample movie hyperparameters ΘVt={μVt,ΛVt}∼p(ΘV∣Vt,Θ0)\Theta_V^t = \{\mu_V^t, \Lambda_V^t\} \sim p(\Theta_V | V^t, \Theta_0)
        for i=1i = 1 to NN in parallel do
            Λi∗←ΛUt+α∑j=1MIijVjt(Vjt)T\Lambda_i^* \leftarrow \Lambda_U^t + \alpha \sum_{j=1}^M I_{ij} V_j^t (V_j^t)^T
            μi∗←(Λi∗)−1(α∑j=1MIijRijVjt+ΛUtμUt)\mu_i^* \leftarrow (\Lambda_i^*)^{-1} (\alpha \sum_{j=1}^M I_{ij} R_{ij} V_j^t + \Lambda_U^t \mu_U^t)
            Sample Uit+1∼N(μi∗,(Λi∗)−1)U_i^{t+1} \sim \mathcal{N}(\mu_i^*, (\Lambda_i^*)^{-1})
        end for
        for j=1j = 1 to MM in parallel do
            Λj∗←ΛVt+α∑i=1NIijUit+1(Uit+1)T\Lambda_j^* \leftarrow \Lambda_V^t + \alpha \sum_{i=1}^N I_{ij} U_i^{t+1} (U_i^{t+1})^T
            μj∗←(Λj∗)−1(α∑i=1NIijRijUit+1+ΛVtμVt)\mu_j^* \leftarrow (\Lambda_j^*)^{-1} (\alpha \sum_{i=1}^N I_{ij} R_{ij} U_i^{t+1} + \Lambda_V^t \mu_V^t)
            Sample Vjt+1∼N(μj∗,(Λj∗)−1)V_j^{t+1} \sim \mathcal{N}(\mu_j^*, (\Lambda_j^*)^{-1})
        end for
    end for
  4. Knowl 4 — Monte Carlo Rating Prediction in BPMF

    model/method

    The predictive distribution for an unobserved rating Rij∗R^*_{ij} given observed ratings RR and prior parameters Θ0\Theta_0 is defined by marginalizing over all latent parameters and hyperparameters:

    p(Rij∗∣R,Θ0)=∬p(Rij∗∣Ui,Vj)p(U,V∣R,ΘU,ΘV)p(ΘU,ΘV∣Θ0) d{U,V} d{ΘU,ΘV}p(R^*_{ij} | R, \Theta_0) = \iint p(R^*_{ij} | U_i, V_j) p(U, V | R, \Theta_U, \Theta_V) p(\Theta_U, \Theta_V | \Theta_0) \, d\{U, V\} \, d\{\Theta_U, \Theta_V\}

    Using KK samples {(U(k),V(k))}k=1K\{(U^{(k)}, V^{(k)})\}_{k=1}^K obtained from the stationary distribution of the Gibbs sampler after discarding burn-in iterations, the predictive distribution is approximated via Monte Carlo integration:

    p(Rij∗∣R,Θ0)≈1K∑k=1Kp(Rij∗∣Ui(k),Vj(k))p(R^*_{ij} | R, \Theta_0) \approx \frac{1}{K} \sum_{k=1}^K p(R^*_{ij} | U_i^{(k)}, V_j^{(k)})

    The point prediction for rating Rij∗R^*_{ij} is the posterior mean computed across the KK generated samples:

    R^ij=E[Rij∗∣R,Θ0]≈1K∑k=1K(Ui(k))TVj(k)\hat{R}_{ij} = \mathbb{E}[R^*_{ij} | R, \Theta_0] \approx \frac{1}{K} \sum_{k=1}^K (U_i^{(k)})^T V_j^{(k)}

  5. Knowl 5 — Predictive Accuracy Comparison across Latent Feature Dimensionality

    data/table

    The following table compares the Root Mean Squared Error (RMSE) on the Netflix validation set (1,408,395 ratings) and test set (2,817,131 user/movie pairs) for linear Probabilistic Matrix Factorization (PMF) trained with MAP estimation versus Bayesian PMF (BPMF) across latent feature dimensions D∈{30,40,60,150,300}D \in \{30, 40, 60, 150, 300\}.

    DD Valid. RMSE PMF Valid. RMSE BPMF % Inc. Test RMSE PMF Test RMSE BPMF % Inc.
    30 0.9154 0.8994 1.74 0.9188 0.9029 1.73
    40 0.9135 0.8968 1.83 0.9170 0.9002 1.83
    60 0.9150 0.8954 2.14 0.9185 0.8989 2.13
    150 0.9178 0.8931 2.69 0.9211 0.8965 2.67
    300 0.9231 0.8920 3.37 0.9265 0.8954 3.36

    For MAP-trained PMF, predictive performance degrades as feature dimensionality grows past D=40D=40 due to severe overfitting, despite regularization tuning. In contrast, BPMF exhibits automatic capacity control: its validation and test RMSE steadily improve as DD increases up to D=300D=300 (representing approximately 150 million parameters), demonstrating that Bayesian integration prevents overfitting even when parameter count substantially exceeds sample constraints.

  6. Knowl 6 — Predictive Performance Comparison on the Netflix Dataset

    empirical result

    Evaluated on the Netflix dataset consisting of 100,480,507 training ratings from 480,189 users across 17,770 movies (where the Netflix baseline system achieves a test RMSE of 0.9514), BPMF with D=30D=30 latent factors achieves a validation RMSE of 0.8994.

    This performance compares to alternative factorization models with D=30D=30 as follows:

    • Unregularized SVD (equivalent to maximum likelihood PMF): reaches a minimum RMSE of approximately 0.9280 and begins heavily overfitting after roughly 10 training epochs.
    • Linear PMF trained via MAP with regularization λU=λV=0.002\lambda_U = \lambda_V = 0.002: achieves an RMSE of 0.9174 (BPMF improves over this by more than 1.7%).
    • Logistic PMF trained via MAP (mapping user-movie dot products through a logistic sigmoid σ(x)=1/(1+exp⁡(−x))\sigma(x) = 1 / (1 + \exp(-x))): achieves an RMSE of 0.9097.

    Integrating out parameters and hyperparameters via MCMC significantly outperforms both point-estimate MAP models and maximum likelihood approximations.

  7. Knowl 7 — Uncertainty Calibration and Superior Accuracy on Infrequent Users

    empirical result

    Analyzing predictive distributions across users grouped by training rating frequency reveals that BPMF naturally quantifies prediction uncertainty and gains its greatest advantage on sparse data:

    1. Prediction Uncertainty: For infrequent users (e.g., a user with 4 ratings), the predictive posterior over ratings has large dispersion/variance across Gibbs samples, reflecting high parameter uncertainty. For frequent users (e.g., a user with 660 ratings), the posterior predictive distribution is sharply peaked, indicating a well-determined latent vector.
    2. Accuracy on Sparse Users: On user subgroups stratified by observed rating count (1–51\text{--}5, 6–106\text{--}10, 11–2011\text{--}20, up to >641>641 ratings), BPMF achieves dramatic RMSE reductions compared to MAP-trained logistic PMF on the lowest-frequency groups (1–51\text{--}5 and 6–106\text{--}10 ratings).
    3. Convergence on Active Users: As the number of observed ratings per user increases, the performance gap between MAP logistic PMF and BPMF narrows, with both exhibiting similar RMSE on users with hundreds of ratings.
  8. Knowl 8 — Computational Complexity and Parallelization of BPMF Gibbs Sampling

    model/method

    The computational cost of each full Gibbs step in BPMF is dominated by solving a D×DD \times D linear system (or inverting the precision matrix Λi∗\Lambda_i^* / Λj∗\Lambda_j^*) for each user and movie, requiring O(D3)O(D^3) arithmetic operations per entity vector.

    Because the full conditional distribution of user feature vectors factorizes as p(U∣R,V,ΘU)=∏i=1Np(Ui∣R,V,ΘU)p(U | R, V, \Theta_U) = \prod_{i=1}^N p(U_i | R, V, \Theta_U) (and analogously for movie feature vectors), updates across users and items are mutually independent within each half-step, allowing trivial parallelization over multiple processor cores.

    On a single core of a Pentium Xeon 3.00 GHz processor, the wall-clock execution time per full Gibbs sweep on the 100M-rating Netflix dataset scales with latent dimension DD as:

    • D=10D = 10: 6.6 minutes per sweep
    • D=30D = 30: 12.9 minutes per sweep
    • D=60D = 60: 31.6 minutes per sweep
    • D=300D = 300: 220.0 minutes per sweep
  9. Knowl 9 — Non-Gaussian Posterior Structure and Limitation of Variational Approximations

    empirical result

    Empirical inspection of two-dimensional projections of latent user and movie factor samples generated by the Gibbs sampler reveals that posterior distributions over individual feature vectors are non-Gaussian and exhibit complex dependencies between user and movie factors.

    This finding explains why MCMC-based Bayesian PMF improves upon MAP-trained models by a much larger margin than variational Bayesian matrix factorization methods. Variational approaches enforce a factorized mean-field assumption q(U,V)=q(U)q(V)q(U, V) = q(U)q(V) where user factors and movie factors are assumed independent and modeled as Gaussian distributions; this overly restrictive assumption fails to capture the true coupled, non-Gaussian posterior geometry.

  10. Knowl 10 — MCMC Convergence Monitoring and Initialization Strategy for BPMF

    model/method

    In applying Gibbs sampling to large-scale matrix factorization, convergence and efficiency are managed through specific initialization and diagnostic procedures:

    1. Initialization: The Markov chain is initialized at the MAP parameter estimates {U1,V1}\{U^1, V^1\} obtained from mini-batch training of a linear PMF model, rather than from a random start. This places the chain in a high-density region immediately, leading to high predictive accuracy from the earliest iterations.
    2. Convergence Diagnostics: Convergence is diagnosed by monitoring the stabilization of the Frobenius norms of the parameter and hyperparameter matrices (∥U∥Fro\|U\|_{\text{Fro}}, ∥V∥Fro\|V\|_{\text{Fro}}, ∥ΛU∥Fro\|\Lambda_U\|_{\text{Fro}}, ∥ΛV∥Fro\|\Lambda_V\|_{\text{Fro}}, ∥μU∥2\|\mu_U\|_2, ∥μV∥2\|\mu_V\|_2). In practice, these norms stabilize within a few hundred Gibbs steps.
    3. Burn-in and Prediction: Discarding the initial samples (e.g., 800 burn-in samples) ensures that predictions averaged across subsequent sweeps accurately reflect the stationary posterior distribution.

Coverage note — None was omitted; all key contributions—including the model formulation, conjugate Gibbs sampling equations, predictive integration, scalability, complexity analysis, uncertainty behavior, non-Gaussian posterior insights, and Netflix experiments—are fully covered.

References

  1. 1.Hinton, G. E., & van Camp, D. (1993). Keeping the neural networks simple by minimizing the description length of the weights. COLT (pp. 5–13).
  2. 2.Hofmann, T. (1999). Probabilistic latent semantic analysis. Proceedings of the 15th Conference on Uncertainty in AI (pp. 289–296). San Fransisco, California: Morgan Kaufmann.
  3. 3.Jordan, M. I., Ghahramani, Z., Jaakkola, T. S., & Saul, L. K. (1999). An introduction to variational methods for graphical models. Machine Learning, 37, 183.
  4. 4.Lim, Y. J., & Teh, Y. W. (2007). Variational Bayesian approach to movie rating prediction. Proceedings of KDD Cup and Workshop.
  5. 5.Marlin, B. (2004). Modeling user rating profiles for collaborative filtering. In S. Thrun, L. Saul and B. Sch¨olkopf (Eds.), Advances in neural information processing systems 16. Cambridge, MA: MIT Press.
  6. 6.Marlin, B., & Zemel, R. S. (2004). The multiple multiplicative factor model for collaborative filtering. Machine Learning, Proceedings of the Twenty-first International Conference (ICML 2004), Banff, Alberta, Canada. ACM.
  7. 7.Neal, R. M. (1993). Probabilistic inference using Markov chain Monte Carlo methods (Technical Report CRG-TR-93-1). Department of Computer Science, University of Toronto.
  8. 8.Nowlan, S. J., & Hinton, G. E. (1992). Simplifying neural networks by soft weight-sharing. Neural Computation, 4, 473–493.
  9. 9.Raiko, T., Ilin, A., & Karhunen, J. (2007). Principal component analysis for large scale problems with lots of missing values. ECML (pp. 691–698).
  10. 10.Rennie, J. D. M., & Srebro, N. (2005). Fast maximum margin matrix factorization for collaborative prediction. Machine Learning, Proceedings of the Twenty-Second International Conference (ICML 2005), Bonn, Germany (pp. 713–719). ACM.
  11. 11.Salakhutdinov, R., & Mnih, A. (2008). Probabilistic matrix factorization. Advances in Neural Information Processing Systems 20. Cambridge, MA: MIT Press.
  12. 12.Srebro, N., & Jaakkola, T. (2003). Weighted low-rank approximations. Machine Learning, Proceedings of the Twentieth International Conference (ICML 2003), Washington, DC, USA (pp. 720–727). AAAI Press.

Citation

MLA
Salakhutdinov, R., and A. Mnih. “Bayesian Probabilistic Matrix Factorization Using Markov Chain Monte Carlo”. Proceedings of the 25th International Conference on Machine Learning - ICML '08, 2008, pp. 880–87, https://doi.org/10.1145/1390156.1390267.
APA
Salakhutdinov, R., & Mnih, A. (2008). Bayesian probabilistic matrix factorization using Markov chain Monte Carlo. Proceedings of the 25th International Conference on Machine Learning - ICML '08, 880–887. https://doi.org/10.1145/1390156.1390267
Chicago
Salakhutdinov, R., and A. Mnih. 2008. “Bayesian Probabilistic Matrix Factorization Using Markov Chain Monte Carlo”. Proceedings of the 25th International Conference on Machine Learning - ICML '08, 880–87. https://doi.org/10.1145/1390156.1390267.
Harvard
Salakhutdinov, R. and Mnih, A. (2008) “Bayesian probabilistic matrix factorization using Markov chain Monte Carlo”, Proceedings of the 25th international conference on Machine learning - ICML '08. ACM Press, pp. 880–887. Available at: https://doi.org/10.1145/1390156.1390267.
Vancouver
1. Salakhutdinov R, Mnih A (2008) Bayesian probabilistic matrix factorization using Markov chain Monte Carlo. In: Proceedings of the 25th international conference on Machine learning - ICML '08. ACM Press, pp 880–887

BibTeX

@inproceedings{Salakhutdinov_2008, series={ICML ’08}, title={Bayesian probabilistic matrix factorization using Markov chain Monte Carlo}, url={http://dx.doi.org/10.1145/1390156.1390267}, DOI={10.1145/1390156.1390267}, booktitle={Proceedings of the 25th international conference on Machine learning - ICML ’08}, publisher={ACM Press}, author={Salakhutdinov, Ruslan and Mnih, Andriy}, year={2008}, pages={880–887}, collection={ICML ’08} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors