Online Learning for Latent Dirichlet Allocation

Matthew D. HoffmanDavid M. BleiFrancis Bach

article2010NeurIPS1,768 citationsTest of Time Award

Presents an online variational Bayes algorithm for Latent Dirichlet Allocation that scales topic modeling to massive streaming document collections by converging to optimal solutions significantly faster than traditional batch methods.

Listen

Modern organizations face massive streams of text data, such as scientific publications, web pages, and customer communications, that cannot be read or categorized by hand. Probabilistic topic modeling, specifically Latent Dirichlet Allocation, is a powerful technique for automatically discovering the latent themes within massive text collections. However, standard methods require analyzing the entire dataset repeatedly in batch passes, making large-scale deployment computationally prohibitive, slow, and poorly suited for real-time data streams.

The article demonstrates an online variational Bayes algorithm that fits topic models to massive and continuous document collections on the fly. The main objective is to establish that this online optimization approach achieves equal or better model quality compared to traditional batch methods while using only a fraction of the time and computational memory.

The authors designed a streaming stochastic optimization approach that updates the global model after analyzing small batches of documents, discarding them immediately after one look. They evaluated the algorithm on static benchmark datasets (352,549 articles from the journal Nature and 100,000 Wikipedia articles) across 288 parameter combinations to determine optimal settings. They also demonstrated real-world feasibility by streaming and fitting a 100-topic model to 3.3 million Wikipedia articles in a single pass.

The evaluation produced several critical findings. First, online learning achieved topic models with quality equal to or better than batch methods while requiring substantially less computation time. Second, in the large-scale test, the online algorithm processed 3.3 million Wikipedia articles at a rate of 60,000 documents per hour and converged to an optimal solution after seeing roughly half the dataset, completing the single pass in under three days—whereas a single iteration of traditional batch methods would have taken days. Third, the method maintained constant memory requirements because documents did not need to be stored locally. Finally, performance proved most robust when using moderate mini-batch sizes of at least 256 documents.

These results demonstrate that organizations can significantly reduce compute costs, infrastructure overhead, and processing timelines when analyzing massive document repositories. Rather than maintaining heavy storage and parallel hardware clusters, teams can deploy lightweight, continuous text processing pipelines that adapt immediately as new information arrives.

Organizations handling large document archives or continuous text streams should adopt online variational inference in place of batch methods. When deploying this approach, practitioners should configure the algorithm using mini-batch sizes between 256 and 4,096 documents to ensure stability and rapid convergence. Teams can also extend this online framework to other hierarchical Bayesian models used for large-scale structured data analysis.

Confidence in these findings is high for standard text modeling tasks based on held-out predictive metrics. However, readers should note that the evaluation relied on perplexity—a standard measure of predictive likelihood—which may not always align perfectly with subjective human judgments of topic coherence. Additionally, while the algorithm theoretically converges to a locally optimal solution, performance depends on proper tuning of the learning rate and mini-batch size parameters.

Cover for Online Learning for Latent Dirichlet Allocation

Abstract

We develop an online variational Bayes (VB) algorithm for Latent Dirichlet Allocation (LDA). Online LDA is based on online stochastic optimization with a natural gradient step, which we show converges to a local optimum of the VB objective function. It can handily analyze massive document collections, including those arriving in a stream. We study the performance of online LDA in several ways, including by fitting a 100-topic topic model to 3.3M articles from Wikipedia in a single pass. We demonstrate that online LDA finds topic models as good or better than those found with batch VB, and in a fraction of the time.

Table of Contents

  • 1 Introduction
  • 2 Online variational Bayes for latent Dirichlet allocation
  • 2.1 Batch variational Bayes for LDA
  • 2.2 Online variational inference for LDA
  • 2.3 Analysis of convergence
  • 3 Related work
  • 4 Experiments
  • 5 Discussion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Online Variational Bayes for Latent Dirichlet Allocation

    algorithm

    Online Variational Bayes (Online VB) for Latent Dirichlet Allocation (LDA) fits global variational topic parameters λ∈RK×W\lambda \in \mathbb{R}^{K \times W} (where KK is the number of topics and WW is vocabulary size) from a continuous stream or large corpus of documents without storing the entire dataset in memory. At each step tt, it processes a mini-batch of S≥1S \ge 1 documents, performs local coordinate ascent over per-document variational Dirichlet parameters γ∈RS×K\gamma \in \mathbb{R}^{S \times K} and per-word assignment distributions ϕ∈RS×W×K\phi \in \mathbb{R}^{S \times W \times K}, computes an intermediate global topic estimate λ~\tilde{\lambda} scaled to the corpus size DD, and updates λ\lambda via a stochastic step with learning rate ρt=(τ0+t)−κ\rho_t = (\tau_0 + t)^{-\kappa}.

    Input: Stream of document word count vectors ntn_t, corpus size DD, mini-batch size SS, topic count KK, vocabulary size WW, priors α∈RK,η∈RW\alpha \in \mathbb{R}^K, \eta \in \mathbb{R}^W, step decay κ∈(0.5,1]\kappa \in (0.5, 1], delay τ0≥0\tau_0 \ge 0
    Output: Variational topic parameters λ∈RK×W\lambda \in \mathbb{R}^{K \times W}
    Initialize λ∈RK×W\lambda \in \mathbb{R}^{K \times W} randomly with positive values
    for t=0,1,2,…t = 0, 1, 2, \dots do
        Sample a mini-batch of SS documents {nts}s=1S\{n_{ts}\}_{s=1}^S
        for each document s∈{1,…,S}s \in \{1, \dots, S\} do
            Initialize γtsk=1\gamma_{tsk} = 1 for all k∈{1,…,K}k \in \{1, \dots, K\}
            repeat
                for each word ww where ntsw>0n_{tsw} > 0 and topic kk do
                    ϕtswk∝exp⁡(Ψ(γtsk)−Ψ(∑i=1Kγtsi)+Ψ(λkw)−Ψ(∑v=1Wλkv))\phi_{tswk} \propto \exp\left( \Psi(\gamma_{tsk}) - \Psi\left(\sum_{i=1}^K \gamma_{tsi}\right) + \Psi(\lambda_{kw}) - \Psi\left(\sum_{v=1}^W \lambda_{kv}\right) \right)
                    Normalize ϕtsw\phi_{tsw} such that ∑k=1Kϕtswk=1\sum_{k=1}^K \phi_{tswk} = 1
                end for
                for each topic kk do
                    γtsk=αk+∑w=1Wntswϕtswk\gamma_{tsk} = \alpha_k + \sum_{w=1}^W n_{tsw} \phi_{tswk}
                end for
            until 1K∑k=1K∣Δγtsk∣<10−5\frac{1}{K} \sum_{k=1}^K |\Delta \gamma_{tsk}| < 10^{-5}
        end for
        for each topic kk and word ww do
            λ~kw=ηw+DS∑s=1Sntswϕtswk\tilde{\lambda}_{kw} = \eta_w + \frac{D}{S} \sum_{s=1}^S n_{tsw} \phi_{tswk}
        end for
        ρt=(τ0+t)−κ\rho_t = (\tau_0 + t)^{-\kappa}
        λ=(1−ρt)λ+ρtλ~\lambda = (1 - \rho_t) \lambda + \rho_t \tilde{\lambda}
    end for
    return λ\lambda

    Here Ψ(⋅)\Psi(\cdot) denotes the digamma function (the first derivative of the log-gamma function, ddxlog⁡Γ(x)\frac{d}{dx} \log \Gamma(x)). In the streaming limit where the collection size is unbounded, D→∞D \to \infty, corresponding to empirical Bayes estimation of topic multinomials β\beta.

  2. Knowl 2 — Stochastic Natural Gradient Interpretation and Convergence of Online LDA

    theoretical result

    Online Variational Bayes for Latent Dirichlet Allocation optimizes the dataset Evidence Lower Bound (ELBO)

    L(n,λ)=∑d=1Dℓ(nd,γ(nd,λ),ϕ(nd,λ),λ)\mathcal{L}(\mathbf{n}, \lambda) = \sum_{d=1}^D \ell(\mathbf{n}_d, \gamma(\mathbf{n}_d, \lambda), \phi(\mathbf{n}_d, \lambda), \lambda)

    by computing stochastic natural gradients with respect to the variational topic Dirichlet parameters λ\lambda.

    For an observed document nt\mathbf{n}_t, the gradient of the per-document ELBO contribution ℓ\ell with respect to λkw\lambda_{kw} is:

    ∂ℓ(nt,γt,ϕt,λ)∂λkw=∑v=1W−∂2log⁡q(βk)∂λkv∂λkw(−λkvD+ηvD+ntvϕtvk)\frac{\partial \ell(\mathbf{n}_t, \gamma_t, \phi_t, \lambda)}{\partial \lambda_{kw}} = \sum_{v=1}^W - \frac{\partial^2 \log q(\beta_k)}{\partial \lambda_{kv} \partial \lambda_{kw}} \left( -\frac{\lambda_{kv}}{D} + \frac{\eta_v}{D} + n_{tv} \phi_{tvk} \right)

    where −∂2log⁡q(βk)∂λk∂λkT-\frac{\partial^2 \log q(\beta_k)}{\partial \lambda_k \partial \lambda_k^T} is the Fisher information matrix of the variational distribution q(βk)=Dirichlet(βk;λk)q(\beta_k) = \text{Dirichlet}(\beta_k; \lambda_k). Multiplying the gradient vector ∇λkℓ\nabla_{\lambda_k} \ell by the inverse of this Fisher information matrix removes the curvature term and yields the natural gradient:

    [−∂2log⁡q(βk)∂λk∂λkT]−1∇λkℓ(nt,γt,ϕt,λ)=−λkD+ηD+ntϕt⋅k\left[ - \frac{\partial^2 \log q(\beta_k)}{\partial \lambda_k \partial \lambda_k^T} \right]^{-1} \nabla_{\lambda_k} \ell(\mathbf{n}_t, \gamma_t, \phi_t, \lambda) = -\frac{\lambda_k}{D} + \frac{\eta}{D} + \mathbf{n}_t \phi_{t\cdot k}

    Multiplying by DD and taking a step of size ρt\rho_t gives the update rule:

    λ←(1−ρt)λ+ρtλ~t\lambda \leftarrow (1 - \rho_t)\lambda + \rho_t \tilde{\lambda}_t

    where λ~kw=ηw+Dntwϕtwk\tilde{\lambda}_{kw} = \eta_w + D n_{tw} \phi_{twk}. Under the Robbins-Monro step-size conditions ∑t=0∞ρt=∞\sum_{t=0}^\infty \rho_t = \infty and ∑t=0∞ρt2<∞\sum_{t=0}^\infty \rho_t^2 < \infty (satisfied when ρt=(τ0+t)−κ\rho_t = (\tau_0 + t)^{-\kappa} with κ∈(0.5,1]\kappa \in (0.5, 1] and τ0≥0\tau_0 \ge 0), the sequence of iterates λ\lambda is guaranteed to converge asymptotically to a stationary point (local maximum) of the variational objective L\mathcal{L}.

  3. Knowl 3 — Per-Document Decomposed Variational Evidence Lower Bound (ELBO) for LDA

    equation

    Under the fully factorized variational distribution family

    q(z,θ,β)=∏k=1KDirichlet(βk;λk)∏d=1D[Dirichlet(θd;γd)∏i=1NdDiscrete(zdi;ϕdwdi)]q(\mathbf{z}, \theta, \beta) = \prod_{k=1}^K \text{Dirichlet}(\beta_k; \lambda_k) \prod_{d=1}^D \left[ \text{Dirichlet}(\theta_d; \gamma_d) \prod_{i=1}^{N_d} \text{Discrete}(z_{di}; \phi_{d w_{di}}) \right]

    the Evidence Lower Bound (ELBO) on the marginal log-likelihood log⁡p(w∣α,η)\log p(\mathbf{w} \mid \alpha, \eta) decomposes as a sum over document-level contributions L(w,ϕ,γ,λ)=∑d=1Dℓ(nd,ϕd,γd,λ)\mathcal{L}(\mathbf{w}, \phi, \gamma, \lambda) = \sum_{d=1}^D \ell(\mathbf{n}_d, \phi_d, \gamma_d, \lambda), where nd=(nd1,…,ndW)\mathbf{n}_d = (n_{d1}, \dots, n_{dW}) is the word count vector for document dd, WW is the vocabulary size, and KK is the number of topics:

    ℓ(nd,ϕd,γd,λ)=∑w=1Wndw∑k=1Kϕdwk(Eq[log⁡θdk]+Eq[log⁡βkw]−log⁡ϕdwk)−log⁡Γ(∑k=1Kγdk)+∑k=1K(αk−γdk)Eq[log⁡θdk]+∑k=1Klog⁡Γ(γdk)+1D∑k=1K[−log⁡Γ(∑w=1Wλkw)+∑w=1W(ηw−λkw)Eq[log⁡βkw]+∑w=1Wlog⁡Γ(λkw)]+log⁡Γ(∑k=1Kαk)−∑k=1Klog⁡Γ(αk)+1D[log⁡Γ(∑w=1Wηw)−∑w=1Wlog⁡Γ(ηw)]\begin{aligned} \ell(\mathbf{n}_d, \phi_d, \gamma_d, \lambda) = & \sum_{w=1}^W n_{dw} \sum_{k=1}^K \phi_{dwk} \left( \mathbb{E}_q[\log \theta_{dk}] + \mathbb{E}_q[\log \beta_{kw}] - \log \phi_{dwk} \right) \\ & - \log \Gamma\left( \sum_{k=1}^K \gamma_{dk} \right) + \sum_{k=1}^K (\alpha_k - \gamma_{dk}) \mathbb{E}_q[\log \theta_{dk}] + \sum_{k=1}^K \log \Gamma(\gamma_{dk}) \\ & + \frac{1}{D} \sum_{k=1}^K \left[ -\log \Gamma\left( \sum_{w=1}^W \lambda_{kw} \right) + \sum_{w=1}^W (\eta_w - \lambda_{kw}) \mathbb{E}_q[\log \beta_{kw}] + \sum_{w=1}^W \log \Gamma(\lambda_{kw}) \right] \\ & + \log \Gamma\left( \sum_{k=1}^K \alpha_k \right) - \sum_{k=1}^K \log \Gamma(\alpha_k) + \frac{1}{D} \left[ \log \Gamma\left( \sum_{w=1}^W \eta_w \right) - \sum_{w=1}^W \log \Gamma(\eta_w) \right] \end{aligned}

    The expectations under qq are evaluated using the digamma function Ψ(x)=ddxlog⁡Γ(x)\Psi(x) = \frac{d}{dx} \log \Gamma(x):

    Eq[log⁡θdk]=Ψ(γdk)−Ψ(∑i=1Kγdi),Eq[log⁡βkw]=Ψ(λkw)−Ψ(∑v=1Wλkv)\mathbb{E}_q[\log \theta_{dk}] = \Psi(\gamma_{dk}) - \Psi\left( \sum_{i=1}^K \gamma_{di} \right), \quad \mathbb{E}_q[\log \beta_{kw}] = \Psi(\lambda_{kw}) - \Psi\left( \sum_{v=1}^W \lambda_{kv} \right)

  4. Knowl 4 — Online Hyperparameter Updates for Latent Dirichlet Allocation

    model/method

    Dirichlet prior hyperparameters α∈RK\alpha \in \mathbb{R}^K (the per-document topic prior) and η∈RW\eta \in \mathbb{R}^W (the per-topic word prior) can be estimated online concurrently with variational parameters using stochastic Newton-Raphson gradient steps:

    α←α−ρtα~(γt),η←η−ρtη~(λ)\alpha \leftarrow \alpha - \rho_t \tilde{\alpha}(\gamma_t), \quad \eta \leftarrow \eta - \rho_t \tilde{\eta}(\lambda)

    where ρt=(τ0+t)−κ\rho_t = (\tau_0 + t)^{-\kappa} is the learning step size, α~(γt)=Hα−1∇αℓ(nt,γt,ϕt,λ)\tilde{\alpha}(\gamma_t) = H_\alpha^{-1} \nabla_\alpha \ell(\mathbf{n}_t, \gamma_t, \phi_t, \lambda) is the product of the inverse Hessian matrix and the gradient with respect to α\alpha of the per-document ELBO contribution ℓ\ell, and η~(λ)=Hη−1∇ηL\tilde{\eta}(\lambda) = H_\eta^{-1} \nabla_\eta \mathcal{L} is the product of the inverse Hessian matrix and the gradient with respect to η\eta of the corpus ELBO L\mathcal{L}. Both Hessian inversions can be evaluated in linear time.

  5. Knowl 5 — Held-Out Perplexity Lower Bound Proxy for Topic Models

    equation

    Model performance and generalization on unseen test documents ntest\mathbf{n}^{\text{test}} are evaluated using held-out perplexity, defined as the geometric mean of the inverse marginal probability of words in the held-out collection:

    perplexity(ntest,λ,α)=exp⁡{−∑ilog⁡p(nitest∣α,β)∑i,wniwtest}\text{perplexity}(\mathbf{n}^{\text{test}}, \lambda, \alpha) = \exp \left\{ - \frac{\sum_i \log p(\mathbf{n}_i^{\text{test}} \mid \alpha, \beta)}{\sum_{i, w} n_{iw}^{\text{test}}} \right\}

    Because the exact marginal log-likelihood log⁡p(nitest∣α,β)\log p(\mathbf{n}_i^{\text{test}} \mid \alpha, \beta) is intractable, a variational upper bound on perplexity is used as a computable proxy:

    perplexity(ntest,λ,α)≤exp⁡{−∑i(Eq[log⁡p(nitest,θi,zi∣α,β)]−Eq[log⁡q(θi,zi)])∑i,wniwtest}\text{perplexity}(\mathbf{n}^{\text{test}}, \lambda, \alpha) \le \exp \left\{ - \frac{\sum_i \left( \mathbb{E}_q[\log p(\mathbf{n}_i^{\text{test}}, \theta_i, \mathbf{z}_i \mid \alpha, \beta)] - \mathbb{E}_q[\log q(\theta_i, \mathbf{z}_i)] \right)}{\sum_{i, w} n_{iw}^{\text{test}}} \right\}

    where variational parameters γi\gamma_i and ϕi\phi_i for each held-out document ii are fit via the variational E-step with global topics λ\lambda held fixed.

  6. Knowl 6 — Impact of Mini-Batch Size and Learning Rate Schedule on Online LDA

    data/table

    The performance of Online LDA is sensitive to the choice of mini-batch size SS, the forgetting rate κ∈(0.5,1.0]\kappa \in (0.5, 1.0], and the early-iteration delay τ0≥0\tau_0 \ge 0. Across 288 parameter evaluations run for five hours of CPU time on the Nature corpus (352,549 documents, vocabulary size 4,253) and a Wikipedia corpus (100,000 documents, vocabulary size 7,995) with K=100K = 100 topics and α=η=0.01\alpha = \eta = 0.01, the best parameter combinations and resulting held-out perplexities are:

    Best parameter settings for Nature corpus
    SS 1 4 16 64 256 1024 4096 16384
    κ\kappa 0.9 0.8 0.8 0.7 0.6 0.5 0.5 0.5
    τ0\tau_0 1024 1024 1024 1024 1024 256 64 1
    Perplexity 1132 1087 1052 1053 1042 1031 1030 1046
    Best parameter settings for Wikipedia corpus
    SS 1 4 16 64 256 1024 4096 16384
    κ\kappa 0.9 0.9 0.8 0.7 0.6 0.5 0.5 0.5
    τ0\tau_0 1024 1024 1024 1024 1024 1024 64 1
    Perplexity 675 640 611 595 588 584 580 584

    The lowest perplexity on both corpora is achieved at S=4096,κ=0.5,τ0=64S = 4096, \kappa = 0.5, \tau_0 = 64. Smaller mini-batches require higher forgetting rates κ\kappa (e.g., 0.80.8--0.90.9) and larger delays τ0=1024\tau_0 = 1024 to maintain stability. Mini-batch sizes of S≥256S \ge 256 consistently yield superior and stable perplexity relative to smaller mini-batches (S≤64S \le 64).

  7. Knowl 7 — Empirical Convergence Speed and Held-Out Perplexity of Online vs. Batch Variational Bayes

    empirical result

    Experiments evaluating 100-topic LDA on both static corpora (Nature, 352,549 articles; Wikipedia, 100,000 articles) and a streaming corpus (3.3 million Wikipedia articles) show that Online Variational Bayes converges significantly faster than standard batch Variational Bayes while finding models of equal or superior predictive quality (measured by held-out perplexity):

    1. On the 352k-document Nature corpus, Online LDA with mini-batch sizes S≥256S \ge 256 reached the final perplexity of 24-hour batch VB in under 1 hour of CPU time.
    2. On the 100k-document Wikipedia corpus, Online LDA achieved a lower final held-out perplexity (~580) than batch VB run on the full 98k training set (~700) or on a 10k subset (~820).
    3. In a streaming setup analyzing 3.3 million Wikipedia articles at a rate of ~60,000 articles per hour in a single pass without storing documents (setting κ=0.5,τ0=1024,S=1024\kappa = 0.5, \tau_0 = 1024, S = 1024), Online LDA converged after viewing approximately 1.65 million documents (under 3 days of processing), whereas a single iteration of batch VB over 3.3 million documents would take multiple days.

Coverage note — Algorithm 1 (standard batch Variational Bayes for LDA) was omitted as a separate knowl because it is standard prior work from Blei et al. (2003) and serves here as the baseline for comparison.

References

  1. 1.M. Braun and J. McAuliffe. Variational inference for large-scale models of discrete choice. arXiv, (0712.2526), 2008.
  2. 2.D. Blei and M. Jordan. Variational methods for the Dirichlet process. In Proc. 21st Int’l Conf. on Machine Learning, 2004.
  3. 3.A. Asuncion, M. Welling, P. Smyth, and Y.W. Teh. On smoothing and inference for topic models. In Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence, 2009.
  4. 4.D. Newman, A. Asuncion, P. Smyth, and M. Welling. Distributed inference for latent Dirichlet allocation. In Neural Information Processing Systems, 2007.
  5. 5.Feng Yan, Ningyi Xu, and Yuan Qi. Parallel inference for latent Dirichlet allocation on graphics processing units. In Advances in Neural Information Processing Systems 22, pages 2134–2142, 2009.
  6. 6.L. Bottou and O. Bousquet. The tradeoffs of large scale learning. In Advances in Neural Information Processing Systems, volume 20, pages 161–168. NIPS Foundation (http://books.nips.cc), 2008.
  7. 7.D. Blei, A. Ng, and M. Jordan. Latent Dirichlet allocation. Journal of Machine Learning Research, 3:993–1022, January 2003.
  8. 8.Hanna Wallach, David Mimno, and Andrew McCallum. Rethinking lda: Why priors matter. In Advances in Neural Information Processing Systems 22, pages 1973–1981, 2009.
  9. 9.W. Buntine. Variational extentions to EM and multinomial PCA. In European Conf. on Machine Learning, 2002.
  10. 10.J. Mairal, F. Bach, J. Ponce, and G. Sapiro. Online learning for matrix factorization and sparse coding. Journal of Machine Learning Research, 11(1):19–60, 2010.
  11. 11.L. Yao, D. Mimno, and A. McCallum. Efficient methods for topic model inference on streaming document collections. In KDD 2009: Proc. 15th ACM SIGKDD int’l Conf. on Knowledge discovery and data mining, pages 937–946, 2009.
  12. 12.M. Jordan, Z. Ghahramani, T. Jaakkola, and L. Saul. Introduction to variational methods for graphical models. Machine Learning, 37:183–233, 1999.
  13. 13.H. Attias. A variational Bayesian framework for graphical models. In Advances in Neural Information Processing Systems 12, 2000.
  14. 14.A. Dempster, N. Laird, and D. Rubin. Maximum likelihood from incomplete data via the EM algorithm. Journal of the Royal Statistical Society, Series B, 39:1–38, 1977.
  15. 15.L. Bottou and N. Murata. Stochastic approximations and efficient learning. The Handbook of Brain Theory and Neural Networks, Second edition. The MIT Press, Cambridge, MA, 2002.
  16. 16.M.A. Sato. Online model selection based on the variational Bayes. Neural Computation, 13(7):1649–1681, 2001.
  17. 17.P. Liang and D. Klein. Online EM for unsupervised models. In Proc. Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics, pages 611–619, 2009.
  18. 18.H. Robbins and S. Monro. A stochastic approximation method. The Annals of Mathematical Statistics, 22(3):400–407, 1951.
  19. 19.L. Bottou. Online learning and stochastic approximations. Cambridge University Press, Cambridge, UK, 1998.
  20. 20.R.M. Neal and G.E. Hinton. A view of the EM algorithm that justifies incremental, sparse, and other variants. Learning in graphical models, 89:355–368, 1998.
  21. 21.M.A. Sato and S. Ishii. On-line EM algorithm for the normalized Gaussian network. Neural Computation, 12(2):407–432, 2000.
  22. 22.T. Griffiths and M. Steyvers. Finding scientific topics. Proc. National Academy of Science, 2004.
  23. 23.X. Song, C.Y. Lin, B.L. Tseng, and M.T. Sun. Modeling and predicting personal information dissemination behavior. In KDD 2005: Proc. 11th ACM SIGKDD int’l Conf. on Knowledge discovery and data mining. ACM, 2005.
  24. 24.K.R. Canini, L. Shi, and T.L. Griffiths. Online inference of topics with latent Dirichlet allocation. In Proceedings of the International Conference on Artificial Intelligence and Statistics, volume 5, 2009.
  25. 25.J. Chang, J. Boyd-Graber, S. Gerrish, C. Wang, and D. Blei. Reading tea leaves: How humans interpret topic models. In Advances in Neural Information Processing Systems 21 (NIPS), 2009.

Citation

MLA
Hoffman, M., et al. “Online Learning for Latent Dirichlet Allocation”. Advances in Neural Information Processing Systems, vol. 23, 2010, https://proceedings.neurips.cc/paper_files/paper/2010/file/71f6278d140af599e06ad9bf1ba03cb0-Paper.pdf.
APA
Hoffman, M., Bach, F., & Blei, D. (2010). Online Learning for Latent Dirichlet Allocation. Advances in Neural Information Processing Systems, 23. https://proceedings.neurips.cc/paper_files/paper/2010/file/71f6278d140af599e06ad9bf1ba03cb0-Paper.pdf
Chicago
Hoffman, M., F. Bach, and D. Blei. 2010. “Online Learning for Latent Dirichlet Allocation”. Advances in Neural Information Processing Systems 23. https://proceedings.neurips.cc/paper_files/paper/2010/file/71f6278d140af599e06ad9bf1ba03cb0-Paper.pdf.
Harvard
Hoffman, M., Bach, F. and Blei, D. (2010) “Online Learning for Latent Dirichlet Allocation”, Advances in Neural Information Processing Systems. Curran Associates, Inc. Available at: https://proceedings.neurips.cc/paper_files/paper/2010/file/71f6278d140af599e06ad9bf1ba03cb0-Paper.pdf.
Vancouver
1. Hoffman M, Bach F, Blei D (2010) Online Learning for Latent Dirichlet Allocation. Advances in Neural Information Processing Systems 23:

BibTeX

@inproceedings{hoffman2010online,
  title = {Online Learning for Latent Dirichlet Allocation},
  author = {Hoffman, Matthew and Bach, Francis and Blei, David},
  year = {2010},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {23},
  url = {https://proceedings.neurips.cc/paper_files/paper/2010/file/71f6278d140af599e06ad9bf1ba03cb0-Paper.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors