Multimodal learning with deep Boltzmann machines

Nitish SrivastavaRuslan Salakhutdinov

article2012JMLR1,822 citations

Proposes a Multimodal Deep Boltzmann Machine that learns a joint generative model across disparate modalities like images and text, enabling effective classification, cross-modal retrieval, and the reconstruction of missing inputs.

Listen

Real-world data increasingly arrives through multiple distinct channels, such as paired images and descriptive text. Integrating these sources is challenging because each channel has fundamentally different statistical properties: text is typically represented as sparse word counts, whereas images are dense, continuous numerical features. Traditional machine learning methods struggle to discover non-linear relationships across these differing formats, often require fully labeled data, and fail when one channel is missing or incomplete.

The article evaluates a unified framework, the Multimodal Deep Boltzmann Machine, designed to extract joint representations from multi-channel data. The objective is to demonstrate that a single probabilistic generative system can effectively combine distinct modalities, leverage vast pools of unlabeled data, handle missing inputs, and improve performance on both classification and information retrieval tasks.

To assess the model, the authors conducted empirical evaluations on the MIR Flickr dataset, utilizing 1 million image-text pairs. The architecture processes each modality through independent specialized pathways before fusing them into a central joint representation layer. Nearly 975,000 unlabeled pairs were used for unsupervised pretraining, while 25,000 labeled items spanning 38 topical classes were split into training and test sets. Performance was measured using Mean Average Precision across classification and retrieval benchmarks against standard baselines and alternative deep architectures.

The findings show that the proposed framework consistently delivers superior performance across evaluated tasks. When classifying multi-channel inputs, the model achieved a Mean Average Precision of 0.609, outperforming traditional Linear Discriminant Analysis (0.492) and Support Vector Machines (0.475), while also surpassing deep autoencoders (0.600) and deep belief networks (0.599). Incorporating unlabeled data during pretraining substantially boosted performance, lifting the model's precision from 0.526 to 0.585. Crucially, when text was entirely missing during testing, the model successfully synthesized proxy text to achieve a score of 0.531, significantly exceeding unimodal image classifiers which scored between 0.375 and 0.469. In cross-modal retrieval experiments, the architecture similarly attained top-tier performance, reaching 0.622 precision for multimodal queries and 0.614 for image-only queries.

These results indicate that organizations do not need to build and maintain separate, siloed algorithms for every possible combination of complete and missing inputs. A single joint generative architecture can serve as a robust general-purpose engine, reducing operational complexity while leveraging abundant, low-cost unlabeled data to improve overall accuracy. By generating plausible substitutes for missing attributes, the system mitigates the operational risks associated with noisy, real-world data collection.

Organizations handling heterogeneous data streams should consider adopting modular, multi-pathway generative architectures rather than purely unimodal or discriminative pipelines. Decision-makers should prioritize unsupervised pretraining on large unlabeled datasets to minimize manual labeling expenses and pilot the deployment of single unified models to streamline infrastructure.

While confidence in the empirical gains is supported by consistent benchmark improvements across multiple splits, several limitations remain. The evaluations focused solely on paired image and text data using fixed feature extractors on a single web dataset. Stakeholders should exercise caution when extrapolating these conclusions to different domains, such as audio or sensor telemetry, where further pilot testing and validation are advised before broad implementation.

  • Paper: Deep Boltzmann Machines, Ruslan Salakhutdinov et al. (2009). Read this account of Deep Boltzmann Machines first to understand the generative architecture and learning methods that the source adapts for joint multimodal representations.
  • Paper: Multimodal Deep Learning, Jiquan Ngiam et al. (2011). This earlier study establishes deep networks for learning shared representations across modalities, a foundation that clarifies the source’s multimodal modeling approach.
Cover for Multimodal learning with deep Boltzmann machines

Abstract

A Deep Boltzmann Machine is described for learning a generative model of data that consists of multiple and diverse input modalities. The model can be used to extract a unified representation that fuses modalities together. We find that this representation is useful for classification and information retrieval tasks. The model works by learning a probability density over the space of multimodal inputs. It uses states of latent variables as representations of the input. The model can extract this representation even when some modalities are absent by sampling from the conditional distribution over them and filling them in. Our experimental results on bi-modal data consisting of images and text show that the Multimodal DBM can learn a good generative model of the joint space of image and text inputs that is useful for information retrieval from both unimodal and multimodal queries. We further demonstrate that this model significantly outperforms SVMs and LDA on discriminative tasks. Finally, we compare our model to other deep learning methods, including autoencoders and deep belief networks, and show that it achieves noticeable gains.

Table of Contents

  • 1 Introduction
  • 2 Background: RBMs and Their Generalizations
  • 2.1 Restricted Boltzmann Machines
  • 2.2 Gaussian RBM
  • 2.3 Replicated Softmax Model
  • 3 Multimodal Deep Boltzmann Machine
  • 3.1 Approximate Learning and Inference
  • 3.2 Salient Features
  • 3.3 Modeling Tasks
  • 4 Experiments
  • 4.1 Dataset and Feature Extraction
  • 4.2 Model Architecture and Learning
  • 4.3 Classification Tasks
  • 4.4 Retrieval Tasks
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Multimodal Deep Boltzmann Machine Architecture and Joint Probability Distribution

    model/method

    The Multimodal Deep Boltzmann Machine (Multimodal DBM) is a deep generative undirected graphical model that integrates disparate data modalities (such as real-valued image descriptors and discrete text word count vectors) into a unified latent representation. The network consists of separate unimodal DBM pathways that feed into a shared joint hidden layer.

    Let vm∈RD\mathbf{v}_m \in \mathbb{R}^D denote a continuous image feature vector of dimension DD, and let vt∈NK\mathbf{v}_t \in \mathbb{N}^K denote a sparse text word count vector over a vocabulary of size KK. The image pathway comprises two layers of binary hidden units hm(1)∈{0,1}F1\mathbf{h}_m^{(1)} \in \{0, 1\}^{F_1} and hm(2)∈{0,1}F2\mathbf{h}_m^{(2)} \in \{0, 1\}^{F_2}, while the text pathway comprises two layers of binary hidden units ht(1)∈{0,1}F1\mathbf{h}_t^{(1)} \in \{0, 1\}^{F_1} and ht(2)∈{0,1}F2\mathbf{h}_t^{(2)} \in \{0, 1\}^{F_2}. The top of the network contains a shared binary joint hidden layer h(3)∈{0,1}F3\mathbf{h}^{(3)} \in \{0, 1\}^{F_3} with undirected connections to both hm(2)\mathbf{h}_m^{(2)} and ht(2)\mathbf{h}_t^{(2)}.

    The image-specific pathway uses a Gaussian Restricted Boltzmann Machine (RBM) energy function for visible-to-hidden interactions and standard binary RBM energy for hidden-to-hidden interactions:

    P(vm;θ)=1Z(θ)∑hm(1),hm(2)exp⁡(−∑i=1D(vmi−bi)22σi2+∑i=1D∑j=1F1vmiσiWij(1)hmj(1)+∑j=1F1∑l=1F2hmj(1)Wjl(2)hml(2))P(\mathbf{v}_m; \theta) = \frac{1}{Z(\theta)} \sum_{\mathbf{h}_m^{(1)}, \mathbf{h}_m^{(2)}} \exp \left( -\sum_{i=1}^D \frac{(v_{mi} - b_i)^2}{2\sigma_i^2} + \sum_{i=1}^D \sum_{j=1}^{F_1} \frac{v_{mi}}{\sigma_i} W_{ij}^{(1)} h_{mj}^{(1)} + \sum_{j=1}^{F_1} \sum_{l=1}^{F_2} h_{mj}^{(1)} W_{jl}^{(2)} h_{ml}^{(2)} \right)

    where {bi}\{b_i\} are visible biases, {σi2}\{\sigma_i^2\} are visible noise variances (typically fixed to 1), and W(1),W(2)\mathbf{W}^{(1)}, \mathbf{W}^{(2)} are symmetric weight matrices.

    The text pathway models word counts via a Replicated Softmax visible-to-hidden energy with document word count M=∑kvtkM = \sum_k v_{tk}, connected to higher binary hidden layers. Joining these pathways through h(3)\mathbf{h}^{(3)}, the overall joint probability distribution over the multimodal inputs is:

    P(vm,vt;θ)=∑hm(2),ht(2),h(3)P(hm(2),ht(2),h(3))(∑hm(1)P(vm,hm(1),hm(2)))(∑ht(1)P(vt,ht(1),ht(2)))P(\mathbf{v}_m, \mathbf{v}_t; \theta) = \sum_{\mathbf{h}_m^{(2)}, \mathbf{h}_t^{(2)}, \mathbf{h}^{(3)}} P(\mathbf{h}_m^{(2)}, \mathbf{h}_t^{(2)}, \mathbf{h}^{(3)}) \left( \sum_{\mathbf{h}_m^{(1)}} P(\mathbf{v}_m, \mathbf{h}_m^{(1)}, \mathbf{h}_m^{(2)}) \right) \left( \sum_{\mathbf{h}_t^{(1)}} P(\mathbf{v}_t, \mathbf{h}_t^{(1)}, \mathbf{h}_t^{(2)}) \right)

  2. Knowl 2 — Mean-Field Inference and Learning in Multimodal Deep Boltzmann Machines

    model/method

    Exact maximum likelihood learning in the Multimodal Deep Boltzmann Machine is intractable due to the partition function and recurrent layer interactions. Model training and inference are conducted using a combination of variational mean-field inference, Markov Chain Monte Carlo (MCMC) stochastic approximation, and greedy layer-wise pretraining.

    For an observed pair v={vm,vt}\mathbf{v} = \{\mathbf{v}_m, \mathbf{v}_t\}, the true posterior over all hidden units h={hm(1),hm(2),ht(1),ht(2),h(3)}\mathbf{h} = \{\mathbf{h}_m^{(1)}, \mathbf{h}_m^{(2)}, \mathbf{h}_t^{(1)}, \mathbf{h}_t^{(2)}, \mathbf{h}^{(3)}\} is approximated by a fully factorized mean-field distribution Q(h∣v;μ)Q(\mathbf{h}|\mathbf{v}; \boldsymbol{\mu}):

    Q(h∣v;μ)=(∏j=1F1q(hmj(1)∣v)∏l=1F2q(hml(2)∣v))(∏j=1F1q(htj(1)∣v)∏l=1F2q(htl(2)∣v))∏k=1F3q(hk(3)∣v)Q(\mathbf{h}|\mathbf{v}; \boldsymbol{\mu}) = \left( \prod_{j=1}^{F_1} q(h_{mj}^{(1)}|\mathbf{v}) \prod_{l=1}^{F_2} q(h_{ml}^{(2)}|\mathbf{v}) \right) \left( \prod_{j=1}^{F_1} q(h_{tj}^{(1)}|\mathbf{v}) \prod_{l=1}^{F_2} q(h_{tl}^{(2)}|\mathbf{v}) \right) \prod_{k=1}^{F_3} q(h_k^{(3)}|\mathbf{v})

    where μ={μm(1),μm(2),μt(1),μt(2),μ(3)}\boldsymbol{\mu} = \{\boldsymbol{\mu}_m^{(1)}, \boldsymbol{\mu}_m^{(2)}, \boldsymbol{\mu}_t^{(1)}, \boldsymbol{\mu}_t^{(2)}, \boldsymbol{\mu}^{(3)}\} are variational parameters with q(hi(l)=1)=μi(l)q(h_i^{(l)} = 1) = \mu_i^{(l)} for l∈{1,2,3}l \in \{1, 2, 3\}.

    Learning proceeds by:

    1. Initializing the model parameters θ\theta via greedy layer-wise pretraining of modified Restricted Boltzmann Machines (RBMs) on each unimodal pathway using Persistent Contrastive Divergence (PCD).
    2. For each mini-batch, computing the optimal variational parameters μ\boldsymbol{\mu} that maximize the variational lower bound via coordinate ascent mean-field fixed-point equations.
    3. Updating model parameters θ\theta along the gradient of the variational bound, approximating the negative phase (model expectations) using persistent MCMC chains.

    During training, text word count vectors are normalized such that ∑kvtk=5\sum_k v_{tk} = 5, eliminating the requirement of maintaining separate Markov chains for varying document lengths.

  3. Knowl 3 — Missing Modality Generation and Latent Representation Inference

    model/method

    A trained Multimodal Deep Boltzmann Machine operates as a generative model capable of cross-modal synthesis (hallucinating missing modalities) and generating unified representations for downstream tasks.

    1. Missing Modality Generation: When one modality is absent (for instance, an image vm\mathbf{v}_m is observed without text vt\mathbf{v}_t), the observed visible input vm\mathbf{v}_m is clamped at the image input layer while all hidden units are initialized randomly. Alternating Gibbs sampling is run across the network layers to sample from the conditional distribution P(vt∣vm,θ)P(\mathbf{v}_t|\mathbf{v}_m, \theta) (a multinomial distribution over the vocabulary) or P(vm∣vt,θ)P(\mathbf{v}_m|\mathbf{v}_t, \theta) (Gaussian distribution over image features).

    2. Fused Representation Inference: To extract a joint feature vector when both modalities are observed, variational inference is run to compute the mean-field activation probabilities μ(3)=Q(h(3)∣vm,vt)\boldsymbol{\mu}^{(3)} = Q(\mathbf{h}^{(3)}|\mathbf{v}_m, \mathbf{v}_t). When the text modality is absent at inference time, mean-field updates are executed either with text units clamped to zero or by leaving the text input layer unclamped so that it updates dynamically to fill in the missing modality. The expected activation vector of the top layer h(3)\mathbf{h}^{(3)} serves as the fused multimodal feature vector for classification and information retrieval.

  4. Knowl 4 — Bidirectional Information Flow in Multimodal DBMs Versus Multimodal DBNs

    model/method

    Multimodal Deep Boltzmann Machines (DBMs) and Multimodal Deep Belief Networks (DBNs) differ fundamentally in how cross-modal information is integrated:

    • Multimodal DBN Architecture: Consists of directed sigmoid belief networks leading down to each separate visible modality, capped by an undirected RBM at the top layer. Because the connections below the top layer are strictly directed top-down, the responsibility for multimodal fusion falls exclusively on the single top RBM layer. Lower-level hidden states in one modality cannot directly affect lower-level representations in another modality during inference.

    • Multimodal DBM Architecture: Built entirely with bidirectional, undirected connections across all layers. Modality fusion is distributed across all hidden layers throughout the network. The bottom hidden layers capture low-level, modality-specific statistical structures, while subsequent layers progressively discard modality-specific correlations to form an increasingly 'modality-free' representation in the top joint layer. During iterative mean-field inference or Gibbs sampling, states of low-level hidden units in one modality directly propagate signals to adjust low-level states in other modality pathways.

  5. Knowl 5 — Multimodal Classification Performance on the MIR Flickr Benchmark

    data/table

    The Multimodal Deep Boltzmann Machine was evaluated on the 38-class MIR Flickr classification task using 1-vs-all logistic regression trained on the 2048-dimensional joint hidden representation h(3)\mathbf{h}^{(3)}. Performance was measured by Mean Average Precision (MAP) and precision at top 50 predictions (Prec@50), averaged over 5 random splits of 15,000 training and 10,000 test images.

    Model MAP Prec@50
    Random 0.124 0.124
    Linear Discriminant Analysis (LDA) 0.492 0.754
    Support Vector Machine (SVM) 0.475 0.758
    DBM-Lab (trained on labeled data only, no SIFT) 0.526 0.791
    DBM-Unlab (pretrained with 975k unlabeled pairs, no SIFT) 0.585 0.836
    Multimodal Deep Belief Network (DBN) 0.599 0.867
    Deep Autoencoder 0.600 0.875
    Multimodal DBM (full model with SIFT features) 0.609 0.873

    DBM-Lab outperforms baseline discriminative models (SVM and LDA) on the identical feature set. Unsupervised pretraining on 975,000 unlabeled image-text pairs (DBM-Unlab) improves MAP by 0.059 over DBM-Lab. Incorporating Pyramid Histogram of Words (PHOW) dense SIFT features further increases MAP to 0.609, exceeding both the Multimodal DBN (0.599) and the Deep Autoencoder (0.600).

  6. Knowl 6 — Unimodal Image Classification with Missing Modality Hallucination

    data/table

    To assess how a multimodal model improves classification when only unimodal image data is present at test time, models trained with both modalities were tested with image inputs alone on the MIR Flickr dataset.

    Model MAP Prec@50
    Image-SVM 0.375 -
    Image-DBN 0.463 0.801
    Image-DBM 0.469 0.803
    DBM-ZeroText (Multimodal DBM, text clamped to zero) 0.522 0.827
    DBM-GenText (Multimodal DBM, text unclamped and inferred) 0.531 0.832

    Key observations:

    1. DBM-ZeroText (MAP 0.522) substantially outperforms models trained purely on image inputs (Image-DBM at 0.469 and Image-SVM at 0.375), demonstrating that multimodal training regularizes feature learning and provides superior representations even if the second modality is absent during inference.
    2. Allowing the Multimodal DBM to dynamically infer missing text via mean-field updates (DBM-GenText) further boosts MAP to 0.531, demonstrating that the generated text serves as a plausible, effective proxy for missing ground truth.
  7. Knowl 7 — Multimodal and Cross-Modal Information Retrieval on MIR Flickr

    empirical result

    The quality of fused representations learned by the Multimodal Deep Boltzmann Machine was evaluated on an information retrieval task constructed from the MIR Flickr test set. The database consisted of 5,000 randomly chosen image-text pairs, and 1,000 disjoint image-text pairs were used as queries. A retrieved database item was defined as relevant to a query if they shared at least one of the 38 class labels. Cosine similarity in the top latent space was used as the retrieval metric.

    • Multimodal Queries (Image and Text present): The Multimodal DBM achieved a Mean Average Precision (MAP) of 0.622, outperforming the Deep Autoencoder (0.612 MAP) and the Multimodal DBN (0.609 MAP).

    • Unimodal Queries (Image-only queries retrieving multimodal pairs): By inferring the missing text representation for unimodal image queries, the Multimodal DBM achieved a MAP of 0.614. This significantly outperformed unimodal retrieval models evaluated on the same image queries, including an Image-DBM (0.587 MAP) and an Image-DBN (0.578 MAP).

  8. Knowl 8 — Layer-Wise Feature Quality Across Multimodal DBM and DBN Depths

    empirical result

    Evaluating 1-vs-all logistic regression classification on the representations extracted from each intermediate layer of the Multimodal DBM and Multimodal DBN on MIR Flickr demonstrates the following representational dynamics:

    1. Monotonic Representation Improvement: For both architectures, classification Mean Average Precision (MAP) starts lowest at the peripheral input layers (approximately 0.40–0.45 for image input and text input) and increases monotonically as depth increases through the first hidden layers (hm(1),ht(1)\mathbf{h}_m^{(1)}, \mathbf{h}_t^{(1)}) and second hidden layers (hm(2),ht(2)\mathbf{h}_m^{(2)}, \mathbf{h}_t^{(2)}), reaching peak performance at the middle joint hidden layer (h(3)\mathbf{h}^{(3)}).

    2. DBM Superiority Across Layers: The representation at every single layer of the Multimodal DBM achieves a higher MAP than the corresponding layer of the Multimodal DBN (e.g., intermediate DBM image and text hidden layers achieve higher MAP than DBN hidden layers), reflecting the regularizing benefit of full bidirectional coupling throughout the network.

  9. Knowl 9 — MIR Flickr Experimental Setup, Feature Pipeline, and Model Hyperparameters

    experimental setup

    The empirical evaluation utilizes the MIR Flickr collection consisting of 1,000,000 Flickr images with user-assigned tags:

    • Data Splits and Labels: 25,000 images are annotated for 24 topics, expanded to 38 binary classes via stricter saliency subcategories. The labeled set is split into 15,000 training images (further partitioned into 10,000 training and 5,000 validation for classifier tuning) and 10,000 test images. The remaining 975,000 unlabeled images are used for unsupervised model pretraining. Approximately 18% of labeled images (4,551) lack user tags.

    • Feature Representation:

      • Image Modality: 3857-dimensional vector constructed by concatenating Pyramid Histogram of Words (PHOW, multiscale dense SIFT bag-of-words), Gist descriptors, and MPEG-7 descriptors (Edge Histogram, Homogeneous Texture, Color Structure, Color Layout, Scalable Color). Each dimension is mean-centered and normalized to unit variance.
      • Text Modality: Word count vector across the 2,000 most frequent Flickr tags (average 5.15 tags per tagged image).
    • Network Layer Dimensions:

      • Image Pathway: Gaussian RBM visible layer (D=3857D = 3857), first hidden layer (F1=1024F_1 = 1024), second hidden layer (F2=1024F_2 = 1024).
      • Text Pathway: Replicated Softmax visible layer (K=2000K = 2000), first hidden layer (F1=1024F_1 = 1024), second hidden layer (F2=1024F_2 = 1024).
      • Joint Top Layer: Binary hidden layer (F3=2048F_3 = 2048).

Coverage note — No substantial contributed material was omitted.

References

  1. 1.R. R. Salakhutdinov and G. E. Hinton. Deep Boltzmann machines. In Proceedings of the International Conference on Artificial Intelligence and Statistics, volume 12, 2009.
  2. 2.Mark J. Huiskes, Bart Thomee, and Michael S. Lew. New trends and ideas in visual concept detection: the MIR flickr retrieval evaluation initiative. In Multimedia Information Retrieval, pages 527–536, 2010.
  3. 3.M. Guillaumin, J. Verbeek, and C. Schmid. Multimodal semi-supervised learning for image classification. In Computer Vision and Pattern Recognition (CVPR), 2010 IEEE Conference on, pages 902 –909, june 2010.
  4. 4.Eric P. Xing, Rong Yan, and Alexander G. Hauptmann. Mining associated text and images with dual-wing harmoniums. In UAI, pages 633–641. AUAI Press, 2005.
  5. 5.Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y. Ng. Multimodal deep learning. In International Conference on Machine Learning (ICML), Bellevue, USA, June 2011.
  6. 6.Ruslan Salakhutdinov and Geoffrey E. Hinton. Replicated softmax: an undirected topic model. In NIPS, pages 1607–1614. Curran Associates, Inc., 2009.
  7. 7.Geoffrey E. Hinton. Training products of experts by minimizing contrastive divergence. Neural Computation, 14(8):1711–1800, 2002.
  8. 8.T. Tieleman. Training restricted Boltzmann machines using approximations to the likelihood gradient. In ICML. ACM, 2008.
  9. 9.L. Younes. On the convergence of Markovian stochastic algorithms with rapidly decreasing ergodicity rates, March 17 2000.
  10. 10.Mark J. Huiskes and Michael S. Lew. The MIR Flickr retrieval evaluation. In MIR ’08: Proceedings of the 2008 ACM International Conference on Multimedia Information Retrieval, New York, NY, USA, 2008. ACM.
  11. 11.A Bosch, Andrew Zisserman, and X Munoz. Image classification using random forests and ferns. IEEE 11th International Conference on Computer Vision (2007), 23:1–8, 2007.
  12. 12.Aude Oliva and Antonio Torralba. Modeling the shape of the scene: A holistic representation of the spatial envelope. International Journal of Computer Vision, 42:145–175, 2001.
  13. 13.B.S. Manjunath, J.-R. Ohm, V.V. Vasudevan, and A. Yamada. Color and texture descriptors. Circuits and Systems for Video Technology, IEEE Transactions on, 11(6):703 –715, 2001.
  14. 14.A. Vedaldi and B. Fulkerson. VLFeat: An open and portable library of computer vision algorithms, 2008.
  15. 15.Muhammet Bastan, Hayati Cam, Ugur Gudukbay, and Ozgur Ulusoy. Bilvideo-7: An mpeg-7-compatible video indexing and retrieval system. IEEE Multimedia, 17:62–73, 2010.

Citation

MLA
Srivastava, N., and R. Salakhutdinov. “Multimodal Learning with Deep Boltzmann Machines”. Journal of Machine Learning Research, vol. 15, no. 1, 2012, pp. 2222–30, http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.1020.1306.
APA
Srivastava, N., & Salakhutdinov, R. (2012). Multimodal Learning with Deep Boltzmann Machines. Journal of Machine Learning Research, 15(1), 2222–2230. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.1020.1306
Chicago
Srivastava, N., and R. Salakhutdinov. 2012. “Multimodal Learning with Deep Boltzmann Machines”. Journal of Machine Learning Research 15 (1): 2222–30. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.1020.1306.
Harvard
Srivastava, N. and Salakhutdinov, R. (2012) “Multimodal Learning with Deep Boltzmann Machines”, Journal of Machine Learning Research, 15(1), pp. 2222–2230. Available at: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.1020.1306.
Vancouver
1. Srivastava N, Salakhutdinov R (2012) Multimodal Learning with Deep Boltzmann Machines. Journal of Machine Learning Research 15:2222–2230

BibTeX

@article{srivastava2012multimodal,
  title = {Multimodal Learning with Deep Boltzmann Machines},
  author = {Srivastava, Nitish and Salakhutdinov, Ruslan},
  year = {2012},
  journal = {Journal of Machine Learning Research},
  volume = {15},
  number = {1},
  pages = {2222-2230},
  url = {http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.1020.1306}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/