Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multimedia information retrieval

Multimedia information retrieval is a subdiscipline of computer science and information retrieval focused on indexing, searching, organizing, and retrieving data across diverse modalities, including text, images, audio, and video. Unlike traditional text-based search systems, it analyzes both low-level perceptual features and high-level semantic content to understand complex media representations. The field encompasses cross-modal and multimodal retrieval techniques, enabling users to submit queries in one or multiple modalities and retrieve relevant items in the same or different formats. Key challenges within the discipline include bridging the semantic gap between raw sensory data and human interpretation, fusing heterogeneous data streams into unified representations, and scaling retrieval algorithms to manage large, unstructured multimedia databases efficiently.

1 item

Multimodal learning with deep Boltzmann machines

Multimodal learning with deep Boltzmann machines

Nitish Srivastava, Ruslan Salakhutdinov

OrganizationsUniversity of Toronto

Why you should read this

Proposes a Multimodal Deep Boltzmann Machine that learns a joint generative model across disparate modalities like images and text, enabling effective classification, cross-modal retrieval, and the reconstruction of missing inputs.

A Deep Boltzmann Machine is described for learning a generative model of data that consists of multiple and diverse input modalities. The model can be used to extract a unified representation that fuses modalities together. We find that this representation is useful for classification and information retrieval tasks. The model works by learning a probability density over the space of multimodal inputs. It uses states of latent variables as representations of the input. The model can extract this representation even when some modalities are absent by sampling from the conditional distribution over them and filling them in. Our experimental results on bi-modal data consisting of images and text show that the Multimodal DBM can learn a good generative model of the joint space of image and text inputs that is useful for information retrieval from both unimodal and multimodal queries. We further demonstrate that this model significantly outperforms SVMs and LDA on discriminative tasks. Finally, we compare our model to other deep learning methods, including autoencoders and deep belief networks, and show that it achieves noticeable gains.

Added

2026-09-18