Built independently by an author, for readers. Read the story and support ChapterPal

keyword

video representations

Video representations are compact mathematical or feature embeddings that encode both the spatial visual appearance and the temporal dynamics of video sequences for computational analysis. While representations of static images capture spatial attributes such as objects, colors, and textures within a single frame, video representations additionally model the motion, sequential patterns, and state changes that occur across consecutive frames over time. These features are extracted using machine learning models, such as three-dimensional convolutional neural networks, recurrent architectures, or spatiotemporal transformers, trained through supervised, unsupervised, or self-supervised methods. By condensing complex, high-dimensional video streams into structured vector spaces, video representations facilitate various downstream applications, including human action recognition, video classification, event detection, frame synthesis, and cross-modal retrieval.

4 items

Unsupervised Learning of Video Representations using LSTMs

Unsupervised Learning of Video Representations using LSTMs

Nitish Srivastava, Elman Mansimov, Ruslan Salakhutdinov

OrganizationsUniversity of Toronto

Why you should read this

Presents an unsupervised LSTM encoder-decoder architecture that learns video representations by reconstructing and predicting sequence frames, substantially improving action recognition accuracy on benchmark datasets when labeled data is scarce.

We use multilayer Long Short Term Memory (LSTM) networks to learn representations of video sequences. Our model uses an encoder LSTM to map an input sequence into a fixed length representation. This representation is decoded using single or multiple decoder LSTMs to perform different tasks, such as reconstructing the input sequence, or predicting the future sequence. We experiment with two kinds of input sequences - patches of image pixels and high-level representations ("percepts") of video frames extracted using a pretrained convolutional net. We explore different design choices such as whether the decoder LSTMs should condition on the generated output. We analyze the outputs of the model qualitatively to see how well the model can extrapolate the learned video representation into the future and into the past. We try to visualize and interpret the learned features. We stress test the model by running it on longer time scales and on out-of-domain data. We further evaluate the representations by finetuning them for a supervised learning problem - human action recognition on the UCF-101 and HMDB-51 datasets. We show that the representations help improve classification accuracy, especially when there are only a few training examples. Even models pretrained on unrelated datasets (300 hours of YouTube videos) can help action recognition performance.

Added

2026-09-14

Generating Videos with Scene Dynamics

Generating Videos with Scene Dynamics

Carl Vondrick, Hamed Pirsiavash, Antonio Torralba

OrganizationsMassachusetts Institute of TechnologyUniversity of Maryland, Baltimore County

Why you should read this

Pioneers the use of Generative Adversarial Networks for video synthesis by introducing a two-stream spatio-temporal architecture that separates static backgrounds from moving foregrounds.

We capitalize on large amounts of unlabeled video in order to learn a model of scene dynamics for both video recognition tasks (e.g. action classification) and video generation tasks (e.g. future prediction). We propose a generative adversarial network for video with a spatio-temporal convolutional architecture that untangles the scene's foreground from the background. Experiments suggest this model can generate tiny videos up to a second at full frame rate better than simple baselines, and we show its utility at predicting plausible futures of static images. Moreover, experiments and visualizations show the model internally learns useful features for recognizing actions with minimal supervision, suggesting scene dynamics are a promising signal for representation learning. We believe generative video models can impact many applications in video understanding and simulation.

Added

2026-03-11

ImageBind One Embedding Space to Bind Them All

ImageBind One Embedding Space to Bind Them All

Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, Ishan Misra

OrganizationsMeta

Why you should read this

Demonstrates that aligning various modalities (audio, depth, thermal) to images automatically aligns them to each other, creating a universal embedding space.

We present ImageBind, an approach to learn a joint embedding across six different modalities - images, text, audio, depth, thermal, and IMU data. We show that all combinations of paired data are not necessary to train such a joint embedding, and only image-paired data is sufficient to bind the modalities together. ImageBind can leverage recent large scale vision-language models, and extends their zero-shot capabilities to new modalities just by using their natural pairing with images. It enables novel emergent applications 'out-of-the-box' including cross-modal retrieval, composing modalities with arithmetic, cross-modal detection and generation. The emergent capabilities improve with the strength of the image encoder and we set a new state-of-the-art on emergent zero-shot recognition tasks across modalities, outperforming specialist supervised models. Finally, we show strong few-shot recognition results outperforming prior work, and that ImageBind serves as a new way to evaluate vision models for visual and non-visual tasks.

Added

2026-01-28