Built independently by an author, for readers. Read the story and support ChapterPal

keyword

video language models

Video language models are artificial intelligence systems designed to jointly process, interpret, and generate both video content and natural language text. These multimodal models typically integrate visual encoders or video tokenizers with large language model architectures, allowing them to comprehend temporal dynamics, spatial relationships, and motion across continuous sequences of visual frames alongside textual context. By learning cross-modal representations that align visual and linguistic data across time, video language models support diverse applications, including video question answering, video captioning, spatio-temporal reasoning, long-form content summarization, and text-conditioned video generation.

4 items

VideoPoet: A Large Language Model for Zero-Shot Video Generation

VideoPoet: A Large Language Model for Zero-Shot Video Generation

Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng, Joshua V. Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso Martinez, David Minnen, Mikhail Sirotenko, Kihyuk Sohn, Xuan Yang, Hartwig Adam, Ming-Hsuan Yang, Irfan Essa, Huisheng Wang, David A. Ross, Bryan Seybold, Lu Jiang

OrganizationsCarnegie Mellon UniversityGoogle

Why you should read this

Demonstrates that a unified decoder-only language model trained on discrete multimodal tokens can match or outperform diffusion approaches across diverse video generation and editing tasks without task-specific architectural changes.

We present VideoPoet, a model for synthesizing high-quality videos from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs – including images, videos, text, and audio. The training protocol follows that of Large Language Models (LLMs), consisting of two stages: pretraining and task-specific adaptation. During pretraining, VideoPoet incorporates a mixture of multimodal generative objectives within an autoregressive Transformer framework. The pretrained LLM serves as a foundation that is adapted to a range of video generation tasks. We present results demonstrating the model’s state-of-the-art capabilities in zero-shot video generation, specifically highlighting the generation of high-fidelity motions. Project page: https://sites.research.google/videopoet/.

Added

2026-09-28

HourVideo: 1-Hour Video-Language Understanding

HourVideo: 1-Hour Video-Language Understanding

Keshigeyan Chandrasegaran, Agrim Gupta, Lea M. Hadzic, Taran Kota, Jimming He, Cristóbal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, Li Fei-Fei

OrganizationsStanford University

Why you should read this

Presents a benchmark of nearly thirteen thousand questions across five hundred hour-long egocentric videos to expose the critical performance gap between state-of-the-art multimodal models and human-level long-form video comprehension.

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, temporal, predictive, causal, counterfactual), and navigation (room-to-room, object retrieval) tasks. HourVideo includes 500 manually curated egocentric videos from the Ego4D dataset, spanning durations of 20 to 120 minutes, and features 12,976 high-quality, five-way multiple-choice questions. Benchmarking results reveal that multimodal models, including GPT-4 and LLaVA-NeXT, achieve marginal improvements over random chance. In stark contrast, human experts significantly outperform the state-of-the-art long-context multimodal model, Gemini Pro 1.5 (85.0% vs. 37.3%), highlighting a substantial gap in multimodal capabilities. Our benchmark, evaluation toolkit, prompts, and documentation are available at hourvideo.stanford.edu.

Added

2026-09-26