Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Video-LLaVA

Video-LLaVA is a multimodal artificial intelligence model designed to understand and reason over both static images and dynamic video content alongside natural language. Unlike conventional vision-language models that process images and videos through separate feature spaces, Video-LLaVA aligns and unifies visual representations into a shared space before projecting them into a foundational large language model. By training on mixed datasets of images and videos, the architecture allows both visual formats to mutually enhance each other, enabling unified visual question answering, dialogue, and reasoning across diverse visual-language tasks.

2 items

Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Ruohong Zhang, Liangke Gui, Zhiqing Sun, Yihao Feng, Keyang Xu, Yuanhan Zhang, Di Fu, Chunyuan Li, Alexander G. Hauptmann, Yonatan Bisk, Yiming Yang

OrganizationsByteDanceCarnegie Mellon UniversityColumbia UniversityNanyang Technological UniversityUniversity of Texas at Austin

Why you should read this

Proposes a cost-effective framework that uses detailed video captions as text proxies for language model reward scoring, enabling direct preference optimization to improve open-ended video question answering performance without relying on expensive multimodal reward models.

Preference modeling techniques, such as direct preference optimization (DPO), has shown effective in enhancing the generalization abilities of large language model (LLM). However, in tasks involving video instruction-following, providing informative feedback, especially for open-ended conversations, remains a significant challenge. While previous studies have explored using large multimodal models (LMMs) as reward models for guiding preference modeling, their ability to accurately assess the quality of generated responses and their alignment with video content has not been conclusively demonstrated. This paper introduces a novel framework that utilizes detailed video captions as a proxy of video content, enabling language models to incorporate this information as supporting evidence for scoring video Question Answering (QA) predictions. Our approach demonstrates robust alignment with OpenAI GPT-4V model’s reward mechanism, which directly takes video frames as input. Furthermore, we show that applying our reward mechanism to DPO algorithm significantly improves model performance on open-ended video QA tasks.

Added

2026-09-26

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, Li Yuan

OrganizationsPandaVilla Tech LimitedPeking UniversityPeng Cheng Laboratory

Why you should read this

Introduces Video-LLaVA, a vision-language model that aligns image and video features into a single representation before projecting them to a large language model, demonstrating that joint multimodal training outperforms specialized single-modality systems across major image and video benchmarks.

The Large Vision-Language Model (LVLM) has enhanced the performance of various downstream tasks in visual-language understanding. Most existing approaches encode images and videos into separate feature spaces, which are then fed as inputs to large language models. However, due to the lack of unified tokenization for images and videos, namely misalignment before projection, it becomes challenging for a Large Language Model (LLM) to learn multi-modal interactions from several poor projection layers. In this work, we unify visual representation into the language feature space to advance the foundational LLM towards a unified LVLM. As a result, we establish a simple but robust LVLM baseline, Video-LLaVA, which learns from a mixed dataset of images and videos, mutually enhancing each other. Video-LLaVA achieves superior performances on a broad range of 9 image benchmarks across 5 image question-answering datasets and 4 image benchmark toolkits. Additionally, our Video-LLaVA also outperforms Video-ChatGPT by 5.8%, 9.9%, 18.6%, and 10.1% on MSRVTT, MSVD, TGIF, and ActivityNet, respectively. Notably, extensive experiments demonstrate that Video-LLaVA mutually benefits images and videos within a unified visual representation, outperforming models designed specifically for images or videos. We aim for this work to provide modest insights into the multi-modal inputs for the LLM. Code address: \href{this https URL}

Added

2026-09-24