Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
Ruohong ZhangLiangke GuiZhiqing SunYihao FengKeyang XuYuanhan ZhangDi FuChunyuan LiAlexander G. HauptmannYonatan Bisk
Proposes a cost-effective framework that uses detailed video captions as text proxies for language model reward scoring, enabling direct preference optimization to improve open-ended video question answering performance without relying on expensive multimodal reward models.
Aligning video large multimodal models to follow human instructions accurately and minimize factual errors remains a critical bottleneck in artificial intelligence. While preference optimization techniques such as direct preference optimization have proven effective for text-only systems, applying them to video understanding is constrained by high costs and data scarcity. Gathering human feedback on videos is prohibitively expensive, and using advanced vision-language models like GPT-4V to score video frames is computationally slow, cost-heavy, and difficult to scale.
The article demonstrates an automated, cost-effective preference optimization framework for video models. The core objective is to evaluate whether detailed text captions can serve as an effective proxy for video content, enabling standard language models to generate reliable reward feedback to train video multimodal models using direct preference optimization.
To accomplish this, the authors created a large-scale dataset, ShareGPTVideo, containing 900,000 detailed video captions generated by prompting GPT-4V with sampled video frames across diverse public video datasets. From this, they produced 900,000 instruction-following question-answer pairs for supervised fine-tuning. For preference optimization, the fine-tuned model generated multiple candidate answers for given questions, and a text language model evaluated these against the detailed captions to assign numerical reward scores and explanations. The resulting 17,000 preference pairs were used to train a model named LLaVA-Hound-DPO, and the validity of using text captions in place of full video frames was evaluated across multiple standard benchmarks.
The findings show that text-based language model rewards align closely with direct vision model evaluations, achieving over 70% preference agreement with GPT-4V frame-based assessments and maintaining scores within one standard deviation in more than 75% of cases. Training with direct preference optimization using these language rewards improved average question answering accuracy to 70.75% across standard benchmarks, an 8.1% improvement over the supervised baseline of 62.65%, while also outperforming prior reinforcement learning methods. Furthermore, direct answer generation from the optimized model consistently outperformed test-time re-ranking of multiple candidates. On an economic level, generating on-policy preference data with this framework cost under 3,000 for equivalent human-annotated data.
These results demonstrate that detailed text representations can bypass the expensive computational bottlenecks of multi-frame video scoring without sacrificing evaluation accuracy. This offers an accessible, high-efficiency path for organizations to reduce hallucinations and improve factual correctness in video AI applications. Additionally, the study established that while benchmark evaluation scores vary significantly across underlying language model versions, relative model rankings remain consistent.
Organizations developing video multimodal systems should adopt caption-proxy reward mechanisms and preference optimization pipelines to improve model alignment at low cost, while ensuring that the visual projector remains frozen during preference training to prevent performance loss. Teams should also clearly document specific evaluator model versions to maintain reproducible benchmarks. Next steps should include expanding training to multiple-choice formats and refining captioning techniques to capture dynamic scene transitions better.
Confidence in these findings is moderate to high based on consistent improvements across multiple in-domain and out-of-domain benchmarks. However, leaders should note key limitations: the distilled captions were found by human auditors to have an accuracy between 80% and 90%, and the evaluation framework relies primarily on automated metrics rather than human corrections.
- Paper: Direct Preference Optimization: Your Language Model is Secretly a Reward Model, Rafael Rafailov et al. (2023). This foundational work introduces Direct Preference Optimization (DPO), the core alignment algorithm that the source adapts to video large multimodal models using language model rewards.
- Paper: RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback, Harrison Lee et al. (2024). This paper establishes the validity of using AI-generated feedback and language models as reward signals for alignment, which directly underpins the source's use of language models to evaluate multimodal predictions.
- Paper: Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models, Muhammad Maaz et al. (2023). This research pioneers instruction-tuning and conversational evaluation benchmarks for video-based large language models, setting the stage for open-ended video QA tasks addressed in the source.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). This paper establishes effective multi-image and video representation strategies for multimodal models, providing critical architectural context for modern video large multimodal models.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). This work establishes how large language models can act as reliable, human-aligned evaluators using chain-of-thought and detailed scoring criteria, motivating the source's proxy-scoring framework.
- Paper: From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models, Tarun Raheja et al. (2026). This work provides a theoretical unification of direct alignment algorithms including DPO, offering deeper insights into the mathematical and design axes governing preference optimization.
- Paper: M-LLM Based Video Frame Selection for Efficient Video Understanding, Kai Hu 0010 et al. (2025). This paper complements video QA preference alignment by exploring adaptive frame selection and caption-based reasoning to efficiently feed relevant visual evidence into multimodal models.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). This technical report shows how advanced video multimodal architectures scale up instruction-tuning and post-training reinforcement learning across comprehensive video reasoning tasks.
