Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Ruohong ZhangLiangke GuiZhiqing SunYihao FengKeyang XuYuanhan ZhangDi FuChunyuan LiAlexander G. HauptmannYonatan Bisk

article2025NAACL164 citations

Proposes a cost-effective framework that uses detailed video captions as text proxies for language model reward scoring, enabling direct preference optimization to improve open-ended video question answering performance without relying on expensive multimodal reward models.

Listen

Aligning video large multimodal models to follow human instructions accurately and minimize factual errors remains a critical bottleneck in artificial intelligence. While preference optimization techniques such as direct preference optimization have proven effective for text-only systems, applying them to video understanding is constrained by high costs and data scarcity. Gathering human feedback on videos is prohibitively expensive, and using advanced vision-language models like GPT-4V to score video frames is computationally slow, cost-heavy, and difficult to scale.

The article demonstrates an automated, cost-effective preference optimization framework for video models. The core objective is to evaluate whether detailed text captions can serve as an effective proxy for video content, enabling standard language models to generate reliable reward feedback to train video multimodal models using direct preference optimization.

To accomplish this, the authors created a large-scale dataset, ShareGPTVideo, containing 900,000 detailed video captions generated by prompting GPT-4V with sampled video frames across diverse public video datasets. From this, they produced 900,000 instruction-following question-answer pairs for supervised fine-tuning. For preference optimization, the fine-tuned model generated multiple candidate answers for given questions, and a text language model evaluated these against the detailed captions to assign numerical reward scores and explanations. The resulting 17,000 preference pairs were used to train a model named LLaVA-Hound-DPO, and the validity of using text captions in place of full video frames was evaluated across multiple standard benchmarks.

The findings show that text-based language model rewards align closely with direct vision model evaluations, achieving over 70% preference agreement with GPT-4V frame-based assessments and maintaining scores within one standard deviation in more than 75% of cases. Training with direct preference optimization using these language rewards improved average question answering accuracy to 70.75% across standard benchmarks, an 8.1% improvement over the supervised baseline of 62.65%, while also outperforming prior reinforcement learning methods. Furthermore, direct answer generation from the optimized model consistently outperformed test-time re-ranking of multiple candidates. On an economic level, generating on-policy preference data with this framework cost under 20,comparedto20, compared to 3,000 for equivalent human-annotated data.

These results demonstrate that detailed text representations can bypass the expensive computational bottlenecks of multi-frame video scoring without sacrificing evaluation accuracy. This offers an accessible, high-efficiency path for organizations to reduce hallucinations and improve factual correctness in video AI applications. Additionally, the study established that while benchmark evaluation scores vary significantly across underlying language model versions, relative model rankings remain consistent.

Organizations developing video multimodal systems should adopt caption-proxy reward mechanisms and preference optimization pipelines to improve model alignment at low cost, while ensuring that the visual projector remains frozen during preference training to prevent performance loss. Teams should also clearly document specific evaluator model versions to maintain reproducible benchmarks. Next steps should include expanding training to multiple-choice formats and refining captioning techniques to capture dynamic scene transitions better.

Confidence in these findings is moderate to high based on consistent improvements across multiple in-domain and out-of-domain benchmarks. However, leaders should note key limitations: the distilled captions were found by human auditors to have an accuracy between 80% and 90%, and the evaluation framework relies primarily on automated metrics rather than human corrections.

Cover for Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Abstract

Preference modeling techniques, such as direct preference optimization (DPO), has shown effective in enhancing the generalization abilities of large language model (LLM). However, in tasks involving video instruction-following, providing informative feedback, especially for open-ended conversations, remains a significant challenge. While previous studies have explored using large multimodal models (LMMs) as reward models for guiding preference modeling, their ability to accurately assess the quality of generated responses and their alignment with video content has not been conclusively demonstrated. This paper introduces a novel framework that utilizes detailed video captions as a proxy of video content, enabling language models to incorporate this information as supporting evidence for scoring video Question Answering (QA) predictions. Our approach demonstrates robust alignment with OpenAI GPT-4V model’s reward mechanism, which directly takes video frames as input. Furthermore, we show that applying our reward mechanism to DPO algorithm significantly improves model performance on open-ended video QA tasks.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Large Multi-Modal Models
  • 2.2 Video-text Datasets
  • 2.3 Preference Modeling for LMMs
  • 3 Method
  • 3.1 Prompting GPT-4V Model for Detailed Video Caption Distillation
  • 3.2 SFT with Generated Video Instruction Data from Detailed Caption
  • 3.3 DPO with Language Model Reward
  • 4 Assessment of Evaluator with GPT-4V Caption as Video Content
  • 5 Experiments
  • 5.1 Benchmark Evaluation
  • 5.2 Open-ended QA Analysis
  • 6 Conclusion
  • 7 Limitations
  • References
  • A Effect of ChatGPT Version on Official Benchmark Evaluation
  • B Evaluation of Captioning Ability from pre-training
  • C GPT-4V Caption Distillation
  • D Human Annotated Examples of Distilled Captions
  • E Video QA Dataset Demonstration
  • F Additional DPO Results
  • G Prompts for GPT-4V and ChatGPT Queries

Knowls

  1. Knowl 1 — Caption-grounded language reward for video preference optimization

    model/method

    The paper introduces a video-alignment method that replaces direct video-frame evaluation with evaluation grounded in a detailed textual caption of the video. A language model receives the detailed caption, the question, the ground-truth answer, and a candidate answer; it produces a natural-language judgment followed by a numerical quality score. The caption supplies evidence about temporal actions, objects, attributes, spatial relations, and visible text, allowing the language model to assess factual consistency without repeatedly processing the video frames. The resulting scores are used to construct preference pairs for direct preference optimization (DPO) of a video large multimodal model.

  2. Knowl 2 — SHAREGPTVIDEO detailed-caption dataset

    data/table

    The paper constructs SHAREGPTVIDEO, a 900,000-video detailed-caption dataset from 400,000 WebVid videos, 450,000 VIDAL videos, and 50,000 ActivityNet videos. Because GPT-4V is queried with images rather than videos, ten frames are sampled uniformly from each video and concatenated into a frame sequence. GPT-4V is prompted to produce a video-level description covering temporal dynamics, actions, object attributes, spatial relationships, counts, aesthetic properties, and text appearing in the video. The resulting captions are substantially more detailed than short video labels and can include world knowledge, such as recognizing a spatula as resembling a Star Wars Stormtrooper helmet. The released resource contains 900,000 captions and is used for video-caption pre-training, instruction generation, and caption-grounded reward modeling.

  3. Knowl 3 — Caption-derived video instruction data

    model/method

    The paper generates video instruction-following data from the detailed captions rather than collecting all questions and answers manually. It randomly samples 300,000 video captions and prompts ChatGPT to create three caption-grounded question-answer pairs for each caption, producing 900,000 instruction pairs. The questions are required to be diverse, answerable from the caption, and concerned with objects, counts, actions, locations, and relations. The full 900,000-pair corpus is released, while a random subset of 240,000 video instruction pairs is combined with 600,000 image-instruction examples for supervised fine-tuning. Conditioning question-answer generation on the detailed caption is intended to keep the instruction data factually consistent with the represented video.

  4. Knowl 4 — Factually enhanced preference-data construction

    algorithm

    The preference-data procedure takes detailed video captions, video questions, ground-truth answers, and a supervised-fine-tuned video LMM as input, and returns preference pairs for DPO.

    1. Randomly select 20,000 video instruction pairs.
    2. For each pair, sample six answers from the supervised-fine-tuned model at temperature 1.01.0, producing 120,000 candidate question-answer instances.
    3. Ask ChatGPT to evaluate each candidate using the detailed caption and ground-truth answer. ChatGPT produces an explanation and an integer reward from 11 to 55, where higher scores indicate better relevance, accuracy, clarity, and completeness.
    4. For each video-question pair, randomly choose one response with score at least 33 as the preferred response ywy_w and one response with score below 33 as the dispreferred response yly_l.
    5. Discard a pair when all six candidate responses have scores at least 33 or all six have scores below 33.

    This process yields approximately 17,000 preference instances of the form (V,x,yw,yl)(V,x,y_w,y_l), where VV is a video and xx is its question. The caption-based reward-generation cost is reported as less than 20atapriceof20 at a price of 1.5permilliontokens;thepapercontraststhiswiththereportedper million tokens; the paper contrasts this with the reported3,000 cost of collecting 10,000 human preference examples and the higher $30-per-million-token cost of direct GPT-4V reward labeling.

  5. Knowl 5 — DPO objective for caption-derived video preferences

    equation

    For the preference dataset DDPO={(V,x,yw,yl)}D_{\mathrm{DPO}}=\{(V,x,y_w,y_l)\}, the paper optimizes the video policy πθ\pi_\theta with the DPO objective

    LDPO(πθ;πref)=−E(V,x,yw,yl)∼DDPO[log⁡σ(βlog⁡πθ(yw∣x,V)πref(yw∣x,V)−βlog⁡πθ(yl∣x,V)πref(yl∣x,V))].\mathcal{L}_{\mathrm{DPO}}(\pi_\theta;\pi_{\mathrm{ref}})=-\mathbb{E}_{(V,x,y_w,y_l)\sim D_{\mathrm{DPO}}}\left[\log\sigma\left(\beta\log\frac{\pi_\theta(y_w\mid x,V)}{\pi_{\mathrm{ref}}(y_w\mid x,V)}-\beta\log\frac{\pi_\theta(y_l\mid x,V)}{\pi_{\mathrm{ref}}(y_l\mid x,V)}\right)\right].

    Here, VV is a video, xx is a question, ywy_w is the preferred answer, yly_l is the dispreferred answer, πθ\pi_\theta is the policy being optimized, πref\pi_{\mathrm{ref}} is the reference policy, and σ(z)=1/(1+e−z)\sigma(z)=1/(1+e^{-z}) is the logistic function. Both policies are initialized from the supervised-fine-tuned model, and the inverse-temperature parameter is fixed at β=0.1\beta=0.1. The objective increases the relative likelihood of caption-supported preferred answers over dispreferred answers while regularizing the policy against the reference model.

  6. Knowl 6 — Agreement between caption-based and frame-based evaluators

    empirical result

    The paper evaluates whether ChatGPT using a detailed caption as a proxy for video content agrees with GPT-4V evaluating the actual video frames. It samples 200 videos from each of WebVid, VIDAL, and ActivityNet; each video has one question and two supervised-fine-tuned-model predictions, with one prediction preferred by the caption-based evaluator. After API filtering, both predictions receive GPT-4V scores for 196 WebVid, 151 VIDAL, and 143 ActivityNet videos.

    The caption-based and frame-based scores have a Pearson correlation of 0.470.47 with p<0.01p<0.01. Their mean scores are 2.9 and 3.5, respectively, indicating that GPT-4V tends to assign somewhat more positive scores. More than 75% of caption-based scores lie within one standard deviation of the GPT-4V scores, with standard deviation σ=1.31\sigma=1.31. Excluding ties, preference agreement is 71.2% on WebVid, 73.9% on VIDAL, and 73.7% on ActivityNet. These results support, with the paper's stated caution, the use of detailed captions as a lower-cost proxy for frame-based video evaluation.

  7. Knowl 7 — Automated open-ended video QA development benchmark

    data/table

    The paper creates an automated development benchmark for long-form video question answering by selecting 2,000 videos from each of WebVid, VIDAL, ActivityNet, MSRVTT, MSVD, TGIF, and Something-Something V2. ChatGPT generates three detailed question-answer pairs per video from the corresponding caption. WebVid, VIDAL, and ActivityNet are treated as in-domain because their captions and instructions participate in training; MSRVTT, MSVD, TGIF, and Something-Something V2 are treated as out-of-domain. ChatGPT scores model answers on a 11--55 scale, and a score of at least 33 counts as correct.

    The benchmark compares the supervised-fine-tuned model with the DPO model and selected ablations under GPT-3.5-turbo-0301 evaluation.

    Method ActivityNet-QA VIDAL-QA WebVid-QA
    Acc. Score Acc. Score Acc. Score
    Video-LLaVA 41.35 2.38 34.30 2.24 42.47 2.39
    LLAVA-HOUND-SFT 66.62 3.05 60.50 2.88 71.07 3.17
    LLAVA-HOUND-DPO 76.62 3.18 70.06 3.04 79.82 3.29
    LLAVA-HOUND-PT + Image Inst. 69.31 3.09 60.57 2.85 68.03 3.02
    LLAVA-HOUND-PT + VChat 67.34 3.02 62.33 2.89 68.98 3.00
    LLAVA-HOUND-DPO + training MLP 71.89 3.10 65.57 2.95 75.37 3.21
    LLAVA-HOUND-SFT + Self-play 64.11 2.85 56.28 2.68 67.89 2.95
    LLAVA-HOUND-DPO w/ lr3e-7 71.13 3.08 64.90 2.92 73.25 3.17
    Method MSVD-QA MSRVTT-QA TGIF-QA SSV2-QA
    Acc. Score Acc. Score Acc. Score Acc. Score
    Video-LLaVA 39.46 2.37 30.78 2.15 32.95 2.18 24.31 1.90
    LLAVA-HOUND-SFT 66.99 3.09 57.82 2.85 66.13 3.07 35.07 2.23
    LLAVA-HOUND-DPO 73.64 3.12 68.29 2.98 74.00 3.12 48.89 2.53
    LLAVA-HOUND-PT + Image Inst. 65.19 2.96 48.66 2.52 53.83 2.62 29.60 2.04

    DPO improves every reported in-domain and out-of-domain accuracy over the supervised-fine-tuned model. The paper reports that self-play preference construction reduces accuracy by about 3 percentage points relative to supervised fine-tuning, unfreezing the MLP projector during DPO decreases performance, and reducing the DPO learning rate from 5×10−75\times10^{-7} to 3×10−73\times10^{-7} also decreases performance. Accuracy generally peaks at approximately 2.5 DPO training epochs, or about 350 steps.

  8. Knowl 8 — Training recipe for LLAVA-HOUND-DPO

    experimental setup

    The paper uses Video-LLaVA as the 7-billion-parameter backbone and trains three successive versions. In caption pre-training, LLAVA-HOUND-PT uses 650,000 image-caption examples from ALLaVA and 900,000 detailed video captions from SHAREGPTVIDEO; the visual encoder is frozen while the MLP projector and language model are fine-tuned with learning rate 2×10−52\times10^{-5} and batch size 128. In supervised fine-tuning, LLAVA-HOUND-SFT uses 600,000 image-instruction examples and 240,000 generated video-instruction examples with learning rate 5×10−65\times10^{-6} and batch size 128. In DPO training, LLAVA-HOUND-DPO uses the approximately 17,000 caption-derived preference pairs, trains the full model for three epochs with learning rate 5×10−75\times10^{-7} and batch size 128, and initializes both policy and reference models from the SFT checkpoint. All experiments use eight A100 GPUs.

  9. Knowl 9 — Performance on established video QA benchmarks

    data/table

    The paper evaluates zero-shot video QA on MSVD-QA, MSRVTT-QA, TGIF-QA, and Next-QA. Accuracy and answer-quality scores are assigned by ChatGPT; the first three datasets use gpt-3.5-turbo-0613, while Next-QA uses gpt-3.5-turbo-0611. LLAVA-HOUND-DPO substantially outperforms its SFT initialization and the reproduced baselines.

    Method MSVD-QA MSRVTT-QA TGIF-QA
    Acc. Score Acc. Score Acc. Score
    Video-ChatGPT 68.6 3.8 58.9 3.4 47.8 3.2
    Chat-UniVi 70.0 3.8 53.1 3.1 46.1 3.1
    Video-LLaVA 71.8 3.9 59.0 3.4 48.4 3.2
    LLaMA-VID-7B 72.6 3.9 58.7 3.4 49.2 3.3
    LLaMA-VID-13B 74.3 4.0 59.8 3.4 50.8 3.3
    VLM-RLAIF 76.4 4.0 63.0 3.4 - -
    LLAVA-HOUND-SFT 75.7 3.9 58.7 3.3 53.5 3.3
    LLAVA-HOUND-DPO 80.7 4.1 70.2 3.7 61.4 3.5
    Method Next-QA Acc. Next-QA Score
    Video-ChatGPT 45.23 2.09
    LLaMA-VID-7B 49.43 3.24
    Chat-UniVi 47.62 3.14
    Video-LLaVA 48.97 3.25
    LLAVA-HOUND-SFT 60.60 3.51
    LLAVA-HOUND-DPO 74.27 3.74

    Across MSVD-QA, MSRVTT-QA, and TGIF-QA, the DPO model has average accuracy 70.75%, compared with 62.65% for the SFT model, an 8.1% improvement. LLAVA-HOUND-DPO also reaches 74.27% accuracy on Next-QA, compared with 60.60% for LLAVA-HOUND-SFT. The paper notes that absolute scores change substantially with the ChatGPT evaluator version, although the relative ranking is comparatively stable.

  10. Knowl 10 — Reliability limits of captions and automated evaluation

    limitation

    Human annotation exposes residual errors in the GPT-4V-generated captions used as evidence. On 75 videos, with 25 videos from each source domain, annotators found 21 inaccurate facts across 14 videos, corresponding to reported caption accuracy of 81%, and 12 missing items across 8 videos, corresponding to reported coverage accuracy of 89%. Observed errors include incorrect text recognition, missed scoreboard transitions, incorrect motion direction, and invented objects such as a nonexistent stool or ladder.

    The proposed development benchmark is fully automated and applies no human correction to captions, questions, or answers. The paper therefore recommends using it for model development and hyperparameter tuning rather than treating its scores as definitive. The study also omits multiple-choice video benchmarks because it focuses on open-ended QA and does not retrain the model on such data. Finally, different ChatGPT evaluator versions produce materially different absolute metrics, so evaluation results require the evaluator version to be reported explicitly.

Coverage note — No substantial contributed component was deliberately omitted; detailed prompt templates and illustrative caption examples were compressed because their operational content is represented by the dataset, instruction-generation, reward, and evaluation knowls.

References

  1. 1.Daechul Ahn, Yura Choi, Youngjae Yu, Dongyeop Kang, and Jonghyun Choi. 2024. Tuning large multimodal models for videos using reinforcement learning from ai feedback. arXiv preprint arXiv:2402.03746.
  2. 2.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966.
  3. 3.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073.
  4. 4.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. 2021a. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1728–1738.
  5. 5.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. 2021b. Frozen in time: A joint video and image encoder for end-to-end retrieval. In ICCV.
  6. 6.David Chen and William B Dolan. 2011. Collecting highly parallel data for paraphrase evaluation. In Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, pages 190–200.
  7. 7.Guiming Hardy Chen, Shunian Chen, Ruifei Zhang, Junying Chen, Xiangbo Wu, Zhiyi Zhang, Zhihong Chen, Jianquan Li, Xiang Wan, and Benyou Wang. 2024a. Allava: Harnessing gpt4v-synthesized data for a lite vision-language model. arXiv preprint arXiv:2402.11684.
  8. 8.Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao. 2023. Shikra: Unleashing multimodal llm’s referential dialogue magic. arXiv preprint arXiv:2306.15195.
  9. 9.Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, et al. 2024b. Sharegpt4video: Improving video understanding and generation with better captions. arXiv preprint arXiv:2406.04325.
  10. 10.Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al. 2024c. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint arXiv:2412.05271.
  11. 11.Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu. 2024d. Self-play fine-tuning converts weak language models to strong language models. arXiv preprint arXiv:2401.01335.
  12. 12.Yihe Deng, Pan Lu, Fan Yin, Ziniu Hu, Sheng Shen, James Zou, Kai-Wei Chang, and Wei Wang. 2024. Enhancing large vision language models with self-training on image comprehension. arXiv preprint arXiv:2405.19716.
  13. 13.Bernard Ghanem Fabian Caba Heilbron, Victor Escorcia and Juan Carlos Niebles. 2015. Activitynet: A large-scale video benchmark for human activity understanding. In CVPR.
  14. 14.Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, Peixian Chen, Yanwei Li, Shaohui Lin, Sirui Zhao, Ke Li, Tong Xu, Xiawu Zheng, Enhong Chen, Rongrong Ji, and Xing Sun. 2024. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. Preprint, arXiv:2405.21075.
  15. 15.Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. 2017. The" something something" video database for learning and evaluating visual common sense. In ICCV.
  16. 16.Madeleine Grunde-McLaughlin, Ranjay Krishna, and Maneesh Agrawala. 2021. Agqa: A benchmark for compositional spatio-temporal reasoning. In CVPR.
  17. 17.Anisha Gunjal, Jihan Yin, and Erhan Bas. 2023. Detecting and preventing hallucinations in large vision language models. arXiv preprint arXiv:2308.06394.
  18. 18.Mingfei Han, Linjie Yang, Xiaojun Chang, and Heng Wang. 2023. Shot2story20k: A new benchmark for comprehensive understanding of multi-shot videos. arXiv preprint arXiv:2312.10300.
  19. 19.Arian Hosseini, Xingdi Yuan, Nikolay Malkin, Aaron Courville, Alessandro Sordoni, and Rishabh Agarwal. 2024. V-star: Training verifiers for self-taught reasoners. arXiv preprint arXiv:2402.06457.
  20. 20.Hamish Ivison, Yizhong Wang, Valentina Pyatkin, Nathan Lambert, Matthew Peters, Pradeep Dasigi, Joel Jang, David Wadden, Noah A Smith, Iz Beltagy, et al. 2023. Camels in a changing climate: Enhancing lm adaptation with tulu 2. arXiv preprint arXiv:2311.10702.
  21. 21.Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017. Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In CVPR.
  22. 22.Peng Jin, Ryuichi Takanobu, Caiwan Zhang, Xiaochun Cao, and Li Yuan. 2023. Chat-univi: Unified visual representation empowers large language models with image and video understanding. arXiv preprint arXiv:2311.08046.
  23. 23.Yang Jin, Zhicheng Sun, Kun Xu, Liwei Chen, Hao Jiang, Quzhe Huang, Chengru Song, Yuliang Liu, Di Zhang, Yang Song, et al. 2024. Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization. arXiv preprint arXiv:2402.03161.
  24. 24.Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Burbune, and Abhinav Rastogi. 2023. Rlaif: Scaling reinforcement learning from human feedback with ai feedback. arXiv preprint arXiv:2309.00267.
  25. 25.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023a. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597.
  26. 26.KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenh ai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. 2023b. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355.
  27. 27.Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. 2023c. Mvbench: A comprehensive multi-modal video understanding benchmark. arXiv preprint arXiv:2311.17005.
  28. 28.Lei Li, Zhihui Xie, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen, Yazheng Yang, Benyou Wang, and Lingpeng Kong. 2023d. Silkie: Preference distillation for large visual language models. arXiv preprint arXiv:2312.10665.
  29. 29.Yanwei Li, Chengyao Wang, and Jiaya Jia. 2023e. Llama-vid: An image is worth 2 tokens in large language models. arXiv preprint arXiv:2311.17043.
  30. 30.Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. 2023a. Video-llava: Learning united visual representation by alignment before projection. Preprint, arXiv:2311.10122.
  31. 31.Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan. 2023b. Video-llava: Learning united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122.
  32. 32.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023a. Improved baselines with visual instruction tuning. arXiv preprint arXiv:2310.03744.
  33. 33.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023b. Visual instruction tuning. arXiv preprint arXiv:2304.08485.
  34. 34.Ruyang Liu, Chen Li, Yixiao Ge, Ying Shan, Thomas H Li, and Ge Li. 2023c. One for all: Video conversation is feasible without video instruction tuning. arXiv preprint arXiv:2309.15785.
  35. 35.Ruipu Luo, Ziwang Zhao, Min Yang, Junwei Dong, Minghui Qiu, Pengcheng Lu, Tao Wang, and Zhongyu Wei. 2023. Valley: Video assistant with large language model enhanced ability. arXiv preprint arXiv:2306.07207.
  36. 36.Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. 2023. Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424.
  37. 37.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  38. 38.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36.
  39. 39.Nina Shvetsova, Anna Kukleva, Xudong Hong, Christian Rupprecht, Bernt Schiele, and Hilde Kuehne. 2023. Howtocaption: Prompting llms to transform video annotations at scale. Preprint, arXiv:2310.04900.
  40. 40.Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, et al. 2023a. Aligning large multimodal models with factually augmented rlhf. arXiv preprint arXiv:2309.14525.
  41. 41.Zhiqing Sun, Yikang Shen, Hongxin Zhang, Qinhong Zhou, Zhenfang Chen, David Cox, Yiming Yang, and Chuang Gan. 2023b. Salmon: Self-alignment with principle-following reward models. arXiv preprint arXiv:2310.05910.
  42. 42.Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. 2024. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191.
  43. 43.Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, et al. 2023. Internvid: A large-scale video-text dataset for multimodal understanding and generation. arXiv preprint arXiv:2307.06942.
  44. 44.Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B Tenenbaum, and Chuang Gan. 2024. Star: A benchmark for situated reasoning in real-world videos. arXiv preprint arXiv:2405.09711.
  45. 45.Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua. 2021. Next-qa: Next phase of question-answering to explaining temporal actions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9777–9786.
  46. 46.Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. 2017. Video question answering via gradually refined attention over appearance and motion. In Proceedings of the 25th ACM international conference on Multimedia, pages 1645–1653.
  47. 47.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5288–5296.
  48. 48.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. 2022. Zero-shot video question answering via frozen bidirectional language models. NeurIPS.
  49. 49.Dongjie Yang, Suyuan Huang, Chengqiang Lu, Xiaodong Han, Haoxin Zhang, Yan Gao, Yao Hu, and Hai Zhao. 2024. Vript: A video is worth thousands of words. arXiv preprint arXiv:2406.06040.
  50. 50.Tianyu Yu, Haoye Zhang, Yuan Yao, Yunkai Dang, Da Chen, Xiaoman Lu, Ganqu Cui, Taiwen He, Zhiyuan Liu, Tat-Seng Chua, et al. 2024. Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness. arXiv preprint arXiv:2405.17220.
  51. 51.Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. 2019. Activitynet-qa: A dataset for understanding complex web videos via question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 9127–9134.
  52. 52.Hang Zhang, Xin Li, and Lidong Bing. 2023a. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858.
  53. 53.Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao. 2023b. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199.
  54. 54.Yuanhan Zhang, Bo Li, haotian Liu, Yong jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and Chunyuan Li. 2024. Llava-next: A strong zero-shot video understanding model.
  55. 55.Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, et al. 2023. Sotopia: Interactive evaluation for social intelligence in language agents. arXiv preprint arXiv:2310.11667.
  56. 56.Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, HongFa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, et al. 2023. Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment. arXiv preprint arXiv:2310.01852.

Citation

MLA
Zhang, R., et al. “Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 694–717, https://doi.org/10.18653/v1/2025.naacl-long.30.
APA
Zhang, R., Gui, L., Sun, Z., Feng, Y., Xu, K., Zhang, Y., Fu, D., Li, C., Hauptmann, A. G., Bisk, Y., & Yang, Y. (2025). Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 694–717. https://doi.org/10.18653/v1/2025.naacl-long.30
Chicago
Zhang, R., L. Gui, Z. Sun, et al. 2025. “Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 694–717. https://doi.org/10.18653/v1/2025.naacl-long.30.
Harvard
Zhang, R. et al. (2025) “Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward”, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 694–717. Available at: https://doi.org/10.18653/v1/2025.naacl-long.30.
Vancouver
1. Zhang R, Gui L, Sun Z, et al (2025) Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 694–717

BibTeX

@inproceedings{Zhang_2025, title={Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward}, url={http://dx.doi.org/10.18653/v1/2025.naacl-long.30}, DOI={10.18653/v1/2025.naacl-long.30}, booktitle={Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)}, publisher={Association for Computational Linguistics}, author={Zhang, Ruohong and Gui, Liangke and Sun, Zhiqing and Feng, Yihao and Xu, Keyang and Zhang, Yuanhan and Fu, Di and Li, Chunyuan and Hauptmann, Alexander G and Bisk, Yonatan and Yang, Yiming}, year={2025}, pages={694–717} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/