Tuning Computer Vision Models With Task Rewards
André Susano PintoAlexander KolesnikovYuge ShiLucas BeyerXiaohua Zhai
Demonstrates that fine-tuning pretrained vision models with REINFORCE using task-specific rewards directly optimizes non-differentiable evaluation metrics across diverse tasks including object detection, panoptic segmentation, and image colorization without task-specific architectural modifications.
Deploying standard computer vision systems often suffers from misalignment between how models are trained and how they are evaluated in real-world use. While models are typically trained to imitate ground-truth data, actual applications require optimizing complex, non-differentiable goals such as comprehensive object detection or visually appealing image styling. To address this mismatch, engineering teams traditionally rely on task-specific heuristics, customized post-processing steps, and tailored network architectures that add considerable development complexity.
The article demonstrates that generic sequence-to-sequence vision models can be effectively aligned with target operational objectives using a two-stage training workflow. It evaluates how initializing generic architectures with standard likelihood pretraining and subsequently fine-tuning them via policy gradient reinforcement learning directly optimizes target task metrics without requiring specialized architectural modifications.
The approach uses standard Vision Transformer encoder-decoder models across diverse visual tasks, including object detection, panoptic segmentation, image colorization, and captioning. Models are first pretrained to predict discrete sequence outputs from visual inputs using standard maximum likelihood estimation. Subsequently, the models are tuned using the REINFORCE algorithm against non-differentiable task rewards, using paired sample baselines to reduce training variance.
The evaluation produced several key findings across benchmark datasets. In object detection on the COCO dataset, task reward tuning improved mean average precision from 39.2% to 54.3% and average recall from 54.4% to 68.4%, outperforming specialized architectures without requiring complex post-processing. In panoptic segmentation, tuning improved the panoptic quality metric from 43.1% to 46.1%, effectively removing incoherent segmentation artifacts on small objects. For image colorization, optimizing a custom color reward increased vividness metrics from 0.46 to 0.97 and color diversity entropy from 1.03 to 1.84, resolving the muted tones common in standard models. In image captioning, consensus metrics improved significantly, rising from 120.0 to 134.5 on base models.
These results show that reinforcement learning provides a viable, general-purpose mechanism to align computer vision models with complex performance goals. By shifting the alignment burden from custom model engineering to reward definition, teams can lower architectural complexity and achieve superior task performance using standard model designs. The findings also indicate that while standard models internally generate high-quality candidate outputs, direct reward tuning is necessary to ensure the model reliably selects these optimal predictions at test time.
Engineering teams should adopt this two-stage framework when standard likelihood objectives fail to reflect real-world priorities, particularly for downstream operations like robotic perception or automated content generation. Organizations must carefully design and validate reward metrics before deployment to mitigate reward-hacking behaviors, such as models exploiting metric loopholes with incomplete text phrases or unnatural over-saturation. Future efforts should evaluate off-policy reinforcement learning methods to lower sampling compute overhead and explore multi-metric reward designs driven by human feedback.
- Paper: Self-Critical Sequence Training for Image Captioning, Steven J. Rennie et al. (2016). This foundational paper establishes self-critical sequence training with policy gradients to optimize non-differentiable sequence metrics, providing the core reinforcement learning baseline mechanism adapted by the source.
- Paper: Sequence Level Training with Recurrent Neural Networks, Marc'Aurelio Ranzato et al. (2015). This work introduces the two-stage training paradigm of initializing sequence models with maximum likelihood estimation before fine-tuning them via policy gradients to optimize non-differentiable task metrics.
- Paper: Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks, Jiasen Lu et al. (2022). This research develops unified sequence-to-sequence Transformer architectures that represent diverse vision tasks as discrete sequence generation, establishing the model formulation tuned by the source.
- Paper: Deformable DETR: Deformable Transformers for End-to-End Object Detection, Xizhou Zhu et al. (2021). This paper presents transformer-based end-to-end object detection, illustrating the specialized architecture and loss designs that the source aims to surpass using generic architectures fine-tuned with task rewards.
- Paper: SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training, Tianzhe Chu et al. (2025). This comparative study directly extends the post-training paradigm by analyzing how reinforcement learning fine-tuning fosters out-of-distribution generalization compared to supervised fine-tuning in vision and multimodal foundation models.
- Paper: Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment, Rui Yang et al. (2024). This work advances foundation model reward tuning by tackling multi-objective alignment and real-time dynamic preference conditioning beyond single-metric task rewards.
- Paper: Demystifying Reinforcement Learning Post-Training of Language Models, Donovan Clay et al. (2026). This paper provides a theoretical and empirical demystification of how reinforcement learning post-training interacts with base model likelihoods and reward signals in sequence models.
- Paper: Visual Planning: Let's Think Only with Images, Yi Xu et al. (2026). This work applies two-stage supervised initialization followed by task reward reinforcement learning to train vision models for sequential visual planning directly in pixel space.
- Paper: Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control, Carles Domingo-Enrich et al. (2025). This paper extends reward-based fine-tuning principles to continuous dynamical generative models like flow and diffusion architectures using stochastic optimal control.
