Tuning Computer Vision Models With Task Rewards

André Susano PintoAlexander KolesnikovYuge ShiLucas BeyerXiaohua Zhai

article2023ICML53 citations

Demonstrates that fine-tuning pretrained vision models with REINFORCE using task-specific rewards directly optimizes non-differentiable evaluation metrics across diverse tasks including object detection, panoptic segmentation, and image colorization without task-specific architectural modifications.

Listen

Deploying standard computer vision systems often suffers from misalignment between how models are trained and how they are evaluated in real-world use. While models are typically trained to imitate ground-truth data, actual applications require optimizing complex, non-differentiable goals such as comprehensive object detection or visually appealing image styling. To address this mismatch, engineering teams traditionally rely on task-specific heuristics, customized post-processing steps, and tailored network architectures that add considerable development complexity.

The article demonstrates that generic sequence-to-sequence vision models can be effectively aligned with target operational objectives using a two-stage training workflow. It evaluates how initializing generic architectures with standard likelihood pretraining and subsequently fine-tuning them via policy gradient reinforcement learning directly optimizes target task metrics without requiring specialized architectural modifications.

The approach uses standard Vision Transformer encoder-decoder models across diverse visual tasks, including object detection, panoptic segmentation, image colorization, and captioning. Models are first pretrained to predict discrete sequence outputs from visual inputs using standard maximum likelihood estimation. Subsequently, the models are tuned using the REINFORCE algorithm against non-differentiable task rewards, using paired sample baselines to reduce training variance.

The evaluation produced several key findings across benchmark datasets. In object detection on the COCO dataset, task reward tuning improved mean average precision from 39.2% to 54.3% and average recall from 54.4% to 68.4%, outperforming specialized architectures without requiring complex post-processing. In panoptic segmentation, tuning improved the panoptic quality metric from 43.1% to 46.1%, effectively removing incoherent segmentation artifacts on small objects. For image colorization, optimizing a custom color reward increased vividness metrics from 0.46 to 0.97 and color diversity entropy from 1.03 to 1.84, resolving the muted tones common in standard models. In image captioning, consensus metrics improved significantly, rising from 120.0 to 134.5 on base models.

These results show that reinforcement learning provides a viable, general-purpose mechanism to align computer vision models with complex performance goals. By shifting the alignment burden from custom model engineering to reward definition, teams can lower architectural complexity and achieve superior task performance using standard model designs. The findings also indicate that while standard models internally generate high-quality candidate outputs, direct reward tuning is necessary to ensure the model reliably selects these optimal predictions at test time.

Engineering teams should adopt this two-stage framework when standard likelihood objectives fail to reflect real-world priorities, particularly for downstream operations like robotic perception or automated content generation. Organizations must carefully design and validate reward metrics before deployment to mitigate reward-hacking behaviors, such as models exploiting metric loopholes with incomplete text phrases or unnatural over-saturation. Future efforts should evaluate off-policy reinforcement learning methods to lower sampling compute overhead and explore multi-metric reward designs driven by human feedback.

arXiv: 2302.08242
  • Paper: Self-Critical Sequence Training for Image Captioning, Steven J. Rennie et al. (2016). This foundational paper establishes self-critical sequence training with policy gradients to optimize non-differentiable sequence metrics, providing the core reinforcement learning baseline mechanism adapted by the source.
  • Paper: Sequence Level Training with Recurrent Neural Networks, Marc'Aurelio Ranzato et al. (2015). This work introduces the two-stage training paradigm of initializing sequence models with maximum likelihood estimation before fine-tuning them via policy gradients to optimize non-differentiable task metrics.
  • Paper: Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks, Jiasen Lu et al. (2022). This research develops unified sequence-to-sequence Transformer architectures that represent diverse vision tasks as discrete sequence generation, establishing the model formulation tuned by the source.
  • Paper: Deformable DETR: Deformable Transformers for End-to-End Object Detection, Xizhou Zhu et al. (2021). This paper presents transformer-based end-to-end object detection, illustrating the specialized architecture and loss designs that the source aims to surpass using generic architectures fine-tuned with task rewards.
Cover for Tuning Computer Vision Models With Task Rewards

Abstract

Misalignment between model predictions and intended usage can be detrimental for the deployment of computer vision models. The issue is exacerbated when the task involves complex structured outputs, as it becomes harder to design procedures which address this misalignment. In natural language processing, this is often addressed using reinforcement learning techniques that align models with a task reward. We adopt this approach and show its surprising effectiveness to improve generic models pretrained to imitate example outputs across multiple computer vision tasks, such as object detection, panoptic segmentation, colorization and image captioning. We believe this approach has the potential to be widely useful for better aligning models with a diverse range of computer vision tasks.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Tuning Models With Rewards
  • 4. Practical Applications
  • 4.1. Panoptic Segmentation
  • 4.2. Object Detection
  • 4.3. Colorization
  • 4.4. Image Captioning
  • 5. Analysis
  • 5.1. Reward Distribution
  • 5.2. Reward-Risk Progression
  • 5.3. Ablation of Reward Baseline
  • 6. Discussion and Limitations
  • 7. Conclusion
  • Acknowledgements
  • References
  • A. Overview of models and hyper-parameters

Knowls

  1. Knowl 1 — Two-Stage MLE and REINFORCE Framework for Vision Sequence Models

    model/method

    Computer vision tasks can be formulated as learning a conditional probability distribution P(y∣x;θ)P(y|x; \theta) parameterized by θ\theta, where input image xx maps to a structured sequence y=[y1,y2,…,ym]y = [y_1, y_2, \dots, y_m] (e.g., text tokens, bounding box coordinates, or discrete per-pixel representations). The framework aligns generic encoder-decoder sequence models with non-differentiable task risk metrics in two sequential stages:

    1. Maximum Likelihood Estimation (MLE) Pretraining: Model parameters θ\theta are trained on a supervised dataset D={(x(i),y(i))}i=1N\mathcal{D} = \{(x^{(i)}, y^{(i)})\}_{i=1}^N by maximizing log-likelihood: max⁡θ∑i=1Nlog⁡P(y(i)∣x(i);θ)\max_{\theta} \sum_{i=1}^N \log P(y^{(i)} | x^{(i)}; \theta)

    2. Reward Optimization with REINFORCE: The pretrained model is tuned to maximize the expectation of an arbitrary, potentially non-differentiable task reward function R(x,y)R(x, y): max⁡θEx∼D[Ey∼P(⋅∣x;θ)[R(x,y)]]\max_{\theta} \mathbb{E}_{x \sim \mathcal{D}} \left[ \mathbb{E}_{y \sim P(\cdot|x; \theta)} [R(x, y)] \right]

    Using the log-derivative policy gradient theorem, the gradient of the expected reward with respect to θ\theta for a given input xx is estimated as: ∇θEy∼P(⋅∣x;θ)[R(x,y)]=Ey∼P(⋅∣x;θ)[(R(x,y)−b(x))∇θlog⁡P(y∣x;θ)]\nabla_{\theta} \mathbb{E}_{y \sim P(\cdot|x; \theta)} [R(x, y)] = \mathbb{E}_{y \sim P(\cdot|x; \theta)} \left[ (R(x, y) - b(x)) \nabla_{\theta} \log P(y|x; \theta) \right] where b(x)b(x) is a baseline reward independent of the sampled prediction yy, introduced to reduce gradient variance. Pretraining with MLE initializes the policy in a viable region of the large sequence space, avoiding reward sparsity issues that prevent training reinforcement learning models from scratch.

  2. Knowl 2 — Reward Optimization with Monte Carlo Sample Baseline

    algorithm

    The policy gradient step updates model parameters θ\theta on a mini-batch of nn inputs x={x1,…,xn}\mathbf{x} = \{x_1, \dots, x_n\} using a baseline-adjusted empirical reward.

    function step_reward(theta, x, alpha):
        # x is a mini-batch of n input images
        # alpha is the learning rate
        for i = 1 to n:
            ysample[i] = sample_model_output(theta, x[i])
            ybaseline[i] = sample_model_output(theta, x[i])
            r[i] = R(x[i], ysample[i]) - R(x[i], ybaseline[i])
        
        loss = (1 / n) * sum_{i=1}^n r[i] * log P(ysample[i] | x[i]; theta)
        Gr = grad_theta loss
        return theta + alpha * Gr

    For each input xix_i, two independent rollouts are sampled from P(⋅∣xi;θ)P(\cdot|x_i; \theta): a target output ysample(i)y_{\text{sample}}^{(i)} and a baseline output ybaseline(i)y_{\text{baseline}}^{(i)}. The reward difference ri=R(xi,ysample(i))−R(xi,ybaseline(i))r_i = R(x_i, y_{\text{sample}}^{(i)}) - R(x_i, y_{\text{baseline}}^{(i)}) serves as an unbiased variance-reduced advantage score that weights the sequence log-likelihood gradient.

  3. Knowl 3 — Object Detection Sequence Representation and Reward Formulations

    model/method

    In sequence-based object detection, bounding box predictions are tokenized into sequences of discretized coordinate tokens (1,000 spatial bins), a categorical class token, and a per-box confidence token.

    Two detection rewards are defined:

    1. Average Recall Reward: Formulated per image as the number of true positive matched ground-truth boxes ∣TP∣|TP| at a given Intersection-over-Union (IoU) threshold minus duplicate matched predictions ∣FPdup∣|FP_{\text{dup}}| with a penalty weight of 0.30.3: Rrecall(x,y)=∣TP∣−0.3∣FPdup∣R_{\text{recall}}(x, y) = |TP| - 0.3 |FP_{\text{dup}}|

    2. Mean Average Precision (mAP) Reward: Because mAP depends on ranking across the entire dataset and does not directly decompose per image, the model optimizes a composite objective: the average recall reward evaluated across multiple IoU thresholds reweighted by inverse class frequencies in the training set, combined with an auxiliary supervised loss on the prediction confidence token to regress the expected IoU between sampled output boxes and ground-truth boxes. This allows the model to opportunistically output dense candidates while scoring their precision rankings accurately.

  4. Knowl 4 — COCO Object Detection Performance under Reward Tuning

    data/table

    Object detection models initialized from MLE pretraining on Objects365 and COCO were tuned with REINFORCE using either recall or mAP rewards.

    Model mAP (%) AR@100 (%)
    Ours (ViT-B/16, MLE Baseline) 39.2 54.4
    Ours (ViT-B/16, Rewardrecall\text{Reward}_{\text{recall}}) 16.8 68.4
    Ours (ViT-B/16, RewardmAP\text{Reward}_{\text{mAP}}) 54.3 67.2
    Pix2seq (ViT-B, Chen et al., 2022) 47.1 N/A
    Pix2seq (ViT-L, Chen et al., 2022) 50.0 N/A

    Tuning solely for recall increases Average Recall (AR@100) from 54.4% to 68.4%, but drops mAP to 16.8% due to the emission of unranked, noisy bounding boxes. Tuning for mAP achieves 54.3% mAP and 67.2% AR@100 using a standard ViT-B/16 backbone at 1280×12801280 \times 1280 resolution, outperforming Pix2seq's ViT-B (47.1% mAP) and larger ViT-L model (50.0% mAP) without requiring custom data augmentation heuristics or non-maximum suppression (NMS).

  5. Knowl 5 — Panoptic Segmentation Reward Formulation and COCO Results

    data/table

    Panoptic Quality (PQ) computes a global metric over classes KK: PQ=1∣K∣∑k∈K∑(p,g)∈TPkIoU(p,g)∣TPk∣+12∣FPk∣+12∣FNk∣PQ = \frac{1}{|K|} \sum_{k \in K} \frac{\sum_{(p,g) \in TP_k} \text{IoU}(p, g)}{|TP_k| + \frac{1}{2}|FP_k| + \frac{1}{2}|FN_k|} where TPkTP_k, FPkFP_k, and FNkFN_k are true positive, false positive, and false negative instance matches for class kk.

    Because PQ is not directly additive across examples, the per-example reward for class kk is defined as: R(y,k)=[∑(p,g)∈TPkIoU(p,g)]−w∣FPk∣R(y, k) = \left[ \sum_{(p,g) \in TP_k} \text{IoU}(p, g) \right] - w |FP_k| with penalty weight w=0.3w = 0.3.

    A UViM architecture (ViT-L/16 encoder, 24-layer autoregressive decoder predicting a 256-token discrete sequence decoded to a 512×512512 \times 512 panoptic segmentation) was tuned using REINFORCE for 30,000 steps with batch size 128 and constant learning rate 10−610^{-6}.

    Model PQ (%)
    UViM512×512\text{UViM}_{512 \times 512} (MLE baseline) 43.1
    UViM512×512→\text{UViM}_{512 \times 512} \rightarrow Ours (PQ Tuned) 46.1
    UViM1280×1280\text{UViM}_{1280 \times 1280} 45.8

    Reward tuning increases PQ from 43.1% to 46.1%, surpassing the higher-resolution UViM1280×1280\text{UViM}_{1280 \times 1280} baseline (45.8%) while eliminating incoherent segmentations on small-scale objects.

  6. Knowl 6 — Image Colorization Multi-Objective Colorfulness Reward Formulation

    model/method

    Standard MLE models for image colorization produce plausible but desaturated, brownish predictions. To generate vivid and varied colors without human supervision, a multi-objective scalar reward is formulated in the CIELAB color space, where LL is lightness and a,ba, b represent chromaticity.

    The reward R(x,y)R(x, y) is defined as the product of a vividness term and a color diversity term: R(x,y)=Rvivid(y)⋅Rentropy(y)R(x, y) = R_{\text{vivid}}(y) \cdot R_{\text{entropy}}(y)

    1. Vividness Term (RvividR_{\text{vivid}}): The fraction of output pixels satisfying a2+b2>10a^2 + b^2 > 10. Applying a constant reward once saturation crosses this threshold prevents reward hacking where the model pushes color saturation to extreme, unnatural levels.

    2. Color Diversity Term (RentropyR_{\text{entropy}}): The Shannon entropy of the per-pixel hue angle θ=arctan⁡(b/a)\theta = \arctan(b/a), computed by binning hue into 7 uniform discrete intervals across the image. This prevents the model from choosing a single saturated color across the entire image.

    Tuning an MLE pretrained UViM colorization model (ViT-L/16 encoder, 24-layer decoder, 256 tokens) for 1,000 steps at learning rate 3⋅10−73 \cdot 10^{-7} increased the vivid pixel fraction from 0.46 to 0.97 and the hue entropy from 1.03 to 1.84.

  7. Knowl 7 — Reward Distribution and Likelihood Ranking Disconnect in Pretrained Generative Models

    empirical result

    Empirical analysis of 10,000 samples per image on COCO image captioning demonstrates why maximum likelihood estimation (MLE) is insufficient for task risk alignment:

    1. High-Reward Samples Exist in MLE Models: In the top 1% quantile of outputs generated by the MLE baseline, high-CIDEr samples exist. However, across all raw MLE samples, fewer than 5% achieve a CIDEr score ≥125\ge 125.

    2. Likelihood Fails to Retrieve High-Reward Outputs: When selecting the candidate with the highest sequence likelihood among NN samples (approximating greedy search, beam search, or nucleus sampling), average CIDEr plateaus well below task-tuned performance, remaining near 120 CIDEr even at N=10,000N = 10,000.

    3. Reward Tuning Shifts Policy Density: Tuning directly with REINFORCE shifts the probability distribution P(y∣x;θ)P(y|x; \theta) such that over 50% of direct samples achieve CIDEr ≥125\ge 125, substantially reducing the likelihood of generating low-quality outputs.

  8. Knowl 8 — Ablation of Baseline Sample Count in REINFORCE Gradient Estimation

    data/table

    The impact of baseline subtraction and the number of Monte Carlo baseline samples on policy gradient tuning performance was evaluated on Panoptic Segmentation (ViT-L, 10k steps) and Image Captioning (ViT-B, 10k steps).

    Task MLE No Baseline 1 Sample 3 Samples 7 Samples
    Panoptic (PQ %) 43.1 44.6 46.0 46.1 46.2
    Caption (CIDEr) 120.0 124.2 132.9 134.2 134.3

    Subtracting a single baseline sample (b=R(x,ybaseline)b = R(x, y_{\text{baseline}})) provides the primary variance reduction benefit (+1.4% PQ and +8.7 CIDEr over no baseline). Increasing the number of samples from 1 to 7 provides diminishing returns (+0.2% PQ and +1.4 CIDEr) while increasing sampling computational cost linearly.

  9. Knowl 9 — Image Captioning CIDEr Reward Tuning on COCO

    data/table

    Image captioning models using ViT encoders (initialized from ImageNet-21k) and 6-layer Transformer decoders (BERT 30k vocabulary, sequence length 128) were pretrained with MLE on COCO captions and tuned with REINFORCE using CIDEr reward and a 7-sample baseline.

    Model CIDEr
    Ours (ViT-B) 120.0 →\rightarrow 134.5
    Ours (ViT-L) 121.7 →\rightarrow 138.7
    Wang et al. (2022) N/A →\rightarrow 138.2
    Hu et al. (2022) (ExpansionNet v2) 128.7 →\rightarrow 143.7

    CIDEr reward optimization on the Karpathy test split yields a +14.5 point gain on ViT-B (120.0 to 134.5) and a +17.0 point gain on ViT-L (121.7 to 138.7), demonstrating that generic sequence models without vision-language specialized pretraining reach performance comparable to specialized captioning architectures.

  10. Knowl 10 — Limitations of Vision Task Reward Tuning: Reward Hacking and Sampling Cost

    limitation

    Tuning vision models with task rewards exhibits two core challenges:

    1. Reward Hacking: Models exploit misalignments in proxy reward functions. In image captioning, optimizing CIDEr leads models to generate ungrammatical trailing phrases (such as ending captions with "with a") to capture frequent dataset n-grams. In colorization, unconstrained saturation optimization causes extreme, single-color artifacts unless explicit entropy and saturation bounds are enforced.

    2. Computational Overhead of Autoregressive Sampling: While MLE pretraining computes sequence losses in parallel via teacher forcing, REINFORCE requires unrolled autoregressive sampling of candidate and baseline sequences at every training step. Hardware acceleration is substantially less efficient for autoregressive generation, making reward tuning computationally dominated by sequence generation rather than gradient backpropagation.

Coverage note — None was omitted; all main contributions—including the two-stage framework, task reward formulations and empirical evaluations for object detection, panoptic segmentation, colorization, and image captioning, reward distribution analyses, baseline ablations, and stated limitations—are fully covered.

References

  1. 1.Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N. Scheduled sampling for sequence prediction with recurrent neural networks. NeurIPS, 2015.
  2. 2.Beyer, L., Zhai, X., and Kolesnikov, A. Big Vision. https://github.com/google-research/big_vision, 2022.
  3. 3.Bhowmik, A., Gumhold, S., Rother, C., and Brachmann, E. Reinforced feature points: Optimizing feature detection and description for a high-level task. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020.
  4. 4.Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. End-to-end object detection with transformers. In ECCV, 2020.
  5. 5.Chen, T., Saxena, S., Li, L., Fleet, D. J., and Hinton, G. Pix2seq: A language modeling framework for object detection. In ICLR, 2022.
  6. 6.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint:1810.04805, 2018.
  7. 7.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  8. 8.Glaese, A., McAleese, N., Trebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al. Improving alignment of dialogue agents via targeted human judgements. arXiv preprint:2209.14375, 2022.
  9. 9.Henderson, P. and Ferrari, V. End-to-end training of object class detectors for mean average precision. In ACCV, 2017.
  10. 10.Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. The curious case of neural text degeneration. In ICLR, 2020.
  11. 11.Hu, J. C., Cavicchioli, R., and Capotondi, A. Expansionnet v2: Block static expansion in fast end to end training for image captioning. arXiv preprint:2208.06551, 2022.
  12. 12.Huang, C., Zhai, S., Guo, P., and Susskind, J. Metricopt: Learning to optimize black-box evaluation metrics. In CVPR, 2021.
  13. 13.Karpathy, A. and Fei-Fei, L. Deep visual-semantic alignments for generating image descriptions. In CVPR, 2015.
  14. 14.Keneshloo, Y., Shi, T., Ramakrishnan, N., and Reddy, C. K. Deep reinforcement learning for sequence-to-sequence models. IEEE TNNLS, 31(7):2469–2489, 2019.
  15. 15.Kirillov, A., He, K., Girshick, R., Rother, C., and Dollar, P. Panoptic segmentation. In CVPR, 2019.
  16. 16.Kolesnikov, A., Guillaumin, M., Ferrari, V., and Lampert, C. H. Closed-form training of conditional random fields for large scale image segmentation. ECCV, 2014.
  17. 17.Kolesnikov, A., Pinto, A. S., Beyer, L., Zhai, X., Harmsen, J. J., and Houlsby, N. UVim: A unified modeling approach for vision with learned guiding codes. In NeurIPS, 2022.
  18. 18.Krähenbóhl, P. and Koltun, V. Efficient inference in fully connected crfs with gaussian edge potentials. NeurIPS, 2011.
  19. 19.Kreutzer, J., Khadivi, S., Matusov, E., and Riezler, S. Can neural machine translation be improved with user feedback? In North American Chapter of the Association for Computational Linguistics: Human Language Technologies, (Industry Papers), June 2018.
  20. 20.Krull, A., Brachmann, E., Nowozin, S., Michel, F., Shotton, J., and Rother, C. Poseagent: Budget-constrained 6d object pose estimation via reinforcement learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  21. 21.Lafferty, J., McCallum, A., and Pereira, F. C. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In ICML, 2001.
  22. 22.Le, N., Rathour, V. S., Yamazaki, K., Luu, K., and Savvides, M. Deep reinforcement learning in computer vision: a comprehensive survey. Artificial Intelligence Review, 2021.
  23. 23.Leblond, R., Alayrac, J.-B., Sifre, L., Pislar, M., Jean-Baptiste, L., Antonoglou, I., Simonyan, K., and Vinyals, O. Machine translation decoding beyond beam search. In EMNLP, 2021.
  24. 24.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In ECCV, 2014.
  25. 25.Lin, T.-Y., Goyal, P., Girshick, R., He, K., and Dollár, P. Focal loss for dense object detection. In ICCV, 2017.
  26. 26.Mathe, S., Pirinen, A., and Sminchisescu, C. Reinforcement learning for visual object detection. In CVPR, 2016.
  27. 27.Nowozin, S., Rother, C., Bagon, S., Sharp, T., Yao, B., and Kohli, P. Decision tree fields. In ICCV, 2011.
  28. 28.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. Training language models to follow instructions with human feedback. In NeurIPS, 2022.
  29. 29.Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. Zero-shot text-to-image generation. In ICML, 2021.
  30. 30.Ranzato, M., Chopra, S., Auli, M., and Zaremba, W. Sequence level training with recurrent neural networks. arXiv preprint:1511.06732, 2015.
  31. 31.Rao, Y., Lin, D., Lu, J., and Zhou, J. Learning globally optimized object detector via policy gradient. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6190–6198, 2018.
  32. 32.Ren, S., He, K., Girshick, R., and Sun, J. Faster r-cnn: Towards real-time object detection with region proposal networks. NeurIPS, 28, 2015.
  33. 33.Rennie, S. J., Marcheret, E., Mroueh, Y., Ross, J., and Goel, V. Self-critical sequence training for image captioning. In CVPR, 2017.
  34. 34.Schmidt, F. Generalization in generation: A closer look at exposure bias. In Proceedings of the 3rd Workshop on Neural Generation and Translation, 2019.
  35. 35.Shao, S., Li, Z., Zhang, T., Peng, C., Yu, G., Zhang, X., Li, J., and Sun, J. Objects365: A large-scale, high-quality dataset for object detection. In ICCV, 2019.
  36. 36.Shazeer, N. and Stern, M. Adafactor: Adaptive learning rates with sublinear memory cost. In ICML, 2018.
  37. 37.Shen, S., Cheng, Y., He, Z., He, W., Wu, H., Sun, M., and Liu, Y. Minimum risk training for neural machine translation. arXiv preprint:1512.02433, 2015.
  38. 38.Song, Y., Schwing, A., Urtasun, R., et al. Training deep neural networks via direct loss minimization. In ICML, 2016.
  39. 39.Stahlberg, F. and Byrne, B. On NMT search errors and model errors: Cat got your tongue? In EMNLP-IJCNLP, 2019.
  40. 40.Steiner, A., Kolesnikov, A., Zhai, X., Wightman, R., Uszkoreit, J., and Beyer, L. How to train your ViT? data, augmentation, and regularization in vision transformers. arXiv preprint:2106.10270, 2021.
  41. 41.Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. Learning to summarize with human feedback. NeurIPS, 2020.
  42. 42.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. In NeurIPS, 2017.
  43. 43.Vedantam, R., Lawrence Zitnick, C., and Parikh, D. Cider: Consensus-based image description evaluation. In CVPR, 2015.
  44. 44.Wang, Y., Xu, J., and Sun, Y. End-to-end transformer based model for image captioning. arXiv preprint:2203.15350, 2022.
  45. 45.Williams, R. J. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8(3):229–256, 1992.
  46. 46.Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L. Scaling vision transformers. In CVPR, 2022.

Citation

MLA
Pinto, A. S., et al. “Tuning Computer Vision Models With Task Rewards”. International Conference on Machine Learning, vol. 202, 2023, pp. 33229–39, https://proceedings.mlr.press/v202/susano-pinto23a.html.
APA
Pinto, A. S., Kolesnikov, A., Shi, Y., Beyer, L., & Zhai, X. (2023). Tuning Computer Vision Models With Task Rewards. International Conference on Machine Learning, 202, 33229–33239. https://proceedings.mlr.press/v202/susano-pinto23a.html
Chicago
Pinto, A. S., A. Kolesnikov, Y. Shi, L. Beyer, and X. Zhai. 2023. “Tuning Computer Vision Models With Task Rewards”. International Conference on Machine Learning 202: 33229–39. https://proceedings.mlr.press/v202/susano-pinto23a.html.
Harvard
Pinto, A.S. et al. (2023) “Tuning Computer Vision Models With Task Rewards”, International Conference on Machine Learning. PMLR, pp. 33229–33239. Available at: https://proceedings.mlr.press/v202/susano-pinto23a.html.
Vancouver
1. Pinto AS, Kolesnikov A, Shi Y, Beyer L, Zhai X (2023) Tuning Computer Vision Models With Task Rewards. In: International Conference on Machine Learning. PMLR, pp 33229–33239

BibTeX

@InProceedings{pmlr-v202-susano-pinto23a,
  title = 	 {Tuning Computer Vision Models With Task Rewards},
  author =       {Susano Pinto, Andr\'{e} and Kolesnikov, Alexander and Shi, Yuge and Beyer, Lucas and Zhai, Xiaohua},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {33229--33239},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/susano-pinto23a/susano-pinto23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/susano-pinto23a.html},
  abstract = 	 {Misalignment between model predictions and intended usage can be detrimental for the deployment of computer vision models. The issue is exacerbated when the task involves complex structured outputs, as it becomes harder to design procedures which address this misalignment. In natural language processing, this is often addressed using reinforcement learning techniques that align models with a task reward. We adopt this approach and show its surprising effectiveness to improve generic models pretrained to imitate example outputs across multiple computer vision tasks, such as object detection, panoptic segmentation, colorization and image captioning. We believe this approach has the potential to be widely useful for better aligning models with a diverse range of computer vision tasks.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/