What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models

Yu PengJingjing FuChuheng ZhangLi ZhaoJiang BianMingyu LiuLing ZhangJun ZhangRui Wang

article2026arXiv1 citations

Proposes PAIR-VLA, a reinforcement learning framework that optimizes vision-language-action models through paired invariance and sensitivity objectives, enabling robotic policies to ignore visual distractors while reacting appropriately to task-critical scene changes under out-of-distribution shifts.

Listen

Vision-Language-Action models show strong promise for general-purpose robotic manipulation, but they frequently fail when deployed in environments with visual changes such as new lighting, unfamiliar table textures, altered camera angles, or distracting objects. Standard reinforcement learning fine-tuning relies on task-success rewards, which improve general policy performance but do not explicitly teach robots whether a visual change should be ignored as irrelevant or acted upon as a critical modification to the task. Traditional methods like visual domain randomization broaden visual exposure, yet they fail to provide direct behavioral instruction on how actions should respond to distinct scene shifts.

The article introduces and evaluates PAIR-VLA, a reinforcement learning fine-tuning framework designed to instill visual robustness by guiding policy responses at the action level. The primary objective is to demonstrate that augmenting standard reinforcement learning with action invariance and sensitivity objectives significantly improves robot manipulation performance under diverse, out-of-distribution visual shifts without increasing operational complexity during deployment.

To achieve this, the approach constructs two paired visual variants of each observation during training. A task-preserving view changes irrelevant factors like background textures and distractor objects while keeping the target object unchanged, whereas a task-altering view modifies the target object's position or orientation. The framework adds two auxiliary objectives to the standard reinforcement learning optimization: an invariance objective that forces action distributions to remain consistent across task-preserving views, and a sensitivity objective that forces action distributions to adapt when the target object changes. The evaluation tested two standard model architectures across pick-and-place tasks within a simulated manipulation environment against unseen textures, lighting, target poses, clutter levels, and camera viewpoints.

The analysis produced several key findings. First, the proposed framework consistently outperformed standard reinforcement learning baselines across all out-of-distribution visual test scenarios, achieving an average absolute success rate improvement of 16.62% on the flow-matching architecture and 9.10% on the autoregressive architecture. Second, the method exhibited strong visual generalization, successfully transferring robustness to unseen lighting conditions and camera angles even when those specific factors were not altered during the construction of training pairs. Third, reinforcement learning fine-tuning efficiency improved by approximately three times, reaching high success thresholds in roughly 80 training steps compared to 240 steps for the baseline. Finally, ablation tests revealed that while invariance guidance provided the primary robustness gain (accounting for a 12.61% baseline improvement), coupling it with sensitivity guidance achieved the strongest overall performance.

These findings indicate that visual robustness can be directly embedded into robotic policies during post-training optimization rather than relying on complex, latency-inducing runtime visual processing modules. Because the auxiliary objectives operate strictly during training, deployed systems run on standard architectures without incurring inference-time latency or computational overhead. This offers an efficient path to reducing failure risks and adaptation costs in automated manipulation workflows.

Decision-makers should consider adopting behavior-level auxiliary objectives when fine-tuning robotic policies to improve operational reliability and training throughput. However, because the current evaluations were conducted in simulation using exact segmentation masks, organizations should validate these gains in physical environments using off-the-shelf visual segmentation tools before full-scale deployment.

arXiv: 2605.13105

No sufficiently relevant recommendations were found.

Cover for What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models

Abstract

Reinforcement learning (RL) fine-tuning has shown promise for Vision-Language-Action (VLA) models in robotic manipulation, but deployment-time visual shifts pose practical challenges. A key difficulty is that standard task rewards supervise task success, but offer limited guidance on whether a visual change is task-irrelevant or changes the behavior required for manipulation. We propose PAIR-VLA (Paired Action Invariance & Sensitivity for Visually Robust VLA), an RL fine-tuning framework to address this difficulty by adding two auxiliary objectives over paired visual variants during PPO optimization: an invariance term that reduces the discrepancy between action distributions for a task-preserving pair (e.g., different distractors), and a sensitivity objective that encourages separable action distributions for a task-altering pair (e.g., target object in a different pose). Together, these objectives turn visual variants from mere observation diversity into behavior-level guidance on policy responses during RL fine-tuning. We evaluate on ManiSkill3 across two representative VLA architectures, OpenVLA and π0.5\pi_{0.5}, under diverse out-of-distribution visual shifts including unseen distractors, texture changes, target object pose variation, viewpoint shifts, and lighting changes. Our method consistently improves over standard PPO, achieving average improvements of 16.62% on π0.5\pi_{0.5} and 9.10% on OpenVLA. Notably, ablations further show generalization across visual shifts: invariance guidance learned from distractor and texture variants transfers to target-pose and lighting shifts, while adding sensitivity guidance on target-pose variants further improves robustness to nuisance shifts, highlighting the broader transferability of behavior-level RL guidance.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 4 Visual Robustness via Invariance and Sensitivity Objectives
  • 4.1 Problem Formulation
  • 4.2 Invariance and Sensitivity Objectives
  • 4.3 Augmented View Construction
  • 5 Experiments
  • 5.1 Experimental Setup
  • 5.2 Main Results
  • 5.3 Ablation Study
  • 5.4 Analysis
  • 5.5 Generalization to Unseen Camera Viewpoints
  • 6 Discussion & Conclusion
  • References
  • A SFT and RL Training Details
  • A.1 SFT Checkpoints
  • A.2 RL Training Details
  • B Benchmark Details
  • B.1 Table Textures and Objects
  • B.2 Scene Lighting
  • B.3 Object Position
  • B.4 Training and Evaluation Tasks
  • C Broader Impacts

Knowls

  1. Knowl 1 — PAIR-VLA Reinforcement Learning Fine-Tuning Framework

    model/method

    PAIR-VLA (Paired Action Invariance & Sensitivity for Visually Robust VLA) is an online reinforcement learning fine-tuning framework for Vision-Language-Action (VLA) policies. It augments the standard Proximal Policy Optimization (PPO) objective with two auxiliary behavior-level action distribution objectives computed over paired visual variants of the current observation:

    L(θ)=LPPO(θ)+αLinv(θ)+βLsens(θ)\mathcal{L}(\theta) = \mathcal{L}_{\text{PPO}}(\theta) + \alpha \mathcal{L}_{\text{inv}}(\theta) + \beta \mathcal{L}_{\text{sens}}(\theta)

    where θ\theta denotes the policy parameters, LPPO(θ)\mathcal{L}_{\text{PPO}}(\theta) is the PPO clipped surrogate objective, Linv(θ)\mathcal{L}_{\text{inv}}(\theta) is the task-preserving invariance loss, Lsens(θ)\mathcal{L}_{\text{sens}}(\theta) is the task-altering sensitivity loss, and α,β>0\alpha, \beta > 0 are weighting hyperparameters.

    The framework applies to autoregressive VLAs (e.g., OpenVLA) using direct action log-likelihoods and to flow-matching VLAs (e.g., π0.5\pi_{0.5}) via a flow-SDE two-layer MDP formulation. Because the auxiliary objectives and paired views are evaluated only during training, the deployed policy architecture and inference runtime remain unchanged.

  2. Knowl 2 — Task-Preserving Invariance Objective

    equation

    The task-preserving invariance loss aligns the action distribution predicted under the original observation with the action distribution predicted under a task-preserving perturbed view:

    Linv(θ)=E(ot,l)[DKL(πθ(⋅∣ot,l) ∥ sg[πθ(⋅∣o~tprev,l)])]\mathcal{L}_{\text{inv}}(\theta) = \mathbb{E}_{(o_t, l)} \left[ D_{\text{KL}}\left(\pi_\theta(\cdot \mid o_t, l) \,\parallel\, \text{sg}\left[\pi_\theta(\cdot \mid \tilde{o}_t^{\text{prev}}, l)\right]\right) \right]

    where:

    • oto_t is the current visual observation at time step tt.
    • ll is the natural language task instruction.
    • o~tprev\tilde{o}_t^{\text{prev}} is the task-preserving view of oto_t, in which task-irrelevant visual factors (such as distractors or background textures) are altered while keeping the task state unchanged.
    • πθ(⋅∣⋅)\pi_\theta(\cdot \mid \cdot) denotes the VLA policy action distribution parameterized by θ\theta.
    • DKL(⋅∥⋅)D_{\text{KL}}(\cdot \parallel \cdot) is the Kullback-Leibler divergence.
    • sg[⋅]\text{sg}[\cdot] denotes the stop-gradient operator, which treats the distribution under the perturbed view as a fixed target to stabilize training.
  3. Knowl 3 — Task-Altering Sensitivity Objective

    equation

    The task-altering sensitivity loss encourages the policy to output distinct action distributions when visual modifications alter the required manipulation behavior:

    Lsens(θ)=−E(ot,l)[min⁡(c,DKL(πθ(⋅∣ot,l) ∥ sg[πθ(⋅∣o~talt,l)]))]\mathcal{L}_{\text{sens}}(\theta) = -\mathbb{E}_{(o_t, l)} \left[ \min\left(c, D_{\text{KL}}\left(\pi_\theta(\cdot \mid o_t, l) \,\parallel\, \text{sg}\left[\pi_\theta(\cdot \mid \tilde{o}_t^{\text{alt}}, l)\right]\right)\right) \right]

    where:

    • oto_t is the current visual observation.
    • ll is the natural language task instruction.
    • o~talt\tilde{o}_t^{\text{alt}} is the task-altering view, in which task-relevant properties (e.g., target object pose) are perturbed.
    • sg[⋅]\text{sg}[\cdot] denotes the stop-gradient operator.
    • c>0c > 0 is a scalar clipping threshold that caps the divergence to prevent unbounded divergence growth from destabilizing PPO policy updates.
  4. Knowl 4 — Task-Preserving and Task-Altering View Generation

    model/method

    PAIR-VLA synthesizes two complementary perturbed views for each observation oto_t during policy training:

    1. Task-Preserving View (o~tprev\tilde{o}_t^{\text{prev}}): Constructed via foreground-background compositing. Prior to training, a set of KK background snapshots {okbg}k=1K\{o_k^{\text{bg}}\}_{k=1}^K is pre-rendered by rendering scene configurations with all objects invisible. During training, a binary segmentation mask mtm_t assigns value 1 to pixels belonging to task-relevant entities (robot, target object, receptacle) and 0 to background and distractor objects. A background snapshot okbgo_k^{\text{bg}} is sampled uniformly, and the view is generated as:

    o~tprev=mt⊙ot+(1−mt)⊙okbg\tilde{o}_t^{\text{prev}} = m_t \odot o_t + (1 - m_t) \odot o_k^{\text{bg}}

    where ⊙\odot denotes element-wise multiplication. Alternatively, when studying viewpoint robustness, o~tprev\tilde{o}_t^{\text{prev}} is generated by re-rendering the identical scene state from an altered camera pose.

    1. Task-Altering View (o~talt\tilde{o}_t^{\text{alt}}): Constructed by perturbing the target object pose while maintaining its semantic identity. Translation perturbations are sampled from a Gaussian distribution and rotation perturbations from a categorical distribution, applied to the target object, and the scene is re-rendered.
  5. Knowl 5 — Definition of Task-Preserving Invariance and Task-Altering Sensitivity

    definition

    In visually robust reinforcement learning for Vision-Language-Action (VLA) policies πθ(a∣o,l)\pi_\theta(a \mid o, l), visual generalization is defined through two complementary behavioral requirements:

    1. Task-Preserving Invariance: The action distribution πθ(⋅∣o,l)\pi_\theta(\cdot \mid o, l) must remain approximately constant under task-irrelevant visual changes, including variations in table textures, background appearance, lighting conditions, novel visual distractors, and camera viewpoint shifts that do not change the underlying manipulation task.

    2. Task-Altering Sensitivity: The action distribution πθ(⋅∣o,l)\pi_\theta(\cdot \mid o, l) must adjust meaningfully when task-relevant properties change, such as variations in the target object pose, location, or orientation that demand different grasping, trajectory, or placement actions.

  6. Knowl 6 — Out-of-Distribution Generalization Performance across Visual Shifts

    data/table

    PAIR-VLA was evaluated on the ManiSkill3 pick-and-place benchmark across four out-of-distribution (OOD) visual shift scenarios against standard PPO fine-tuning using both autoregressive (OpenVLA) and flow-matching (π0.5\pi_{0.5}) backbones. Success rates (%) are averaged across three seeds (128 evaluation episodes per run):

    Model Method Table Texture Lighting Target Pose Clutter Avg. ( Avg.)
    OpenVLA PPO 86.98 72.14 83.59 68.88 77.90 (–)
    OpenVLA PAIR-VLA (Ours) 94.53 80.47 90.63 82.36 87.00 (+9.10)
    π0.5\pi_{0.5} PPO 63.54 28.54 56.46 36.46 46.25 (–)
    π0.5\pi_{0.5} PAIR-VLA (Ours) 80.21 51.67 69.38 50.21 62.87 (+16.62)

    The clutter column reflects the average success rate across test scenes with 2, 4, 6, and 8 distractors containing held-out object categories. PAIR-VLA achieves consistent absolute performance gains across all visual shifts on both backbones, including unseen lighting shifts (+8.33% on OpenVLA, +23.13% on π0.5\pi_{0.5}) despite lighting not being explicitly varied in the paired views.

  7. Knowl 7 — Ablation Analysis of Invariance and Sensitivity Auxiliary Objectives

    data/table

    An ablation study evaluating the individual and joint effects of the invariance loss (Linv\mathcal{L}_{\text{inv}}) and sensitivity loss (Lsens\mathcal{L}_{\text{sens}}) was conducted on ManiSkill3 using the π0.5\pi_{0.5} backbone across four OOD visual shift settings:

    Method Table Texture Lighting Target Pose Clutter Avg. Avg.
    PPO Baseline 63.54 28.54 56.46 36.46 46.25 –
    PPO w/ Linv\mathcal{L}_{\text{inv}} 72.92 44.58 67.92 50.00 58.86 +12.61
    PPO w/ Lsens\mathcal{L}_{\text{sens}} 63.96 33.54 53.33 41.10 47.98 +1.73
    Ours (Linv+Lsens\mathcal{L}_{\text{inv}} + \mathcal{L}_{\text{sens}}) 80.21 51.67 69.38 50.21 62.87 +16.62

    The results indicate that Linv\mathcal{L}_{\text{inv}} provides the primary robustness benefit (+12.61% overall gain over PPO), while Lsens\mathcal{L}_{\text{sens}} provides a complementary effect that yields an additional +4.01% average gain when combined with Linv\mathcal{L}_{\text{inv}}.

  8. Knowl 8 — RL Fine-Tuning Sample Efficiency and Clutter Scalability

    empirical result

    PAIR-VLA demonstrates enhanced training efficiency and scalability under visual clutter compared to standard PPO on the ManiSkill3 benchmark:

    1. Training Efficiency: In an in-distribution pick-and-place task with 1 distractor, PAIR-VLA with OpenVLA reaches a 90% success rate within 80 PPO training steps, whereas baseline PPO requires approximately 240 steps to achieve 90% success, representing an approximate 3×3\times sample efficiency improvement. A similar acceleration profile is maintained in OOD clutter scenarios (4 distractors).

    2. Robustness to Increasing Clutter: When evaluating performance under 2, 4, 6, and 8 distractors (where half the distractors are sampled from an unseen held-out category set), PAIR-VLA exhibits a more graceful degradation profile than PPO. In the densest clutter condition with 8 distractors, PAIR-VLA achieves 72% success on OpenVLA (vs. 56% for PPO) and 33% on π0.5\pi_{0.5} (vs. 21% for PPO).

  9. Knowl 9 — Camera Viewpoint Extrapolation via Viewpoint-Based Invariance

    empirical result

    When instantiated with task-preserving views generated by varying camera rotation angles during RL fine-tuning on π0.5\pi_{0.5}:

    • In-Distribution Viewpoints: Over the training range of camera rotation angles [0∘,20∘][0^\circ, 20^\circ] (sampled at 4∘4^\circ increments), PAIR-VLA achieves an average success rate of 70.36%, matching the baseline PPO success rate of 70.97%.
    • Out-of-Distribution Extrapolation: At unseen camera rotation angles beyond the training distribution, PAIR-VLA significantly outperforms baseline PPO, achieving 38.34% vs. 27.92% success at 24∘24^\circ, and 32.92% vs. 20.63% success at 28∘28^\circ.
  10. Knowl 10 — Limitations of PAIR-VLA

    limitation

    The PAIR-VLA framework has two primary stated limitations:

    1. Simulation Evaluation: The experimental validation is restricted to simulated manipulation tasks within the ManiSkill3 simulator, leaving real-world transfer and physical deployment untested.

    2. Dependence on Ground-Truth Simulator Information: The construction of task-preserving views relies on exact object segmentation masks and simulator APIs for background compositing and target pose rendering. Transfer to real-world environments requires approximating masks with zero-shot vision models (e.g., SAM 3) and image inpainting/synthesis techniques, the noise of which was not evaluated.

Coverage note — None was omitted; all key theoretical formulations, algorithmic mechanisms, empirical results, ablations, and limitations are covered.

References

  1. 1.Abby O'Neill, Abdul Rehman, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 6892–6903. IEEE, 2024.
  2. 2.Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, et al. Droid: A large-scale in-the-wild robot manipulation dataset. arXiv preprint arXiv:2403.12945, 2024.
  3. 3.Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, Pete Florence, Chuyuan Fu, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Kehang Han, Karol Hausman, Alexander Herzog, Jasmine Hsu, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Lisa Lee, Tsang-Wei Edward Lee, Sergey Levine, Yao Lu, Henryk Michalewski, Igor Mordatch, Karl Pertsch, Kanishka Rao, Krista Reymann, Michael Ryoo, Grecia Salazar, Pannag Sanketi, Pierre Sermanet, Jaspiar Singh, Anikait Singh, Radu Soricut, Huong Tran, Vincent Vanhoucke, Quan Vuong, Ayzaan Wahid, Stefan Welker, Paul Wohlhart, Jialin Wu, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, and Brianna Zitkovich. Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023.
  4. 4.Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, Jianlan Luo, You Liang Tan, Lawrence Yunliang Chen, Pannag Sanketi, Quan Vuong, Ted Xiao, Dorsa Sadigh, Chelsea Finn, and Sergey Levine. Octo: An open-source generalist robot policy, 2024.
  5. 5.Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al. π0\pi_0: A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164, 2024.
  6. 6.Kaylee Burns, Zach Witzel, Jubayer Ibn Hamid, Tianhe Yu, Chelsea Finn, and Karol Hausman. What makes pre-trained visual representations successful for robust manipulation?, 2023.
  7. 7.Youguang Xing, Xu Luo, Junlin Xie, Lianli Gao, Hengtao Shen, and Jingkuan Song. Shortcut learning in generalist robot policies: The role of dataset diversity and fragmentation, 2025.
  8. 8.Haozhan Li, Yuxin Zuo, Jiale Yu, Yuhao Zhang, Zhaohui Yang, Kaiyan Zhang, Xuekai Zhu, Yuchen Zhang, Tianxing Chen, Ganqu Cui, Dehui Wang, Dingxiang Luo, Yuchen Fan, Youbang Sun, Jia Zeng, Jiangmiao Pang, Shanghang Zhang, Yu Wang, Yao Mu, Bowen Zhou, and Ning Ding. Simplevla-rl: Scaling vla training via reinforcement learning, 2025.
  9. 9.Chao Yu, Yuanqing Wang, Zhen Guo, Hao Lin, Si Xu, Hongzhi Zang, Quanlu Zhang, Yongji Wu, Chunyang Zhu, Junhao Hu, et al. Rlinf: Flexible and efficient large-scale reinforcement learning via macro-to-micro flow transformation. arXiv preprint arXiv:2509.15965, 2025.
  10. 10.Jijia Liu, Feng Gao, Bingwen Wei, Xinlei Chen, Qingmin Liao, Yi Wu, Chao Yu, and Yu Wang. What can rl bring to vla generalization? an empirical study, 2026.
  11. 11.Kang Chen, Zhihao Liu, Tonghe Zhang, Zhen Guo, Si Xu, Hao Lin, Hongzhi Zang, Xiang Li, Quanlu Zhang, Zhaofei Yu, Guoliang Fan, Tiejun Huang, Yu Wang, and Chao Yu. πRL\pi_{\text{RL}}: Online rl fine-tuning for flow-based vision-language-action models, 2026.
  12. 12.Hongzhi Zang, Mingjie Wei, Si Xu, Yongji Wu, Zhen Guo, Yuanqing Wang, Hao Lin, Peihong Wang, Liangzhi Shi, Yuqing Xie, Zhexuan Xu, Zhihao Liu, Kang Chen, Wenhao Tang, Quanlu Zhang, Weinan Zhang, Chao Yu, and Yu Wang. Rlinf-vla: A unified and efficient framework for reinforcement learning of vision-language-action models, 2026.
  13. 13.Hongyin Zhang, Shuo Zhang, Junxi Jin, Qixin Zeng, Runze Li, and Donglin Wang. Robustvla: Robustness-aware reinforcement post-training for vision-language-action models. arXiv preprint arXiv:2511.01331, 2025.
  14. 14.Ilya Kostrikov, Denis Yarats, and Rob Fergus. Image augmentation is all you need: Regularizing deep reinforcement learning from pixels, 2021.
  15. 15.Michael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto, Pieter Abbeel, and Aravind Srinivas. Reinforcement learning with augmented data, 2020.
  16. 16.Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto. Mastering visual continuous control: Improved data-augmented reinforcement learning, 2021.
  17. 17.Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel. Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS), pages 23–30. IEEE, 2017.
  18. 18.Jonathan Tremblay, Aayush Prakash, David Acuna, Mark Brophy, Varun Jampani, Cem Anil, Thang To, Eric Cameracci, Shaad Boochoon, and Stan Birchfield. Training deep networks with synthetic data: Bridging the reality gap by domain randomization. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pages 969–977, 2018.
  19. 19.Bhairav Mehta, Manfred Diaz, Florian Golemo, Christopher J Pal, and Liam Paull. Active domain randomization. In Conference on Robot Learning, pages 1162–1176. PMLR, 2020.
  20. 20.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  21. 21.Stone Tao, Fanbo Xiang, Arth Shukla, Yuzhe Qin, Xander Hinrichsen, Xiaodi Yuan, Chen Bao, Xinsong Lin, Yulin Liu, Tse-kai Chan, et al. Maniskill3: Gpu parallelized robotics simulation and rendering for generalizable embodied ai. arXiv preprint arXiv:2410.00425, 2024.
  22. 22.Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024.
  23. 23.Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, et al. π0.5\pi_{0.5}: A vision-language-action model with open-world generalization. arXiv preprint arXiv:2504.16054, 2025.
  24. 24.Jiaheng Hu, Rose Hendrix, Ali Farhadi, Aniruddha Kembhavi, Roberto Martín-Martín, Peter Stone, Kuo-Hao Zeng, and Kiana Ehsani. Flare: Achieving masterful and adaptive robot policies with large-scale reinforcement learning fine-tuning. In 2025 IEEE International Conference on Robotics and Automation (ICRA), pages 3617–3624. IEEE, 2025.
  25. 25.Moo Jin Kim, Chelsea Finn, and Percy Liang. Fine-tuning vision-language-action models: Optimizing speed and success, 2025.
  26. 26.Bram Grooten, Tristan Tomilin, Gautham Vasan, Matthew E Taylor, A Rupam Mahmood, Meng Fang, Mykola Pechenizkiy, and Decebal Constantin Mocanu. Madi: Learning to mask distractions for generalization in visual deep reinforcement learning. arXiv preprint arXiv:2312.15339, 2023.
  27. 27.Sangmim Song, Sarath Kodagoda, Marc Carmichael, and Karthick Thiyagarajan. Overcoming visual clutter in vision language action models via concept-gated visual distillation. arXiv preprint arXiv:2603.10340, 2026.
  28. 28.Shihan Wu, Xu Luo, Ji Zhang, Junlin Xie, Jingkuan Song, Heng Tao Shen, and Lianli Gao. Policy contrastive decoding for robotic foundation models. arXiv preprint arXiv:2505.13255, 2025.
  29. 29.Asher J. Hancock, Allen Z. Ren, and Anirudha Majumdar. Run-time observation interventions make vision-language-action models more visually robust, 2024.
  30. 30.Tianhe Yu, Ted Xiao, Austin Stone, Jonathan Tompson, Anthony Brohan, Su Wang, Jaspiar Singh, Clayton Tan, Jodilyn Peralta, Brian Ichter, et al. Scaling robot learning with semantically imagined experience. arXiv preprint arXiv:2302.11550, 2023.
  31. 31.Daniel Ho, Kanishka Rao, Zhuo Xu, Eric Jang, Mohi Khansari, and Yunfei Bai. Retinagan: An object-aware approach to sim-to-real transfer. In 2021 IEEE International Conference on Robotics and Automation (ICRA), pages 10920–10926. IEEE, 2021.
  32. 32.Amy Zhang, Rowan McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine. Learning invariant representations for reinforcement learning without reconstruction. arXiv preprint arXiv:2006.10742, 2020.
  33. 33.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In Hal Daumé III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 1597–1607. PMLR, 13–18 Jul 2020.
  34. 34.Michael Laskin, Aravind Srinivas, and Pieter Abbeel. Curl: Contrastive unsupervised representations for reinforcement learning. In International conference on machine learning, pages 5639–5650. PMLR, 2020.
  35. 35.David Bertoin, Adil Zouitine, Mehdi Zouitine, and Emmanuel Rachelson. Look where you look! saliency-guided q-networks for generalization in visual reinforcement learning. Advances in neural information processing systems, 35:30693–30706, 2022.
  36. 36.Jingbo Sun, Songjun Tu, Qichao Zhang, Ke Chen, and Dongbin Zhao. Salience-invariant consistent policy learning for generalization in visual reinforcement learning, 2025.
  37. 37.Jonathan Yang, Chelsea Finn, and Dorsa Sadigh. Invariance co-training for robot visual generalization. arXiv preprint arXiv:2512.05230, 2025.
  38. 38.Stephen James and Andrew J Davison. Q-attention: Enabling efficient learning for vision-based robotic manipulation. IEEE Robotics and Automation Letters, 7(2):1612–1619, 2022.
  39. 39.Vincent Pacelli and Anirudha Majumdar. Learning task-driven control policies via information bottlenecks, 2020.
  40. 40.Shuoshuo Zhang, Yizhen Zhang, Jingjing Fu, Lei Song, Jiang Bian, Yujiu Yang, and Rui Wang. See less, see right: Bi-directional perceptual shaping for multimodal reasoning. arXiv preprint arXiv:2512.22120, 2025.
  41. 41.John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel. High-dimensional continuous control using generalized advantage estimation. arXiv preprint arXiv:1506.02438, 2015.
  42. 42.Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Didac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, et al. Sam 3: Segment anything with concepts. arXiv preprint arXiv:2511.16719, 2025.
  43. 43.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Liang Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models. Iclr, 1(2):3, 2022.

Citation

MLA
Peng, Y., et al. “What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models”. arXiv, 2026, http://arxiv.org/abs/2605.13105v1.
APA
Peng, Y., Fu, J., Zhang, C., Zhao, L., Bian, J., Liu, M., Zhang, L., Zhang, J., & Wang, R. (2026). What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models. arXiv. http://arxiv.org/abs/2605.13105v1
Chicago
Peng, Y., J. Fu, C. Zhang, et al. 2026. “What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models”. arXiv. http://arxiv.org/abs/2605.13105v1.
Harvard
Peng, Y. et al. (2026) “What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2605.13105v1.
Vancouver
1. Peng Y, Fu J, Zhang C, Zhao L, Bian J, Liu M, Zhang L, Zhang J, Wang R (2026) What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models. arXiv

BibTeX

@article{peng2026what,
  title = {What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models},
  author = {Peng, Yuanfang and Fu, Jingjing and Zhang, Chuheng and Zhao, Li and Bian, Jiang and Liu, Mingyu and Zhang, Ling and Zhang, Jun and Wang, Rui},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2605.13105v1},
  eprint = {2605.13105}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission