What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models
Yu PengJingjing FuChuheng ZhangLi ZhaoJiang BianMingyu LiuLing ZhangJun ZhangRui Wang
Proposes PAIR-VLA, a reinforcement learning framework that optimizes vision-language-action models through paired invariance and sensitivity objectives, enabling robotic policies to ignore visual distractors while reacting appropriately to task-critical scene changes under out-of-distribution shifts.
Vision-Language-Action models show strong promise for general-purpose robotic manipulation, but they frequently fail when deployed in environments with visual changes such as new lighting, unfamiliar table textures, altered camera angles, or distracting objects. Standard reinforcement learning fine-tuning relies on task-success rewards, which improve general policy performance but do not explicitly teach robots whether a visual change should be ignored as irrelevant or acted upon as a critical modification to the task. Traditional methods like visual domain randomization broaden visual exposure, yet they fail to provide direct behavioral instruction on how actions should respond to distinct scene shifts.
The article introduces and evaluates PAIR-VLA, a reinforcement learning fine-tuning framework designed to instill visual robustness by guiding policy responses at the action level. The primary objective is to demonstrate that augmenting standard reinforcement learning with action invariance and sensitivity objectives significantly improves robot manipulation performance under diverse, out-of-distribution visual shifts without increasing operational complexity during deployment.
To achieve this, the approach constructs two paired visual variants of each observation during training. A task-preserving view changes irrelevant factors like background textures and distractor objects while keeping the target object unchanged, whereas a task-altering view modifies the target object's position or orientation. The framework adds two auxiliary objectives to the standard reinforcement learning optimization: an invariance objective that forces action distributions to remain consistent across task-preserving views, and a sensitivity objective that forces action distributions to adapt when the target object changes. The evaluation tested two standard model architectures across pick-and-place tasks within a simulated manipulation environment against unseen textures, lighting, target poses, clutter levels, and camera viewpoints.
The analysis produced several key findings. First, the proposed framework consistently outperformed standard reinforcement learning baselines across all out-of-distribution visual test scenarios, achieving an average absolute success rate improvement of 16.62% on the flow-matching architecture and 9.10% on the autoregressive architecture. Second, the method exhibited strong visual generalization, successfully transferring robustness to unseen lighting conditions and camera angles even when those specific factors were not altered during the construction of training pairs. Third, reinforcement learning fine-tuning efficiency improved by approximately three times, reaching high success thresholds in roughly 80 training steps compared to 240 steps for the baseline. Finally, ablation tests revealed that while invariance guidance provided the primary robustness gain (accounting for a 12.61% baseline improvement), coupling it with sensitivity guidance achieved the strongest overall performance.
These findings indicate that visual robustness can be directly embedded into robotic policies during post-training optimization rather than relying on complex, latency-inducing runtime visual processing modules. Because the auxiliary objectives operate strictly during training, deployed systems run on standard architectures without incurring inference-time latency or computational overhead. This offers an efficient path to reducing failure risks and adaptation costs in automated manipulation workflows.
Decision-makers should consider adopting behavior-level auxiliary objectives when fine-tuning robotic policies to improve operational reliability and training throughput. However, because the current evaluations were conducted in simulation using exact segmentation masks, organizations should validate these gains in physical environments using off-the-shelf visual segmentation tools before full-scale deployment.
- Paper: OpenVLA: An Open-Source Vision-Language-Action Model, Moo Jin Kim et al. (2024). Introduces the open-source Vision-Language-Action architecture that serves as one of the primary base manipulation models fine-tuned and benchmarked in this work.
- Paper: π0: A Vision-Language-Action Flow Model for General Robot Control, Kevin Black et al. (2024). Presents the foundational vision-language-action flow matching model whose variant family forms the core architecture evaluated under reinforcement learning fine-tuning.
- Paper: Proximal Policy Optimization Algorithms, John Schulman et al. (2017). Establishes the Proximal Policy Optimization algorithm on top of which the paired invariance and sensitivity auxiliary objectives are integrated.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). Demonstrates the paradigm of translating web-scale vision-language models into robotic action spaces that modern VLA fine-tuning frameworks build upon.
- Paper: SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training, Tianzhe Chu et al. (2025). Analyzes how reinforcement learning post-training confers out-of-distribution visual generalization compared to supervised fine-tuning in multimodal foundation models.
No sufficiently relevant recommendations were found.
