OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
Huang HuangFangchen LiuLetian FuTingfan WuMustafa MukadamJitendra MalikKen GoldbergPieter Abbeel
Proposes a vision-language-action architecture that extracts task-relevant visual tokens aligned with text instructions using frozen pre-trained encoders, achieving superior zero-shot generalization in robot manipulation without degrading multi-modal semantic alignments through fine-tuning.
Deploying robotic systems that understand natural language instructions and adapt reliably to new environments is a central challenge in automated manipulation. Vision-Language-Action models combine visual input and text instructions to predict robotic actions. However, mainstream models typically fine-tune pre-trained vision-language models on robotic datasets or pass visual and text tokens directly to downstream control networks. Because robotic datasets lack the scale and semantic diversity of broad multimodal datasets, this fine-tuning causes models to overfit, degrading pre-trained visual-language alignments and leading to poor generalization when encountering novel objects and scenes.
The article develops and evaluates OTTER, a Vision-Language-Action model designed to retain pre-trained semantic knowledge by freezing its vision-language encoders and performing text-aware visual feature extraction. The primary objective is to demonstrate that selectively filtering visual features based on task instructions achieves superior zero-shot generalization across novel robotic manipulation tasks compared to standard fine-tuning approaches.
The researchers assessed the approach through controlled physical robot experiments on a Franka arm and simulation benchmarks using the LIBERO manipulation suite. The evaluation tested single-primitive pick-and-place actions and multi-primitive operations—including poking, pouring, and opening/closing drawers—across seen training scenarios and unseen test configurations with novel target objects and distractors. The architecture utilizes frozen visual and language components from CLIP, applying an attention mechanism to extract visual patch features aligned with task tokens before passing them alongside robot state representations into a causal transformer policy.
The empirical findings establish that OTTER markedly improves robotic control and zero-shot generalization over baseline models like Octo and OpenVLA. In physical pick-and-place tests with unseen objects, OTTER achieved a 62% success rate, which increased to 73% when pre-trained on the Open X-Embodiment dataset, whereas fine-tuned baselines achieved 12% or lower. In multi-primitive physical trials, an enlarged model variant reached a 77% overall success rate across novel tasks, maintaining strong performance on demanding primitives like pouring (77%) and poking (93%) where baseline models entirely failed. Simulation trials echoed these results, with OTTER leading unseen task benchmarks at 61% accuracy. Ablation analyses confirmed that freezing pre-trained vision features and including robot state feedback are vital, as fine-tuning the vision encoder degraded unseen task performance to 15%.
These results show that decoupling semantic task planning from low-level action control produces substantially more data-efficient and robust robotic policies. Preserving pre-trained representations eliminates the extensive data collection typically required to teach models how to identify everyday objects. For organizations building automated workflows, this design reduces operational risk and deployment costs while enabling faster adaptation to unseen items. The article also shows that system performance scales favorably when pairing larger vision-language models with expanded robotic training data.
Based on these findings, teams developing vision-language robotic systems should adopt frozen, text-aware feature extractors rather than fully fine-tuning pre-trained multimodal encoders. Next steps supported by the article include pre-training policies across broader robotic datasets and scaling visual backbones to capture larger performance gains. However, leadership should note limitations: the current architecture evaluates arms parameterized by standard spatial transformations and has not yet addressed complex robot hands or long-horizon manipulation sequences, warranting focused pilot testing before broad industrial deployment.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). RT-2 establishes the vision-language-action approach of transferring pretrained visual and linguistic knowledge into robot control, framing OTTER’s effort to preserve that knowledge while improving generalization.
- Paper: Octo: An Open-Source Generalist Robot Policy, O. Team et al. (2024). Octo provides a key generalist-policy baseline and illustrates how multimodal inputs and large-scale robot demonstrations are combined in the control architectures OTTER evaluates.
- Paper: OpenVLA: An Open-Source Vision-Language-Action Model, Moo Jin Kim et al. (2024). OpenVLA clarifies the fine-tuned, pretrained-backbone approach that OTTER contrasts with its frozen, text-aware feature extraction.
- Paper: TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers, Bin Yu et al. (2026). TwinBrainVLA carries OTTER’s strategy of preserving pretrained vision-language knowledge into a dual-pathway architecture that lets a trainable control pathway query a frozen semantic model.
