OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction

Huang HuangFangchen LiuLetian FuTingfan WuMustafa MukadamJitendra MalikKen GoldbergPieter Abbeel

article2025ICML63 citations

Proposes a vision-language-action architecture that extracts task-relevant visual tokens aligned with text instructions using frozen pre-trained encoders, achieving superior zero-shot generalization in robot manipulation without degrading multi-modal semantic alignments through fine-tuning.

Listen

Deploying robotic systems that understand natural language instructions and adapt reliably to new environments is a central challenge in automated manipulation. Vision-Language-Action models combine visual input and text instructions to predict robotic actions. However, mainstream models typically fine-tune pre-trained vision-language models on robotic datasets or pass visual and text tokens directly to downstream control networks. Because robotic datasets lack the scale and semantic diversity of broad multimodal datasets, this fine-tuning causes models to overfit, degrading pre-trained visual-language alignments and leading to poor generalization when encountering novel objects and scenes.

The article develops and evaluates OTTER, a Vision-Language-Action model designed to retain pre-trained semantic knowledge by freezing its vision-language encoders and performing text-aware visual feature extraction. The primary objective is to demonstrate that selectively filtering visual features based on task instructions achieves superior zero-shot generalization across novel robotic manipulation tasks compared to standard fine-tuning approaches.

The researchers assessed the approach through controlled physical robot experiments on a Franka arm and simulation benchmarks using the LIBERO manipulation suite. The evaluation tested single-primitive pick-and-place actions and multi-primitive operations—including poking, pouring, and opening/closing drawers—across seen training scenarios and unseen test configurations with novel target objects and distractors. The architecture utilizes frozen visual and language components from CLIP, applying an attention mechanism to extract visual patch features aligned with task tokens before passing them alongside robot state representations into a causal transformer policy.

The empirical findings establish that OTTER markedly improves robotic control and zero-shot generalization over baseline models like Octo and OpenVLA. In physical pick-and-place tests with unseen objects, OTTER achieved a 62% success rate, which increased to 73% when pre-trained on the Open X-Embodiment dataset, whereas fine-tuned baselines achieved 12% or lower. In multi-primitive physical trials, an enlarged model variant reached a 77% overall success rate across novel tasks, maintaining strong performance on demanding primitives like pouring (77%) and poking (93%) where baseline models entirely failed. Simulation trials echoed these results, with OTTER leading unseen task benchmarks at 61% accuracy. Ablation analyses confirmed that freezing pre-trained vision features and including robot state feedback are vital, as fine-tuning the vision encoder degraded unseen task performance to 15%.

These results show that decoupling semantic task planning from low-level action control produces substantially more data-efficient and robust robotic policies. Preserving pre-trained representations eliminates the extensive data collection typically required to teach models how to identify everyday objects. For organizations building automated workflows, this design reduces operational risk and deployment costs while enabling faster adaptation to unseen items. The article also shows that system performance scales favorably when pairing larger vision-language models with expanded robotic training data.

Based on these findings, teams developing vision-language robotic systems should adopt frozen, text-aware feature extractors rather than fully fine-tuning pre-trained multimodal encoders. Next steps supported by the article include pre-training policies across broader robotic datasets and scaling visual backbones to capture larger performance gains. However, leadership should note limitations: the current architecture evaluates arms parameterized by standard spatial transformations and has not yet addressed complex robot hands or long-horizon manipulation sequences, warranting focused pilot testing before broad industrial deployment.

arXiv: 2503.03734FangchenLiu/otter_jax
Cover for OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction

Abstract

Vision-Language-Action (VLA) models aim to predict robotic actions based on visual observations and language instructions. Existing approaches require fine-tuning pre-trained vision-language models (VLMs) as visual and language features are independently fed into downstream policies, degrading the pre-trained semantic alignments. We propose OTTER, a novel VLA architecture that leverages these existing alignments through explicit, text-aware visual feature extraction. Instead of processing all visual features, OTTER selectively extracts and passes only task-relevant visual features that are semantically aligned with the language instruction to the policy transformer. This allows OTTER to keep the pre-trained vision-language encoders frozen. Thereby, OTTER preserves and utilizes the rich semantic understanding learned from large-scale pre-training, enabling strong zero-shot generalization capabilities. In simulation and real-world experiments, OTTER significantly outperforms existing VLA models, demonstrating strong zero-shot generalization to novel objects and environments. Video, code, checkpoints, and dataset: https://ottervla.github.io/.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Vision Language Pre-training
  • 2.2 Vision Language Action Models
  • 3 Method
  • 3.1 Text-Aware Visual Feature Extraction
  • 3.2 Model Architecture
  • 4 Experiments
  • 4.1 Environment Setup
  • 4.2 Baselines
  • 5 Results
  • 5.1 Real-world Experiments
  • 5.2 Simulation Experiments
  • 5.3 Ablations
  • 5.4 Scaling up Vision and Language Encoders
  • 5.4.1 Scaling up Pre-training Dataset with Human Videos
  • 6 Limitations and Conclusions
  • References
  • A Environment Setup
  • A.1 Simulation Tasks
  • A.2 Real-world Tasks
  • B Model and Training Details
  • B.1 Model Architecture for OTTER and Baselines
  • B.2 Training Hyper-parameters
  • C Vision-Language Attention Visualization
  • C.1 Pre-training on the Human Dataste
  • D More Ablations

Knowls

  1. Knowl 1 — Text-aware visual features from frozen CLIP

    model/method

    OTTER uses a frozen, pre-trained CLIP model to select visual patch features according to the language instruction. It takes visual patch features fvf_v from the final vision self-attention block’s attention output, rather than the block’s final output feature map, and per-token language features flf_l. After modality-specific projection and normalization, both lie in a shared dd-dimensional space: fv∈Rn×df_v\in\mathbb{R}^{n\times d} for nn patches and fl∈Rm×df_l\in\mathbb{R}^{m\times d} for mm language tokens. The resulting text-aware features are

    fvl=softmax⁡n ⁣(f^lf^v⊤τ)(f^v+PE),f_{vl}=\operatorname{softmax}_{n}\!\left(\frac{\hat f_l\hat f_v^{\top}}{\tau}\right)(\hat f_v+PE),

    where hats denote unit-L2L_2-normalized features, PE∈Rn×dPE\in\mathbb{R}^{n\times d} is the two-dimensional sine-cosine position embedding for the visual patches, and τ\tau is a learnable temperature clipped to the interval [0,100][0,100]. The softmax is over the nn visual patches, so each language token produces a weighted combination of patch features. The CLIP encoders remain frozen during robot-policy training; τ\tau is the only learned parameter in this feature-fusion operation.

  2. Knowl 2 — Policy architecture for combining text-aware perception and robot state

    model/method

    At each timestep, OTTER computes text-aware visual features fvlf_{vl} for each camera. A learnable cross-attention pooling layer uses four learned queries to produce four outputs, which are concatenated into one visual token per camera. A separate learnable cross-attention pooling layer compresses the language features into one text token. An MLP encodes the robot’s proprioception, and the text token, visual token(s), and embodiment feature are concatenated into the timestep representation supplied to a causal transformer. The standard policy transformer has four layers, eight attention heads, and hidden dimension 512; it uses a 12-step context in the real-robot setup. A feed-forward action head predicts future actions from each timestep representation. This architecture keeps CLIP frozen and separates language-guided selection of task-relevant visual information from downstream action prediction.

  3. Knowl 3 — Zero-shot performance across real-robot manipulation primitives

    empirical result

    In real-robot evaluation on unseen tasks across pouring, drawer manipulation, poking, and pick-and-place, OTTER-based policies substantially outperform the listed baselines. The evaluation uses 1,185 human tele-operated demonstrations across the four primitives and 150 trials of unseen tasks. Success rates are reported in the order pouring, drawer, poking, pick-and-place, then mean ± standard error:

    • π0\pi_0-Fast-Droid: 0%, 0%, 0%, 61%; mean 29% ± 3.5%.
    • Fine-tuned π0\pi_0-Fast-Droid: 0%, 45%, 27%, 51%; mean 35% ± 3.8%.
    • Fine-tuned Octo: 0%, 0%, 0%, 5%; mean 4% ± 1.2%.
    • Fine-tuned OpenVLA: 0%, 0%, 0%, 1%; mean 0.6% ± 0.5%.
    • OTTER: 63%, 50%, 93%, 61%; mean 67% ± 3.8%.
    • OTTER-L, the larger policy: 65%, 55%, 90%, 69%; mean 71% ± 3.5%.
    • OTTER-OXE, pre-trained on the Open X-Embodiment robot dataset: 60%, 65%, 93%, 66%; mean 70% ± 3.6%.
    • OTTER-OXE-L, combining Open X-Embodiment pre-training with the larger policy: 77%, 75%, 93%, 75%; mean 77% ± 3.3%.

    Octo and OpenVLA achieve no success on the unseen pouring, drawer, or poking tasks in this evaluation. The OTTER results show strong performance across all four primitives; the larger and robot-data-pre-trained variants also perform better on average than OTTER without those additions.

  4. Knowl 4 — Real-world single-primitive pick-and-place results

    empirical result

    For physical pick-and-place evaluation, models were tested on 100 trials of in-distribution training tasks and 70 trials of unseen tasks. The table below reports success rate ± standard error, in the order training tasks and unseen tasks:

    • π0\pi_0-Fast-Droid: —; 61% ± 5.3%.
    • Fine-tuned Octo: 15% ± 3.4%; 12% ± 3.6%.
    • OTTER without a pre-trained CLIP vision encoder: 17% ± 2.9%; 11% ± 2.5%.
    • Fine-tuned OpenVLA: 30% ± 3.9%; 9% ± 3.1%.
    • Direct Feature Passing OTTER (DFP-OTTER), which passes text and vision tokens independently: 29% ± 3.7%; 4% ± 1.6%.
    • OTTER with CLIP fine-tuned end-to-end: 26% ± 4.0%; 15% ± 3.9%.
    • OTTER without its embodiment feature: 40% ± 4.0%; 29% ± 4.3%.
    • OTTER without its pooled text token: 57% ± 4.4%; 53% ± 4.6%.
    • OTTER: 68% ± 4.3%; 62% ± 4.2%.
    • OTTER-OXE, pre-trained on the Open X-Embodiment dataset and fine-tuned on the pick-and-place data: 72% ± 3.9%; 73% ± 2.8%.

    OTTER’s training-task and unseen-task success rates are close, whereas several alternatives have much lower unseen-task performance. Pre-training OTTER on the larger robot dataset is associated with higher success rates on both splits.

  5. Knowl 5 — LIBERO simulation performance on in-distribution and novel tasks

    empirical result

    In simulation, OTTER was evaluated on 30 in-distribution tasks from LIBERO-Spatial, LIBERO-Object, and LIBERO-Goal, and on 10 constructed unseen tasks; each unseen task had 50 evaluation trials. The listed in-distribution metrics are success rates ± standard error for Spatial, Object, Goal, and their reported average, followed by the unseen-task rate:

    • Fine-tuned Octo: 79% ± 1.0%, 86% ± 0.9%, 85% ± 0.9%, 83% ± 1.0%; unseen 26% ± 1.1%.
    • Fine-tuned OpenVLA: 85% ± 0.9%, 88% ± 0.8%, 79% ± 1.0%, 84% ± 0.9%; unseen 48% ± 1.0%.
    • DFP-OTTER: 79% ± 0.9%, 80% ± 1.0%, 78% ± 1.1%, 79% ± 1.0%; unseen 45% ± 1.0%.
    • OTTER: 84% ± 1.0%, 89% ± 1.2%, 79% ± 0.9%, 84% ± 1.1%; unseen 61% ± 1.1%.

    The models have similar reported average performance on the in-distribution tasks, while OTTER achieves the highest rate on the constructed unseen tasks. In particular, OTTER exceeds the listed direct-feature-passing variant by 16 percentage points and fine-tuned OpenVLA by 13 percentage points on unseen tasks.

  6. Knowl 6 — Ablations isolate the contributions of pretrained vision, embodiment, and text

    empirical result

    Ablations show that components of OTTER’s feature extraction and policy input matter for generalization. On LIBERO-Object tasks, the reported in-distribution and unseen success rates, respectively, are: OTTER without CLIP vision, 80% ± 0.7% and 29% ± 0.9%; OTTER without the embodiment feature, 79% ± 0.9% and 48% ± 0.8%; OTTER without the pooled text token, 71% ± 1.2% and 49% ± 1.0%; full OTTER, 89% ± 1.2% and 61% ± 1.1%. Thus, removing pretrained CLIP vision is especially damaging on unseen tasks, while removing either the embodiment or text token also lowers performance.

    On 70 physical unseen pick-and-place trials, full OTTER achieved 62% ± 4.2% success. Two alternative fusion designs performed much worse: DFP-OTTER (CLS), which uses CLIP’s class token rather than text-aware visual extraction, achieved 6% ± 0.8%; OTTER (xattn), which uses standard cross-attention between language and vision tokens instead of the proposed extraction, achieved 2% ± 0.5%. These results support the proposed text-aware visual feature extraction over those alternatives.

  7. Knowl 7 — Robot state, action representation, and execution horizon

    model/method

    OTTER represents proprioception as a 10-dimensional vector containing absolute end-effector translation (x,y,z)(x,y,z), a six-dimensional rotation representation formed by flattening the first two rows of the end-effector’s SO(3)SO(3) rotation matrix, and a continuous gripper state. Actions are also 10-dimensional: for each future target end-effector transform TiT_i and current end-effector transform TeeT_{ee}, the predicted pose change is represented by the relative transform Tee−1TiT_{ee}^{-1}T_i, encoded as relative translation and a six-dimensional rotation representation, with the continuous absolute gripper position appended. In the real-robot configuration, the policy predicts a chunk of 12 future actions per output timestep. During execution, OTTER combines temporal ensembling with receding-horizon control; experiments selected an execution action horizon of 8 steps.

  8. Knowl 8 — Evaluation data and training protocol

    experimental setup

    The evaluation covers simulation and physical manipulation. In simulation, OTTER uses LIBERO-Spatial, LIBERO-Object, LIBERO-Goal, and LIBERO-90 data, with 50 demonstrations per simulation task. The in-distribution evaluation consists of 30 original Spatial, Object, and Goal tasks; 10 constructed unseen tasks modify instructions and objects from LIBERO-90. In the physical setting, the pick-and-place dataset DS-PnP contains 724 demonstrations across 10 tasks, and DS-ALL contains 1,185 demonstrations across pick-and-place, poking, pouring, and drawer opening/closing. The multi-primitive evaluation covers 19 in-distribution and 15 unseen tasks, with 10 trials per task and 2–3 random distractor objects in trials where objects are manipulated. The single-primitive pick-and-place comparison uses 100 training-task trials and 70 unseen-task trials.

    Training uses AdamW, learning rate 3×10−43\times10^{-4}, 2,000 warm-up steps, weight decay 0.01, cosine learning-rate decay, gradient clipping at 1, batch size 64, and 40,000 gradient steps (60,000 for the larger configuration). Input images are 224×224224\times224. Models are trained on four NVIDIA A100 80GB GPUs.

  9. Knowl 9 — Performance improves with larger CLIP vision-language encoders

    empirical result

    On physical pick-and-place tasks, OTTER was evaluated with CLIP ViT-B/32, ViT-B/16, and ViT-L/14 encoders, increasing the encoder’s inference computation. The paper reports that moving from ViT-B/32 to ViT-L/14 raises OTTER’s success rate by 27.5% on training tasks and 39.3% on unseen tasks. The result indicates that OTTER can benefit from scaling the frozen vision-language encoder, including on unseen pick-and-place tasks.

  10. Knowl 10 — Stated limitations on morphology and task complexity

    limitation

    OTTER’s robot-state and action representations rely on end-effector poses that can be parameterized with SE(3)SE(3) transforms. The authors identify this as a limitation for robots with different morphologies, particularly multi-finger hands that cannot be readily represented this way. They also state that the study does not extensively evaluate scaling to long-horizon tasks or more complex scenes, leaving those capabilities unresolved.

Coverage note — No substantial contributed material was omitted; supplementary attention-map visualizations and fine-grained image-augmentation settings are not separate knowls because they serve as supporting diagnostics or implementation details.

References

  1. 1.Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Alayrac, J.-B., Recasens, A., Schneider, R., Arandjelovic, R., Ramapuram, J., De Fauw, J., Smaira, L., Dieleman, S., and Zisserman, A. Self-supervised multimodal versatile networks. Advances in neural information processing systems, 33:25–37, 2020.
  3. 3.Ba, J. L. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  4. 4.Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966, 2023.
  5. 5.Beyer, L., Steiner, A., Pinto, A. S., Kolesnikov, A., Wang, X., Salz, D., Neumann, M., Alabdulmohsin, I., Tschannen, M., Bugliarello, E., Unterthiner, T., Keysers, D., Koppula, S., Liu, F., Grycner, A., Gritsenko, A., Houlsby, N., Kumar, M., Rong, K., Eisenschlos, J., Kabra, R., Bauer, M., Bosnjak, M., Chen, X., Minderer, M., Voigtlaender, P., Bica, I., Balazevic, I., Puigcerver, J., Papalampidi, P., Henaff, O., Xiong, X., Soricut, R., Harmsen, J., and Zhai, X. Paligemma: A versatile 3b vlm for transfer, 2024. URL https://arxiv.org/abs/2407.07726.
  6. 6.Black, K., Brown, N., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Groom, L., Hausman, K., Ichter, B., Jakubczak, S., Jones, T., Ke, L., Levine, S., Li-Bell, A., Mothukuri, M., Nair, S., Pertsch, K., Shi, L. X., Tanner, J., Vuong, Q., Walling, A., Wang, H., and Zhilinsky, U. π0: A vision-language-action flow model for general robot control, 2024. URL https://arxiv.org/abs/2410.24164.
  7. 7.Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al. Rt-1: Robotics transformer for real-world control at scale. arXiv:2212.06817, 2022.
  8. 8.Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al. RT-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818, 2023.
  9. 9.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877–1901, 2020.
  10. 10.Cherti, M., Beaumont, R., Wightman, M., Wortsman, M., Ilharco, G., Gordon, C., Schuhmann, C., Schmidt, L., and Jitsev, J. Reproducible scaling laws for contrastive language-image learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2818–2829, 2023.
  11. 11.Chi, C., Feng, S., Du, Y., Xu, Z., Cousineau, E., Burchfiel, B., and Song, S. Diffusion policy: Visuomotor policy learning via action diffusion. arXiv preprint arXiv:2303.04137, 2023.
  12. 12.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113, 2023.
  13. 13.Collaboration, E., O’Neill, A., Rehman, A., Maddukuri, A., Gupta, A., Padalkar, A., Lee, A., Pooley, A., Gupta, A., Mandlekar, A., Jain, A., Tung, A., Bewley, A., Herzog, A., Irpan, A., Khazatsky, A., Rai, A., Gupta, A., Wang, A., Kolobov, A., Singh, A., Garg, A., Kembhavi, A., Xie, A., Brohan, A., Raffin, A., Sharma, A., Yavary, A., Jain, A., Balakrishna, A., Wahid, A., Burgess-Limerick, B., Kim, B., Scholkopf, B., Wulfe, B., Ichter, B., Lu, C., Xu, C., Le, C., Finn, C., Wang, C., Xu, C., Chi, C., Huang, C., Chan, C., Agia, C., Pan, C., Fu, C., Devin, C., Xu, D., Morton, D., Driess, D., Chen, D., Pathak, D., Shah, D., Buchler, D., Jayaraman, D., Kalashnikov, D., Sadigh, D., Johns, E., Foster, E., Liu, F., Ceola, F., Xia, F., Zhao, F., Frujeri, F. V., Stulp, F., Zhou, G., Sukhatme, G. S., Salhotra, G., Yan, G., Feng, G., Schiavi, G., Berseth, G., Kahn, G., Wang, G., Su, H., Fang, H.-S., Shi, H., Bao, H., Amor, H. B., Christensen, H. I., Furuta, H., Walke, H., Fang, H., Ha, H., Mordatch, I., Radosavovic, I., Leal, I., Liang, J., Abou-Chakra, J., Kim, J., Drake, J., Peters, J., Schneider, J., Hsu, J., Bohg, J., Bingham, J., Wu, J., Gao, J., Hu, J., Wu, J., Wu, J., Sun, J., Luo, J., Gu, J., Tan, J., Oh, J., Wu, J., Lu, J., Yang, J., Malik, J., Silverio, J., Hejna, J., Booher, J., Tompson, J., Yang, J., Salvador, J., Lim, J. J., Han, J., Wang, K., Rao, K., Pertsch, K., Hausman, K., Go, K., Gopalakrishnan, K., Goldberg, K., Byrne, K., Oslund, K., Kawaharazuka, K., Black, K., Lin, K., Zhang, K., Ehsani, K., Lekkala, K., Ellis, K., Rana, K., Srinivasan, K., Fang, K., Singh, K. P., Zeng, K.-H., Hatch, K., Hsu, K., Itti, L., Chen, L. Y., Pinto, L., Fei-Fei, L., Tan, L., Fan, L. J., Ott, L., Lee, L., Weihs, L., Chen, M., Lepert, M., Memmel, M., Tomizuka, M., Itkina, M., Castro, M. G., Spero, M., Du, M., Ahn, M., Yip, M. C., Zhang, M., Ding, M., Heo, M., Srirama, M. K., Sharma, M., Kim, M. J., Kanazawa, N., Hansen, N., Heess, N., Joshi, N. J., Suenderhauf, N., Liu, N., Palo, N. D., Shafiullah, N. M. M., Mees, O., Kroemer, O., Bastani, O., Sanketi, P. R., Miller, P. T., Yin, P., Wohlhart, P., Xu, P., Fagan, P. D., Mitrano, P., Sermanet, P., Abbeel, P., Sundaresan, P., Chen, Q., Vuong, Q., Rafailov, R., Tian, R., Doshi, R., Martin-Martin, R., Baijal, R., Scalise, R., Hendrix, R., Lin, R., Qian, R., Zhang, R., Mendonca, R., Shah, R., Hoque, R., Julian, R., Bustamante, S., Kirmani, S., Levine, S., Lin, S., Moore, S., Bahl, S., Dass, S., Sonawani, S., Song, S., Xu, S., Haldar, S., Karamcheti, S., Adebola, S., Guist, S., Nasiriany, S., Schaal, S., Welker, S., Tian, S., Ramamoorthy, S., Dasari, S., Belkhale, S., Park, S., Nair, S., Mirchandani, S., Osa, T., Gupta, T., Harada, T., Matsushima, T., Xiao, T., Kollar, T., Yu, T., Ding, T., Davchev, T., Zhao, T. Z., Armstrong, T., Darrell, T., Chung, T., Jain, V., Vanhoucke, V., Zhan, W., Zhou, W., Burgard, W., Chen, X., Chen, X., Wang, X., Zhu, X., Geng, X., Liu, X., Liangwei, X., Li, X., Pang, Y., Lu, Y., Ma, Y. J., Kim, Y., Chebotar, Y., Zhou, Y., Zhu, Y., Wu, Y., Xu, Y., Wang, Y., Bisk, Y., Cho, Y., Lee, Y., Cui, Y., Cao, Y., Wu, Y.-H., Tang, Y., Zhu, Y., Zhang, Y., Jiang, Y., Li, Y., Li, Y., Iwasawa, Y., Matsuo, Y., Ma, Z., Xu, Z., Cui, Z. J., Zhang, Z., Fu, Z., and Lin, Z. Open x-embodiment: Robotic learning datasets and rt-x models, 2024.
  14. 14.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  15. 15.Dong, X., Bao, J., Zheng, Y., Zhang, T., Chen, D., Yang, H., Zeng, M., Zhang, W., Yuan, L., Chen, D., et al. Maskclip: Masked self-distillation advances contrastive language-image pretraining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10995–11005, 2023.
  16. 16.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2020.
  17. 17.Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.
  18. 18.Fu, L., Huang, H., Datta, G., Chen, L. Y., Panitch, W. C.-H., Liu, F., Li, H., and Goldberg, K. In-context imitation learning via next-token prediction. arXiv preprint arXiv:2408.15980, 2024.
  19. 19.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In CVPR, 2016.
  20. 20.Jang, E., Irpan, A., Khansari, M., Kappler, D., Ebert, F., Lynch, C., Levine, S., and Finn, C. Bc-z: Zero-shot task generalization with robotic imitation learning. In Conference on Robot Learning, 2022.
  21. 21.Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q., Sung, Y.-H., Li, Z., and Duerig, T. Scaling up visual and vision-language representation learning with noisy text supervision. In International conference on machine learning, pp. 4904–4916. PMLR, 2021.
  22. 22.Jiang, Y., Gupta, A., Zhang, Z., Wang, G., Dou, Y., Chen, Y., Fei-Fei, L., Anandkumar, A., Zhu, Y., and Fan, L. VIMA: General robot manipulation with multimodal prompts. International Conference on Machine Learning (ICML), 2023.
  23. 23.Karamcheti, S., Nair, S., Balakrishna, A., Liang, P., Kollar, T., and Sadigh, D. Prismatic vlms: Investigating the design space of visually-conditioned language models. arXiv preprint arXiv:2402.07865, 2024.
  24. 24.Kerr, J., Kim, C. M., Goldberg, K., Kanazawa, A., and Tancik, M. Lerf: Language embedded radiance fields. In International Conference on Computer Vision (ICCV), 2023.
  25. 25.Khazatsky, A., Pertsch, K., Nair, S., Balakrishna, A., Dasari, S., Karamcheti, S., Nasiriany, S., Srirama, M. K., Chen, L. Y., Ellis, K., Fagan, P. D., Hejna, J., Itkina, M., Lepert, M., Ma, Y. J., Miller, P. T., Wu, J., Belkhale, S., Dass, S., Ha, H., Jain, A., Lee, A., Lee, Y., Memmel, M., Park, S., Radosavovic, I., Wang, K., Zhan, A., Black, K., Chi, C., Hatch, K. B., Lin, S., Lu, J., Mercat, J., Rehman, A., Sanketi, P. R., Sharma, A., Simpson, C., Vuong, Q., Walke, H. R., Wulfe, B., Xiao, T., Yang, J. H., Yavary, A., Zhao, T. Z., Agia, C., Baijal, R., Castro, M. G., Chen, D., Chen, Q., Chung, T., Drake, J., Foster, E. P., Gao, J., Herrera, D. A., Heo, M., Hsu, K., Hu, J., Jackson, D., Le, C., Li, Y., Lin, K., Lin, R., Ma, Z., Maddukuri, A., Mirchandani, S., Morton, D., Nguyen, T., O’Neill, A., Scalise, R., Seale, D., Son, V., Tian, S., Tran, E., Wang, A. E., Wu, Y., Xie, A., Yang, J., Yin, P., Zhang, Y., Bastani, O., Berseth, G., Bohg, J., Goldberg, K., Gupta, A., Gupta, A., Jayaraman, D., Lim, J. J., Malik, J., Martin-Martin, R., Ramamoorthy, S., Sadigh, D., Song, S., Wu, J., Yip, M. C., Zhu, Y., Kollar, T., Levine, S., and Finn, C. Droid: A large-scale in-the-wild robot manipulation dataset, 2024.
  26. 26.Kim, M. J., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E., Lam, G., Sanketi, P., Vuong, Q., Kollar, T., Burchfiel, B., Tedrake, R., Sadigh, D., Levine, S., Liang, P., and Finn, C. Openvla: An open-source vision-language-action model, 2024. URL https://arxiv.org/abs/2406.09246.
  27. 27.Lan, M., Chen, C., Ke, Y., Wang, X., Feng, L., and Zhang, W. Clearclip: Decomposing clip representations for dense vision-language inference. In ECCV, 2024.
  28. 28.Laurençon, H., Tronchon, L., Cord, M., and Sanh, V. What matters when building vision-language models? arXiv preprint arXiv:2405.02246, 2024.
  29. 29.Li, J., Li, D., Savarese, S., and Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023.
  30. 30.Liu, B., Zhu, Y., Gao, C., Feng, Y., Liu, Q., Zhu, Y., and Stone, P. Libero: Benchmarking knowledge transfer for lifelong robot learning. Advances in Neural Information Processing Systems, 36, 2024.
  31. 31.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. In NeurIPS, 2023.
  32. 32.Octo Model Team, Ghosh, D., Walke, H., Pertsch, K., Black, K., Mees, O., Dasari, S., Hejna, J., Xu, C., Luo, J., Kreiman, T., Tan, Y., Chen, L. Y., Sanketi, P., Vuong, Q., Xiao, T., Sadigh, D., Finn, C., and Levine, S. Octo: An open-source generalist robot policy. In Proceedings of Robotics: Science and Systems, Delft, Netherlands, 2024.
  33. 33.OpenAI. Gpt-4o system card. https://cdn.openai.com/gpt-4o-system-card.pdf, 2024. Accessed: 2024-09-14.
  34. 34.Perez, E., Strub, F., De Vries, H., Dumoulin, V., and Courville, A. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI conference on artificial intelligence, volume 32, 2018.
  35. 35.Pertsch, K., Stachowicz, K., Ichter, B., Driess, D., Nair, S., Vuong, Q., Mees, O., Finn, C., and Levine, S. Fast: Efficient action tokenization for vision-language-action models, 2025. URL https://arxiv.org/abs/2501.09747.
  36. 36.Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al. Improving language understanding by generative pre-training. 2018.
  37. 37.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI Blog, 1(8):9, 2019.
  38. 38.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PMLR, 2021.
  39. 39.Rao, Y., Zhao, W., Chen, G., Tang, Y., Zhu, Z., Huang, G., Zhou, J., and Lu, J. Denseclip: Language-guided dense prediction with context-aware prompting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  40. 40.Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al. A generalist agent. arXiv preprint arXiv:2205.06175, 2022.
  41. 41.Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35:25278–25294, 2022.
  42. 42.Shah, D., Sridhar, A., Dashora, N., Stachowicz, K., Black, K., Hirose, N., and Levine, S. ViNT: A Foundation Model for Visual Navigation. In 7th Annual Conference on Robot Learning (CoRL), 2023.
  43. 43.Yao, L., Huang, R., Hou, L., Lu, G., Niu, M., Xu, H., Liang, X., Li, Z., Jiang, X., and Xu, C. Filip: Fine-grained interactive language-image pre-training. arXiv preprint arXiv:2111.07783, 2021.
  44. 44.Yuan, L., Chen, D., Chen, Y.-L., Codella, N., Dai, X., Gao, J., Hu, H., Huang, X., Li, B., Li, C., et al. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432, 2021.
  45. 45.Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 11975–11986, 2023.
  46. 46.Zhao, T. Z., Kumar, V., Levine, S., and Finn, C. Learning fine-grained bimanual manipulation with low-cost hardware. arXiv preprint arXiv:2304.13705, 2023.

Citation

MLA
Huang, H., et al. “OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction”. arXiv, 2025, https://doi.org/10.48550/arxiv.2503.03734.
APA
Huang, H., Liu, F., Fu, L., Wu, T., Mukadam, M., Malik, J., Goldberg, K., & Abbeel, P. (2025). OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction. arXiv. https://doi.org/10.48550/arxiv.2503.03734
Chicago
Huang, H., F. Liu, L. Fu, et al. 2025. “OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2503.03734.
Harvard
Huang, H. et al. (2025) “OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction”. arXiv. Available at: https://doi.org/10.48550/arxiv.2503.03734.
Vancouver
1. Huang H, Liu F, Fu L, Wu T, Mukadam M, Malik J, Goldberg K, Abbeel P (2025) OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction. https://doi.org/10.48550/arxiv.2503.03734

BibTeX

@misc{https://doi.org/10.48550/arxiv.2503.03734,
  doi = {10.48550/ARXIV.2503.03734},
  url = {https://arxiv.org/abs/2503.03734},
  author = {Huang, Huang and Liu, Fangchen and Fu, Letian and Wu, Tingfan and Mukadam, Mustafa and Malik, Jitendra and Goldberg, Ken and Abbeel, Pieter},
  keywords = {Robotics (cs.RO), Computer Vision and Pattern Recognition (cs.CV), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/