UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Jianke ZhangYanjiang GuoYucheng HuXiaoyu ChenXiang ZhuJianyu Chen

article2025ICML105 citations

Proposes a unified training paradigm for Vision-Language-Action models that couples multi-modal understanding with future visual prediction to capture both high-level semantics and low-level spatial details, boosting robot manipulation success on the Calvin benchmark by 33%.

Listen

Generalist robotic manipulation increasingly relies on vision-language-action models that transfer broad semantic knowledge from pretrained vision-language models into robotic control policies. However, while vision-language models excel at high-level reasoning and understanding linguistic concepts, they often struggle with low-level visual details, precise spatial relationships, and physical dynamics. These spatial and physical dynamics are essential for fine-grained manipulation tasks, creating a critical performance bottleneck for current robotic policies.

The main objective of the article is to demonstrate and evaluate UP-VLA, a unified vision-language-action framework designed to overcome these limitations. The approach integrates high-level multi-modal understanding with low-level future visual prediction within a single autoregressive model to optimize robotic action planning.

To establish a balanced training pipeline, the researchers built their model on a 1.3-billion-parameter foundation and implemented a two-stage training approach. In the first phase, they pre-trained the model on a mix of 665,000 image-text pairs to preserve semantic reasoning alongside 25,000 robotic demonstrations to learn visual prediction. In the second phase, they fine-tuned the model on downstream manipulation tasks, prompting the system to generate actions alongside scene descriptions and predicted future visual frames. The authors evaluated the system across standard simulated multi-task manipulation benchmarks as well as real-world robotic arm experiments spanning over 2,000 demonstrations of basic and complex tabletop skills.

The evaluation yielded several key findings. First, on the standard long-horizon simulation benchmark, UP-VLA achieved an average task completion length of 4.08 consecutive tasks, outperforming the previous state of the art by approximately 33%. Second, in real-world evaluations on unseen objects and fine-grained spatial tasks—such as cable routing and picking up small items—the proposed model achieved higher success rates than pure prediction models and conventional vision-language policies. Third, ablation experiments revealed that removing visual prediction caused performance on simulation tasks to drop sharply from 4.08 to 1.44, while omitting multi-modal understanding significantly reduced the robot's success rate on previously unseen objects from 58% to 20%.

These results indicate that combining generative visual prediction with language understanding allows embodied agents to overcome the historical trade-off between broad semantic generalization and physical precision. In practice, this dual-capability architecture improves task success and operational reliability in unfamiliar environments, reducing the risk of failure when deploying automated robotic systems in dynamic real-world settings.

Stakeholders and engineering teams developing embodied automation should consider adopting unified pre-training objectives that integrate visual future prediction into language-action models. For future development, the authors suggest exploring broader datasets and larger model backbones, noting that current limitations include occasional object misidentifications and background visual artifacts in newly generated frames. Overall, the findings demonstrate a high level of confidence in the effectiveness of joint visual prediction and semantic understanding for robotic control.

arXiv: 2501.18867
Cover for UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Preliminaries
  • 4. Methodology
  • 4.1. Backbone
  • 4.2. Bridging Visual Prediction and Multi-modal Understanding
  • 4.3. Enhancing Action Learning with Joint Prediction and Understanding
  • 4.4. Training Strategy
  • 4.4.1. TRAINING PIPELINE
  • 4.4.2. TRAINING OBJECTIVE
  • 5. Experiments
  • 5.1. Experiment Setup and baseline
  • 5.2. Simulation Evaluation
  • 5.3. Real Robot Evaluation
  • 5.4. Ablation Studies
  • 5.5. Quantitative Results
  • 6. Conclusion
  • Impact Statement
  • References
  • A. Implementation Details
  • B. Manipulation Dataset Details

Knowls

  1. Knowl 1 — UP-VLA Unified Architecture for Multi-Modal Understanding, Prediction, and Action

    model/method

    UP-VLA is a Vision-Language-Action architecture that unifies multi-modal understanding, future visual prediction, and robotic action generation within a single autoregressive transformer backbone based on Show-o (utilizing the 1.3B parameter Phi-1.5 language model).

    The model incorporates two complementary vision pipelines alongside an action generation head:

    1. Continuous Vision Pipeline for Multi-Modal Understanding (MMU): Visual observations II are encoded via a continuous vision encoder (CLIP-ViT) and projected through an MLP connection layer E1(I)E_1(I) to produce continuous visual embedding tokens u={ui}i=1Mu = \{u_i\}_{i=1}^M aligned with the language embedding space.

    2. Discrete Vision Pipeline for Future Visual Prediction (PRE): Input observations OtO_t are encoded into discrete image tokens vt={vi}i=1Mv_t = \{v_i\}_{i=1}^M via a discrete VQ-GAN / MagVIT encoder E2(Ot)E_2(O_t). The model autoregressively or directly predicts discrete future tokens v^t+Δt\hat{v}_{t+\Delta t}, which can be reconstructed into predicted future frames O^t+Δt\hat{O}_{t+\Delta t} using the discrete VQ-GAN decoder.

    3. Action Policy Head: The final layer hidden representations at designated action token positions, denoted A^t:t+Δt\hat{A}_{t:t+\Delta t}, are pooled using a single-layer Multi-Head Attention Pooling (MAP) module and projected via an MLP to generate end-effector control actions: a^t:t+Δt=MLP(MAP(A^t:t+Δt))\hat{a}_{t:t+\Delta t} = \text{MLP}(\text{MAP}(\hat{A}_{t:t+\Delta t})) Actions encompass relative end-effector displacements a^pos\hat{a}_{\text{pos}} and discrete gripper open/close status a^end\hat{a}_{\text{end}}.

  2. Knowl 2 — Action Learning via Joint Prediction and Self-Generated Scene Conditioning

    model/method

    To combine semantic scene understanding with low-level spatial-physical dynamics, UP-VLA executes action prediction by jointly generating future visual frames and conditioning on self-generated scene captions.

    Given human task instruction LL and current visual observation OtO_t, the input instruction is augmented using the model's own multi-modal understanding capability: L′=[E1(Ot′), πθMMU(Ot,Lprompt), L]L' = \left[ E_1(O'_t), \, \pi_\theta^{\text{MMU}}(O_t, L_{\text{prompt}}), \, L \right] where E1(Ot′)=MLP(ViT(Ot′))E_1(O'_t) = \text{MLP}(\text{ViT}(O'_t)) provides continuous visual token embeddings, LpromptL_{\text{prompt}} is a descriptive prompt (e.g., "describe this image"), and πθMMU(Ot,Lprompt)\pi_\theta^{\text{MMU}}(O_t, L_{\text{prompt}}) represents the text description of the scene generated by the model.

    Conditioned on the observation OtO_t and expanded prompt L′L', the model jointly predicts future visual frames and action features: (O^t+Δt,A^t:t+Δt)=πθPRE(Ot,L′)(\hat{O}_{t+\Delta t}, \hat{A}_{t:t+\Delta t}) = \pi_\theta^{\text{PRE}}(O_t, L') where A^t:t+Δt\hat{A}_{t:t+\Delta t} is then mapped by an attention pooling and linear layer policy head into low-level continuous actions a^t:t+Δt\hat{a}_{t:t+\Delta t}.

  3. Knowl 3 — Unified Prompting and Flexible Attention Mechanism

    model/method

    UP-VLA formats multi-modal understanding, future visual prediction, and robot action learning into unified token sequences using special delimiter tokens and task-specific attention structures:

    • Special Tokens: Sequences use SOI/EOI\text{SOI}/\text{EOI} for start/end of continuous image tokens, SOT/EOT\text{SOT}/\text{EOT} for start/end of text, SOIdisc/EOIdisc\text{SOI}_{\text{disc}}/\text{EOI}_{\text{disc}} for discrete image tokens, and SOA/EOA\text{SOA}/\text{EOA} for action tokens.

    • Multi-Modal Understanding (MMU): Formatted as [SOI,u1,…,uM,EOI,SOT,l1,…,lN,EOT][\text{SOI}, u_1, \dots, u_M, \text{EOI}, \text{SOT}, l_1, \dots, l_N, \text{EOT}]. Continuous image tokens precede text tokens. Full bidirectional self-attention is permitted among image tokens so they attend to each other, followed by standard causal attention over text tokens to predict the next token.

    • Future Visual Prediction (PRE): Formatted as [SOT,l1,…,lN,EOT,SOIdisc,v1,…,vM,EOIdisc][\text{SOT}, l_1, \dots, l_N, \text{EOT}, \text{SOI}_{\text{disc}}, v_1, \dots, v_M, \text{EOI}_{\text{disc}}] with special task identifier PRE\text{PRE}. Discrete image tokens are positioned after language instruction tokens, allowing them to attend to all prompt tokens. Future image tokens v1:M′v'_{1:M} are directly predicted at the corresponding positional slots.

    • Joint Prediction-Action Learning (ACT): Concatenates prompt tokens, continuous image embeddings, generated descriptive text, discrete observation tokens, and action tokens [SOA,A1,…,AK,EOA][\text{SOA}, A_1, \dots, A_K, \text{EOA}], enabling actions to attend simultaneously to high-level semantic descriptions and fine-grained visual prediction representations.

  4. Knowl 4 — Multi-Task Training Loss of UP-VLA

    equation

    The complete training objective of UP-VLA combines multi-modal understanding language modeling loss LMMU\mathcal{L}_{\text{MMU}}, future image discrete cross-entropy loss LPRE\mathcal{L}_{\text{PRE}}, and embodied action regression loss LACT\mathcal{L}_{\text{ACT}}: L=λ1LMMU+λ2LPRE+λ3LACT\mathcal{L} = \lambda_1 \mathcal{L}_{\text{MMU}} + \lambda_2 \mathcal{L}_{\text{PRE}} + \lambda_3 \mathcal{L}_{\text{ACT}} where λ1,λ2,λ3\lambda_1, \lambda_2, \lambda_3 are weighting hyperparameters.

    The constituent losses are formulated as:

    1. Language Modeling for Multi-Modal Understanding: LMMU=∑i=1Nlog⁡pθ(li∣u,l1,…,li−1)\mathcal{L}_{\text{MMU}} = \sum_{i=1}^N \log p_\theta(l_i \mid u, l_1, \dots, l_{i-1}) where u={ui}i=1Mu = \{u_i\}_{i=1}^M are continuous visual tokens and l={li}i=1Nl = \{l_i\}_{i=1}^N are text tokens.

    2. Image Modeling for Future Visual Prediction: LPRE=∑j=1Mlog⁡pθ(vj′∣l,v1,…,vj,…,vM)\mathcal{L}_{\text{PRE}} = \sum_{j=1}^M \log p_\theta(v'_j \mid l, v_1, \dots, v_j, \dots, v_M) where v={vi}i=1Mv = \{v_i\}_{i=1}^M are input discrete image tokens, v′={vj′}j=1Mv' = \{v'_j\}_{j=1}^M are target future discrete image tokens, and ll denotes instruction tokens.

    3. Action Modeling for Embodied Control: LACT=∑(∥a^pos−apos∥22+BCE(a^end,aend))\mathcal{L}_{\text{ACT}} = \sum \left( \|\hat{a}_{\text{pos}} - a_{\text{pos}}\|_2^2 + \text{BCE}(\hat{a}_{\text{end}}, a_{\text{end}}) \right) where a^pos\hat{a}_{\text{pos}} and aposa_{\text{pos}} denote predicted and ground-truth continuous relative end-effector poses, and BCE(a^end,aend)\text{BCE}(\hat{a}_{\text{end}}, a_{\text{end}}) is binary cross-entropy on gripper open/close status.

  5. Knowl 5 — Two-Stage Pre-training and Action Fine-Tuning Pipeline

    model/method

    UP-VLA utilizes a two-stage training scheme while keeping the continuous CLIP-ViT and discrete VQ-GAN encoders frozen and fine-tuning the transformer backbone:

    1. Prediction and Understanding Pre-training Stage:

      • Initialized from Show-o (1.3B parameters, Phi-1.5 backbone).
      • Pre-trained for 20,000 steps with a batch size of 64 and a 1,000-step linear learning rate warmup.
      • Co-trains across two datasets: the Bridge dataset (25,000 robotic manipulation demonstrations) for future visual prediction, and LLaVA-tuning-665k (665,000 image-text pairs) for multi-modal vision-language understanding.
    2. Prediction with Action Tuning Stage:

      • Fine-tuned on downstream robot manipulation datasets with a batch size of 64 using the joint prediction-and-understanding action learning mechanism.
      • Continues co-training on image-text pairs alongside robot trajectory data to prevent catastrophic forgetting of multi-modal semantic comprehension capabilities.
  6. Knowl 6 — Zero-Shot Long-Horizon Evaluation on the CALVIN Benchmark

    data/table

    UP-VLA was evaluated on the CALVIN benchmark for zero-shot long-horizon language-conditioned robot manipulation. Policies are evaluated on sequential execution of up to 5 chained sub-tasks in an unseen environment D after training on environments A, B, and C (ABC →\rightarrow D), as well as in-distribution evaluation where training includes all environments (ABCD →\rightarrow D).

    Method Type Tasks Completed in a Row Avg. Len ↑\uparrow
    1 2 3 4 5
    CALVIN ABC →\rightarrow D (Unseen Environment D)
    RT-1 other 0.533 0.222 0.094 0.038 0.013 0.90
    Diffusion Policy other 0.402 0.123 0.026 0.008 0.000 0.56
    3D Diffuser Actor other 0.938 0.803 0.662 0.533 0.412 3.35
    3D-VLA VLA 0.447 0.163 0.081 0.016 0.000 0.71
    UP-VLA-RT-2 VLA 0.612 0.389 0.236 0.138 0.062 1.44
    Robo-Flamingo VLA 0.824 0.619 0.466 0.331 0.235 2.47
    Uni-Pi Prediction 0.560 0.160 0.080 0.080 0.040 0.92
    Susie Prediction 0.870 0.690 0.490 0.380 0.260 2.69
    GR-1 Prediction 0.854 0.712 0.596 0.497 0.401 3.06
    UP-VLA-phi-w/o-mmu Prediction 0.844 0.705 0.604 0.520 0.430 3.13
    UP-VLA (Ours) Prediction VLA 0.928 0.865 0.815 0.769 0.699 4.08
    CALVIN ABCD →\rightarrow D (In-Distribution Scene D)
    RT-1 other 0.844 0.617 0.438 0.323 0.227 2.45
    Robo-Flamingo VLA 0.964 0.896 0.824 0.740 0.660 4.09
    GR-1 Prediction 0.949 0.896 0.844 0.789 0.731 4.21
    UP-VLA (Ours) Prediction VLA 0.962 0.921 0.879 0.842 0.812 4.42

    UP-VLA achieves an average completion length of 4.08 on the ABC →\rightarrow D generalization split (a 33.3% relative improvement over GR-1's 3.06 and outperforming 3D Diffuser Actor's 3.35) and 4.42 on the in-distribution ABCD →\rightarrow D split.

  7. Knowl 7 — Real-World Robot Manipulation Evaluation

    empirical result

    UP-VLA was evaluated on a physical Franka-Emika Panda manipulator across 6 tabletop manipulation skills (picking, placing, routing cables, pressing buttons, opening drawers) trained on >2k demonstrations and evaluated over 20 rollout attempts per task:

    • Seen Tasks: UP-VLA achieved an 80% success rate, outperforming UP-VLA-RT-2 (65%), RT-1 (40%), and Diffusion Policy (30%), while matching UP-VLA-phi-w/o-mmu (75%).

    • Unseen Tasks (New Objects / Semantic Generalization): When tested on grasping completely new objects (e.g., unseen vegetable shapes, arrow-shaped paper, toy pizza, novel colored blocks), UP-VLA attained a 60% success rate, exceeding UP-VLA-RT-2 (50%), UP-VLA-phi-w/o-mmu (30%), RT-1 (10%), and Diffusion Policy (5%). Multi-modal understanding pretraining provided substantial benefit for semantic object transfer.

    • Unseen Tasks (Precise Operations / Spatial Details): On fine-grained manipulation tasks requiring precise spatial localization (cable routing, picking small blocks, picking up a pen), UP-VLA achieved a 55% success rate, outperforming UP-VLA-phi-w/o-mmu (35%), UP-VLA-RT-2 (20%), RT-1 (5%), and Diffusion Policy (0%). Visual future prediction provided critical spatial detail modeling for precise physical interactions.

  8. Knowl 8 — Ablation of Understanding and Prediction Components

    data/table

    Ablation experiments quantify the relative impact of multi-modal understanding (MMU), future visual prediction pretraining on Bridge data, visual prediction co-generation, and self-generated MMU scene description conditioning.

    Method CALVIN ABC →\rightarrow D Real World
    Avg. Len ↑\uparrow Seen ↑\uparrow Unseen ↑\uparrow
    w/o MMU 3.89 0.85 0.20
    w/o Bridge-Pretrain 2.74 0.65 0.30
    w/o Prediction 1.44 0.65 0.35
    w/o MMU-Condition 3.99 0.80 0.50
    Full UP-VLA 4.08 0.80 0.58

    Key findings include:

    1. Removing the future visual prediction objective (w/o Prediction) causes the largest drop in CALVIN completion length from 4.08 to 1.44, indicating visual dynamics supervision is critical for policy generalization.
    2. Removing Bridge dataset pretraining (w/o Bridge-Pretrain) degrades CALVIN completion length to 2.74 and real-world unseen performance to 0.30, demonstrating the necessity of video pretraining for physical dynamics modeling.
    3. Removing multi-modal understanding pretraining (w/o MMU) causes a severe drop in real-world unseen task success from 0.58 to 0.20, showing that MMU prevents overfitting and enables semantic zero-shot transfer.
    4. Removing MMU scene description prompt expansion (w/o MMU-Condition) reduces real-world unseen task success from 0.58 to 0.50.
  9. Knowl 9 — Visual Prediction Artifacts and Object Grounding Limitations

    limitation

    UP-VLA displays two primary limitations during qualitative and quantitative evaluations:

    1. Visual Generation Domain Artifacts: When predicting future frames in unseen environments (such as CALVIN test scene D), the generated discrete image tokens occasionally reproduce the background color schemes of training environments (scenes ABC) rather than preserving the input scene's true background. This is caused by limited scale and diversity during visual generative pre-training.

    2. Fine-Grained Semantic Object Identification: During Visual Question Answering and descriptive prompt generation, the model occasionally misidentifies specific objects in heavily cluttered scenes, reflecting capacity and pretraining data bounds of the 1.3B parameter model backbone.

Coverage note — None was omitted; all contributed methodology, training pipelines, loss functions, benchmark data, real-world evaluations, ablations, and stated limitations are included.

References

  1. 1.Balazadeh, V., Ataei, M., Cheong, H., Khasahmadi, A. H., and Krishnan, R. G. Synthetic vision: Training vision-language models to understand physics. arXiv preprint arXiv:2412.08619, 2024.
  2. 2.Black, K., Nakamoto, M., Atreya, P., Walke, H., Finn, C., Kumar, A., and Levine, S. Zero-shot robotic manipulation with pretrained image-editing diffusion models. arXiv preprint arXiv:2310.10639, 2023.
  3. 3.Black, K., Brown, N., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Groom, L., Hausman, K., Ichter, B., et al. pi 0: A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164, 2024.
  4. 4.Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023.
  5. 5.Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817, 2022.
  6. 6.Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818, 2023.
  7. 7.Chen, B., Xu, Z., Kirmani, S., Ichter, B., Sadigh, D., Guibas, L., and Xia, F. Spatialvlm: Endowing vision-language models with spatial reasoning capabilities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14455–14465, 2024a.
  8. 8.Chen, X., Guo, J., He, T., Zhang, C., Zhang, P., Yang, D. C., Zhao, L., and Bian, J. Igor: Image-goal representations are the atomic control units for foundation models in embodied ai, 2024b. URL https://arxiv.org/abs/2411.00785.
  9. 9.Chi, C., Feng, S., Du, Y., Xu, Z., Cousineau, E., Burchfiel, B., and Song, S. Diffusion policy: Visuomotor policy learning via action diffusion. In Proceedings of Robotics: Science and Systems (RSS), 2023.
  10. 10.Cui, C., Ding, P., Song, W., Bai, S., Tong, X., Ge, Z., Suo, R., Zhou, W., Liu, Y., Jia, B., et al. Openhelix: A short survey, empirical analysis, and open-source dual-system vla model for robotic manipulation. arXiv preprint arXiv:2505.03912, 2025.
  11. 11.Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P. N., and Hoi, S. Instructblip: Towards general-purpose vision-language models with instruction tuning. Advances in Neural Information Processing Systems, 36, 2024.
  12. 12.Ding, P., Zhao, H., Zhang, W., Song, W., Zhang, M., Huang, S., Yang, N., and Wang, D. Quar-vla: Vision-language-action model for quadruped robots. In European Conference on Computer Vision, pp. 352–367. Springer, 2024.
  13. 13.Ding, P., Ma, J., Tong, X., Zou, B., Luo, X., Fan, Y., Wang, T., Lu, H., Mo, P., Liu, J., et al. Humanoid-vla: Towards universal humanoid control with visual integration. arXiv preprint arXiv:2502.14795, 2025.
  14. 14.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  15. 15.Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023.
  16. 16.Du, Y., Yang, S., Dai, B., Dai, H., Nachum, O., Tenenbaum, J., Schuurmans, D., and Abbeel, P. Learning universal policies via text-guided video generation. Advances in Neural Information Processing Systems, 36, 2024.
  17. 17.Esser, P., Rombach, R., and Ommer, B. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 12873–12883, 2021.
  18. 18.Ge, Y., Zhao, S., Zhu, J., Ge, Y., Yi, K., Song, L., Li, C., Ding, X., and Shan, Y. Seed-x: Multimodal models with unified multi-granularity comprehension and generation. arXiv preprint arXiv:2404.14396, 2024.
  19. 19.Ghaffari, S. and Krishnaswamy, N. Exploring failure cases in multimodal reasoning about physical dynamics. In Proceedings of the AAAI Symposium Series, volume 3, pp. 105–114, 2024.
  20. 20.Guo, Y., Hu, Y., Zhang, J., Wang, Y.-J., Chen, X., Lu, C., and Chen, J. Prediction with action: Visual policy learning via joint denoising process. arXiv preprint arXiv:2411.18179, 2024.
  21. 21.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  22. 22.Jang, E., Irpan, A., Khansari, M., Kappler, D., Ebert, F., Lynch, C., Levine, S., and Finn, C. Bc-z: Zero-shot task generalization with robotic imitation learning. In Conference on Robot Learning, pp. 991–1002. PMLR, 2022.
  23. 23.Ke, T.-W., Gkanatsios, N., and Fragkiadaki, K. 3d diffuser actor: Policy diffusion with 3d scene representations. arXiv preprint arXiv:2402.10885, 2024.
  24. 24.Kim, M., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E., Lam, G., Sanketi, P., Vuong, Q., Kollar, T., Burchfiel, B., Tedrake, R., Sadigh, D., Levine, S., Liang, P., and Finn, C. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024.
  25. 25.Li, J., Li, D., Savarese, S., and Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pp. 19730–19742. PMLR, 2023a.
  26. 26.Li, X., Liu, M., Zhang, H., Yu, C., Xu, J., Wu, H., Cheang, C., Jing, Y., Zhang, W., Liu, H., et al. Vision-language foundation models as effective robot imitators. arXiv preprint arXiv:2311.01378, 2023b.
  27. 27.Li, Y., Bubeck, S., Eldan, R., Del Giorno, A., Gunasekar, S., and Lee, Y. T. Textbooks are all you need ii: phi-1.5 technical report. arXiv preprint arXiv:2309.05463, 2023c.
  28. 28.Li, Z., Wang, H., Liu, D., Zhang, C., Ma, A., Long, J., and Cai, W. Multimodal causal reasoning benchmark: Challenging vision large language models to infer causal links between siamese images. arXiv preprint arXiv:2408.08105, 2024.
  29. 29.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. Advances in neural information processing systems, 36, 2024.
  30. 30.Mees, O., Hermann, L., Rosete-Beas, E., and Burgard, W. Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks. IEEE Robotics and Automation Letters (RA-L), 7(3):7327–7334, 2022.
  31. 31.O’Neill, A., Rehman, A., Gupta, A., Maddukuri, A., Gupta, A., Padalkar, A., Lee, A., Pooley, A., Gupta, A., Mandlekar, A., et al. Open x-embodiment: Robotic learning datasets and rt-x models. arXiv preprint arXiv:2310.08864, 2023.
  32. 32.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PMLR, 2021.
  33. 33.Song, W., Chen, J., Ding, P., Zhao, H., Zhao, W., Zhong, Z., Ge, Z., Ma, J., and Li, H. Accelerating vision-language-action model integrated with action chunking via parallel decoding. arXiv preprint arXiv:2503.02310, 2025.
  34. 34.Tan, M. and Le, Q. Efficientnet: Rethinking model scaling for convolutional neural networks. In International conference on machine learning, pp. 6105–6114. PMLR, 2019.
  35. 35.Walke, H., Black, K., Lee, A., Kim, M. J., Du, M., Zheng, C., Zhao, T., Hansen-Estruch, P., Vuong, Q., He, A., Myers, V., Fang, K., Finn, C., and Levine, S. Bridgedata v2: A dataset for robot learning at scale. In Conference on Robot Learning (CoRL), 2023.
  36. 36.Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O. K., Singhal, S., Som, S., et al. Image as a foreign language: Beit pretraining for all vision and vision-language tasks. arXiv preprint arXiv:2208.10442, 2022.
  37. 37.Wang, Y.-J., Zhang, B., Chen, J., and Sreenath, K. Prompt a robot to walk with large language models. arXiv preprint arXiv:2309.09969, 2023.
  38. 38.Wen, C., Jayaraman, D., and Gao, Y. Can transformers capture spatial relations between objects? arXiv preprint arXiv:2403.00729, 2024.
  39. 39.Wu, H., Jing, Y., Cheang, C., Chen, G., Xu, J., Li, X., Liu, M., Li, H., and Kong, T. Unleashing large-scale video generative pre-training for visual robot manipulation. arXiv preprint arXiv:2312.13139, 2023.
  40. 40.Xie, J., Mao, W., Bai, Z., Zhang, D. J., Wang, W., Lin, K. Q., Gu, Y., Chen, Z., Yang, Z., and Shou, M. Z. Show-o: One single transformer to unify multimodal understanding and generation. arXiv preprint arXiv:2408.12528, 2024.
  41. 41.Yu, L., Cheng, Y., Sohn, K., Lezama, J., Zhang, H., Chang, H., Hauptmann, A. G., Yang, M.-H., Hao, Y., Essa, I., et al. Magvit: Masked generative video transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10459–10469, 2023.
  42. 42.Zhang, J., Guo, Y., Chen, X., Wang, Y.-J., Hu, Y., Shi, C., and Chen, J. Hirt: Enhancing robotic control with hierarchical robot transformers. arXiv preprint arXiv:2410.05273, 2024.
  43. 43.Zhao, W., Ding, P., Zhang, M., Gong, Z., Bai, S., Zhao, H., and Wang, D. Vlas: Vision-language-action model with speech instructions for customized robot manipulation. arXiv preprint arXiv:2502.13508, 2025.
  44. 44.Zhen, H., Qiu, X., Chen, P., Yang, J., Yan, X., Du, Y., Hong, Y., and Gan, C. 3d-vla: A 3d vision-language-action generative world model. arXiv preprint arXiv:2403.09631, 2024.
  45. 45.Zheng, B., Gu, J., Li, S., and Dong, C. Lm4lv: A frozen large language model for low-level vision tasks. arXiv preprint arXiv:2405.15734, 2024a.
  46. 46.Zheng, R., Liang, Y., Huang, S., Gao, J., Daumé III, H., Kolobov, A., Huang, F., and Yang, J. Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies. arXiv preprint arXiv:2412.10345, 2024b.

Citation

MLA
Zhang, J., et al. “UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent”. arXiv, 2025, https://doi.org/10.48550/arxiv.2501.18867.
APA
Zhang, J., Guo, Y., Hu, Y., Chen, X., Zhu, X., & Chen, J. (2025). UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent. arXiv. https://doi.org/10.48550/arxiv.2501.18867
Chicago
Zhang, J., Y. Guo, Y. Hu, X. Chen, X. Zhu, and J. Chen. 2025. “UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2501.18867.
Harvard
Zhang, J. et al. (2025) “UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent”. arXiv. Available at: https://doi.org/10.48550/arxiv.2501.18867.
Vancouver
1. Zhang J, Guo Y, Hu Y, Chen X, Zhu X, Chen J (2025) UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent. https://doi.org/10.48550/arxiv.2501.18867

BibTeX

@misc{https://doi.org/10.48550/arxiv.2501.18867,
  doi = {10.48550/ARXIV.2501.18867},
  url = {https://arxiv.org/abs/2501.18867},
  author = {Zhang, Jianke and Guo, Yanjiang and Hu, Yucheng and Chen, Xiaoyu and Zhu, Xiang and Chen, Jianyu},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/