Behavior Generation with Latent Actions
Seungjae LeeYibin WangHaritheja EtukuruH. Jin KimNur Muhammad (Mahi) ShafiullahLerrel Pinto
Proposes Vector-Quantized Behavior Transformer (VQ-BeT), a decision-making model that uses hierarchical vector quantization to capture complex multimodal action distributions while achieving competitive generation quality at five times the inference speed of diffusion policies.
Generating complex, multi-modal physical behaviors from demonstration data remains a major challenge in robotics and autonomous systems. Traditional behavior cloning models struggle because physical action distributions are continuous and highly varied, where small sequential errors can compound rapidly and lead to catastrophic operational failures. While previous methods like Behavior Transformers (BeT) used simple k-means clustering to tokenize actions and diffusion policies used iterative denoising to handle diversity, both face severe limitations: k-means clustering fails to scale to high-dimensional or long-horizon tasks, and diffusion-based models require significant computational time during live execution.
The article introduces the Vector-Quantized Behavior Transformer (VQ-BeT) to overcome these limitations. The objective of the work is to develop and evaluate a versatile behavior-generation framework that can accurately capture diverse action distributions in continuous domains while maintaining the inference speed required for real-time control.
The authors approach this problem using a two-stage architecture evaluated across simulated benchmarks, an autonomous driving dataset, and physical robotic hardware. The first stage uses a Residual Vector-Quantized Variational Autoencoder (Residual VQ-VAE) to compress continuous action chunks into structured, discrete latent codes using a coarse-to-fine hierarchy. The second stage trains a Transformer network to predict these discrete action tokens from observation sequences—optionally conditioned on specific goals—alongside a continuous offset head that fine-tunes precision. Evaluation spans eight diverse domains: seven robotic manipulation and locomotion benchmarks, trajectory planning on the nuScenes autonomous driving dataset, and twelve real-world manipulation tasks using a mobile robot arm.
The key findings demonstrate clear advantages in performance, behavioral diversity, and computational efficiency. In simulated goal-conditioned tasks, VQ-BeT outperformed leading baselines in six out of seven environments. In unconditional tasks, it achieved state-of-the-art results in five of seven environments and produced higher behavioral entropy, successfully generating multiple valid trajectories rather than collapsing to a single mode. Crucially, because VQ-BeT generates actions in a single forward pass, it delivers a 5-fold speedup in simulation and a 25-fold speedup on real robotic hardware compared to diffusion-based alternatives. In physical robot trials, VQ-BeT matched or exceeded baselines on simple tasks and outperformed diffusion policies by 73% on two-phase sequences, maintaining more than triple the completion rate on extended, long-horizon tasks. On the nuScenes driving benchmark, it achieved the lowest average trajectory error (0.73 meters) among comparable models despite receiving only partial scene information.
These findings suggest that combining learned discrete action spaces with Transformer architectures provides a scalable, computationally efficient foundation for embodied artificial intelligence. For practical operations, the dramatic reduction in inference latency lowers hardware compute requirements and makes closed-loop real-time execution feasible on cost-effective, noisy robotic platforms. Unlike diffusion policies that struggle when open-loop receding-horizon assumptions fail, VQ-BeT operates effectively in direct closed-loop control without sacrificing execution speed.
Based on these results, engineering teams developing real-time autonomous systems or robotic manipulation policies should consider adopting residual vector quantization instead of k-means clustering or computationally intensive diffusion heads. Future efforts should focus on testing the architecture across larger multi-robot datasets to explore cross-embodiment latent action spaces and integrating learned discrete action representations with online reinforcement learning.
Decision-makers should note certain limitations: the autonomous driving evaluation relied on pre-processed object tracks lacking lane boundary data, which moderately elevated collision rates relative to full-information models. Additionally, while the model is robust across most tasks, hyperparameter choices such as codebook sizes and autoregressive decoding require careful tuning depending on whether the domain is simulated or running on physical hardware.
- Paper: Neural Discrete Representation Learning, Aäron van den Oord et al. (2017). Introduces Vector-Quantized Variational Autoencoders (VQ-VAE) and discrete codebook tokenization, which provides the foundational vector quantization framework adapted by VQ-BeT for tokenizing continuous action spaces.
- Paper: Generating Diverse High-Fidelity Images with VQ-VAE-2, Ali Razavi et al. (2019). Presents hierarchical vector quantization to capture multiscale structure in discrete representation learning, directly inspiring the hierarchical VQ architecture used in VQ-BeT.
- Paper: Diffusion policy: Visuomotor policy learning via action diffusion, Cheng Chi et al. (2023). Establishes generative action modeling for continuous multimodal robot policies, serving as a primary benchmark and comparative baseline for VQ-BeT.
- Paper: Decision Transformer: Reinforcement Learning via Sequence Modeling, Lili Chen et al. (2021). Demonstrates the efficacy of autoregressive transformer architectures for modeling trajectory sequences in sequential decision-making and control.
- Paper: Planning with Diffusion for Flexible Behavior Synthesis, Michael Janner et al. (2022). Pioneers generative modeling for continuous behavioral trajectories in decision-making, providing context for generative policy generation.
- Paper: Regularized Vector Quantization for Tokenized Image Synthesis, Jiahui Zhang et al. (2023). Examines codebook collapse and regularization strategies in vector quantization, offering key insights into optimizing discrete latent token representations.
- Paper: DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation, Haozhe Xie et al. (2026). Extends latent action representations to real-time, dynamic object manipulation through streaming and latency-aware vision-language-action policies.
- Paper: OpenVLA: An Open-Source Vision-Language-Action Model, Moo Jin Kim et al. (2024). Applies discrete action tokenization and transformer-based policy architectures at scale within an open-source vision-language-action model across diverse robot embodiments.
- Paper: Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations, Yucheng Hu et al. (2025). Combines predictive visual dynamic representations with robot policy learning to guide continuous action generation in embodied tasks.
- Paper: History-Guided Video Diffusion, Kiwhan Song et al. (2025). Develops flexible history-conditioned generative diffusion modeling for extended horizons, advancing sequential modeling for long-range physical dynamics.
