Behavior Generation with Latent Actions

Seungjae LeeYibin WangHaritheja EtukuruH. Jin KimNur Muhammad (Mahi) ShafiullahLerrel Pinto

article2024ICML223 citations

Proposes Vector-Quantized Behavior Transformer (VQ-BeT), a decision-making model that uses hierarchical vector quantization to capture complex multimodal action distributions while achieving competitive generation quality at five times the inference speed of diffusion policies.

Listen

Generating complex, multi-modal physical behaviors from demonstration data remains a major challenge in robotics and autonomous systems. Traditional behavior cloning models struggle because physical action distributions are continuous and highly varied, where small sequential errors can compound rapidly and lead to catastrophic operational failures. While previous methods like Behavior Transformers (BeT) used simple k-means clustering to tokenize actions and diffusion policies used iterative denoising to handle diversity, both face severe limitations: k-means clustering fails to scale to high-dimensional or long-horizon tasks, and diffusion-based models require significant computational time during live execution.

The article introduces the Vector-Quantized Behavior Transformer (VQ-BeT) to overcome these limitations. The objective of the work is to develop and evaluate a versatile behavior-generation framework that can accurately capture diverse action distributions in continuous domains while maintaining the inference speed required for real-time control.

The authors approach this problem using a two-stage architecture evaluated across simulated benchmarks, an autonomous driving dataset, and physical robotic hardware. The first stage uses a Residual Vector-Quantized Variational Autoencoder (Residual VQ-VAE) to compress continuous action chunks into structured, discrete latent codes using a coarse-to-fine hierarchy. The second stage trains a Transformer network to predict these discrete action tokens from observation sequences—optionally conditioned on specific goals—alongside a continuous offset head that fine-tunes precision. Evaluation spans eight diverse domains: seven robotic manipulation and locomotion benchmarks, trajectory planning on the nuScenes autonomous driving dataset, and twelve real-world manipulation tasks using a mobile robot arm.

The key findings demonstrate clear advantages in performance, behavioral diversity, and computational efficiency. In simulated goal-conditioned tasks, VQ-BeT outperformed leading baselines in six out of seven environments. In unconditional tasks, it achieved state-of-the-art results in five of seven environments and produced higher behavioral entropy, successfully generating multiple valid trajectories rather than collapsing to a single mode. Crucially, because VQ-BeT generates actions in a single forward pass, it delivers a 5-fold speedup in simulation and a 25-fold speedup on real robotic hardware compared to diffusion-based alternatives. In physical robot trials, VQ-BeT matched or exceeded baselines on simple tasks and outperformed diffusion policies by 73% on two-phase sequences, maintaining more than triple the completion rate on extended, long-horizon tasks. On the nuScenes driving benchmark, it achieved the lowest average trajectory error (0.73 meters) among comparable models despite receiving only partial scene information.

These findings suggest that combining learned discrete action spaces with Transformer architectures provides a scalable, computationally efficient foundation for embodied artificial intelligence. For practical operations, the dramatic reduction in inference latency lowers hardware compute requirements and makes closed-loop real-time execution feasible on cost-effective, noisy robotic platforms. Unlike diffusion policies that struggle when open-loop receding-horizon assumptions fail, VQ-BeT operates effectively in direct closed-loop control without sacrificing execution speed.

Based on these results, engineering teams developing real-time autonomous systems or robotic manipulation policies should consider adopting residual vector quantization instead of k-means clustering or computationally intensive diffusion heads. Future efforts should focus on testing the architecture across larger multi-robot datasets to explore cross-embodiment latent action spaces and integrating learned discrete action representations with online reinforcement learning.

Decision-makers should note certain limitations: the autonomous driving evaluation relied on pre-processed object tracks lacking lane boundary data, which moderately elevated collision rates relative to full-information models. Additionally, while the model is robust across most tasks, hyperparameter choices such as codebook sizes and autoregressive decoding require careful tuning depending on whether the domain is simulated or running on physical hardware.

arXiv: 2403.03181
Cover for Behavior Generation with Latent Actions

Abstract

Generative modeling of complex behaviors from labeled datasets has been a longstanding problem in decision-making. Unlike language or image generation, decision-making requires modeling actions — continuous-valued vectors that are multimodal in their distribution, potentially drawn from uncurated sources, where generation errors can compound in sequential prediction. A recent class of models called Behavior Transformers (BeT) addresses this by discretizing actions using k-means clustering to capture different modes. However, k-means struggles to scale for high-dimensional action spaces or long sequences, and lacks gradient information, and thus BeT suffers in modeling long-range actions. In this work, we present Vector-Quantized Behavior Transformer (VQ-BeT), a versatile model for behavior generation that handles multimodal action prediction, conditional generation, and partial observations. VQ-BeT augments BeT by tokenizing continuous actions with a hierarchical vector quantization module. Across seven environments including simulated manipulation, autonomous driving, and robotics, VQ-BeT improves on state-of-the-art models such as BeT and Diffusion Policies. Importantly, we demonstrate VQ-BeT's improved ability to capture behavior modes while accelerating inference speed 5× over Diffusion Policies. Videos can be found https://sjlee.cc/vq-bet/

Table of Contents

  • 1. Introduction
  • 2. Background and Preliminaries
  • 2.1. Behavior cloning
  • 2.2. Behavior Transformers
  • 2.3. Residual Vector Quantization
  • 3. Vector-Quantized Behavior Transformers
  • 3.1. Sequential prediction on behavior data
  • 3.2. Action (chunk) discretization via Residual VQ
  • 3.3. Weighted update for code prediction
  • 3.4. Conditional and non-conditional task formulation
  • 4. Experiments
  • 4.1. Environments, datasets, and baselines
  • 4.2. Performance of behavior generated by VQ-BeT
  • 4.3. How well does VQ-BeT capture multimodality?
  • 4.4. Inference-time efficiency of VQ-BeT
  • 4.5. Adapting VQ-BeT for autonomous driving
  • 4.6. Design decisions that matter for VQ-BeT
  • 4.7. Adapting VQ-BeT to real-world robots
  • 5. Related Works
  • 6. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Experimental and Dataset
  • A.1. Simulated environments
  • A.2. Real-world environments
  • B. Additional Results
  • B.1. VQ-BeT with larger Residual VQ Codebook
  • C. Implementation Details
  • C.1. Model Design Choises
  • C.2. VQ-BeT for Driving Dataset

Knowls

  1. Knowl 1 — Vector-Quantized Behavior Transformer Framework

    model/method

    Vector-Quantized Behavior Transformer (VQ-BeT) is a behavior cloning architecture designed to model continuous, multimodal action distributions over long horizons without suffering from the scaling limitations of kk-means discretization or the inference latency of diffusion models.

    VQ-BeT operates in two sequential training stages:

    1. Action Discretization Stage: A Residual Vector-Quantized Variational Autoencoder (Residual VQ-VAE) is trained offline on action sequences (or action chunks) at:t+na_{t:t+n}. The encoder ϕ\phi maps continuous actions to a continuous latent space, which is quantized hierarchically across NqN_q discrete codebook layers into discrete code indices. The decoder ψ\psi reconstructs continuous actions from the sum of codebook embeddings. This hierarchical structure assigns coarse spatial clustering to the primary code layer and fine-grained action refinement to secondary layers.
    2. Policy Learning Stage: A causal transformer (GPT-style architecture) processes history sequences of observations ot−h:to_{t-h:t} (and optionally goal observation sequences oN−g:No_{N-g:N}). The transformer outputs categorical distribution logits via code prediction heads ζcode\zeta_{\text{code}} to predict the discrete hierarchical action tokens, as well as a continuous offset vector via an offset head ζoffset\zeta_{\text{offset}} that adjusts the decoded action centers to preserve high-fidelity continuous control.

    At test time, the policy performs a single forward pass through the transformer and the frozen Residual VQ decoder to produce action chunk predictions, enabling high-frequency closed-loop execution.

  2. Knowl 2 — Residual Vector Quantization for Continuous Action Tokenization

    model/method

    To discretize continuous actions or action chunks at:t+n∈Rn×daa_{t:t+n} \in \mathbb{R}^{n \times d_a} (where n≥1n \ge 1 is the action horizon and dad_a is the action dimension), VQ-BeT uses a Residual Vector-Quantized Autoencoder parameterized by an encoder ϕ\phi and a decoder ψ\psi.

    The input action chunk is first mapped by the encoder to a latent embedding vector x=ϕ(at:t+n)x = \phi(a_{t:t+n}). The latent vector is quantized across NqN_q cascaded discrete codebook layers, where each layer i∈{1,…,Nq}i \in \{1, \dots, N_q\} contains a codebook {e1i,e2i,…,eki}⊂Rde\{e^i_1, e^i_2, \dots, e^i_k\} \subset \mathbb{R}^{d_e} of kk learned embedding vectors.

    Quantization proceeds recursively via nearest-neighbor lookup on residuals:

    zq1=ec11,where c1=arg⁡min⁡j∥x−ej1∥2z_q^1 = e^1_{c_1}, \quad \text{where } c_1 = \arg\min_j \|x - e^1_j\|_2

    For subsequent layers i=2,…,Nqi = 2, \dots, N_q, the residual vector is quantized:

    zqi=ecii,where ci=arg⁡min⁡j∥x−∑m=1i−1zqm−eji∥2z_q^i = e^i_{c_i}, \quad \text{where } c_i = \arg\min_j \left\| x - \sum_{m=1}^{i-1} z_q^m - e^i_j \right\|_2

    The overall quantized latent representation is the sum of vectors from all codebooks:

    zq(x)=∑i=1Nqzqiz_q(x) = \sum_{i=1}^{N_q} z_q^i

    The continuous action chunk is reconstructed by decoding this sum: a^t:t+n=ψ(zq(x))\hat{a}_{t:t+n} = \psi(z_q(x)).

    The training objective combines an L1L_1 reconstruction loss with the vector quantization commitment loss:

    LRecon=∥at:t+n−ψ(zq(ϕ(at:t+n)))∥1\mathcal{L}_{\text{Recon}} = \|a_{t:t+n} - \psi(z_q(\phi(a_{t:t+n})))\|_1 LRVQ=LRecon+∥SG[ϕ(at:t+n)]−zq(ϕ(at:t+n))∥22+λcommit∥ϕ(at:t+n)−SG[zq(ϕ(at:t+n))]∥22\mathcal{L}_{\text{RVQ}} = \mathcal{L}_{\text{Recon}} + \|\text{SG}[\phi(a_{t:t+n})] - z_q(\phi(a_{t:t+n}))\|_2^2 + \lambda_{\text{commit}}\|\phi(a_{t:t+n}) - \text{SG}[z_q(\phi(a_{t:t+n}))]\|_2^2

    where SG[⋅]\text{SG}[\cdot] denotes the stop-gradient operator, and λcommit=1\lambda_{\text{commit}} = 1 is the commitment weight. Codebook vectors are updated using exponential moving averages rather than direct gradient descent. In practice, Nq=2N_q = 2 layers provide sufficient fidelity, where layer 1 provides coarse clustering (primary code) and layer 2 provides fine-grained corrections (secondary code).

  3. Knowl 3 — VQ-BeT Policy Training Objective and Loss Formulation

    equation

    The policy network in VQ-BeT is trained to predict the hierarchical discrete action codes and a continuous residual offset vector given observation tokens oto_t.

    The policy objective consists of a weighted Focal classification loss over the hierarchical code predictions and an L1L_1 loss over the continuous offset:

    Lcode=Lfocal(ζcodei=1(ot))+βLfocal(ζcodei>1(ot))\mathcal{L}_{\text{code}} = \mathcal{L}_{\text{focal}}(\zeta_{\text{code}}^{i=1}(o_t)) + \beta \mathcal{L}_{\text{focal}}(\zeta_{\text{code}}^{i>1}(o_t)) ⌊at:t+n⌋=ψ(∑j=1k∑i=1Nqeji⋅I[ζcodei(ot)=j])\lfloor a_{t:t+n} \rfloor = \psi\left(\sum_{j=1}^k \sum_{i=1}^{N_q} e_j^i \cdot \mathbb{I}[\zeta_{\text{code}}^i(o_t) = j]\right) Loffset=∥at:t+n−(⌊at:t+n⌋+ζoffset(ot))∥1\mathcal{L}_{\text{offset}} = \|a_{t:t+n} - (\lfloor a_{t:t+n} \rfloor + \zeta_{\text{offset}}(o_t))\|_1 LVQ-BeT=Lcode+Loffset\mathcal{L}_{\text{VQ-BeT}} = \mathcal{L}_{\text{code}} + \mathcal{L}_{\text{offset}}

    where:

    • ζcodei=1(ot)\zeta_{\text{code}}^{i=1}(o_t) denotes the predicted categorical logits for the primary codebook layer (i=1i=1).
    • ζcodei>1(ot)\zeta_{\text{code}}^{i>1}(o_t) denotes the predicted categorical logits for the secondary codebook layers (i>1i > 1).
    • Lfocal\mathcal{L}_{\text{focal}} is the multi-class focal loss measuring cross-entropy against the ground-truth quantized code indices obtained from the trained Residual VQ-VAE.
    • β∈(0,1]\beta \in (0, 1] is a hyperparameter (typically set between 0.10.1 and 0.60.6) that discounts the secondary code loss relative to the primary code loss.
    • ⌊at:t+n⌋\lfloor a_{t:t+n} \rfloor is the decoded action chunk reconstructed by the frozen Residual VQ decoder ψ\psi using the predicted code vectors ejie_j^i.
    • ζoffset(ot)\zeta_{\text{offset}}(o_t) is the output of the continuous offset head, predicting an action-space displacement from the codebook cluster centroid to match the ground-truth continuous action chunk at:t+na_{t:t+n}.
  4. Knowl 4 — Unconditional and Goal-Conditional Policy Formulations in VQ-BeT

    definition

    VQ-BeT formalizes behavior generation under two operational settings without requiring reward annotations:

    1. Unconditional Policy Formulation: Given a dataset of observations and actions D={(ot,at)}\mathcal{D} = \{(o_t, a_t)\}, the unconditional policy models the distribution of future action sequences at:t+n∈Ana_{t:t+n} \in \mathcal{A}^n of length nn conditioned on a history of hh past observations ot−h:t∈Oho_{t-h:t} \in \mathcal{O}^h:
    π:Oh→An,π(at:t+n∣ot−h:t)\pi: \mathcal{O}^h \to \mathcal{A}^n, \quad \pi(a_{t:t+n} \mid o_{t-h:t})
    1. Goal-Conditional Policy Formulation: For goal-directed tasks, the policy is conditioned on both past observations ot−h:t∈Oho_{t-h:t} \in \mathcal{O}^h and a future goal sequence oN−g:N∈Ogo_{N-g:N} \in \mathcal{O}^g representing gg goal observation tokens:
    π:Oh×Og→An,π(at:t+n∣ot−h:t,oN−g:N)\pi: \mathcal{O}^h \times \mathcal{O}^g \to \mathcal{A}^n, \quad \pi(a_{t:t+n} \mid o_{t-h:t}, o_{N-g:N})

    Both formulations tokenize continuous actions via Residual VQ-VAE and use a causal transformer with code prediction heads and an offset head to predict multi-modal trajectories in a single forward pass.

  5. Knowl 5 — Goal-Conditional Behavior Generation Performance Across Simulated Benchmarks

    data/table

    VQ-BeT was evaluated on seven goal-conditional simulation benchmarks against Goal-Conditioned Behavior Cloning (GCBC), Conditional Behavior Transformer (C-BeT), Conditional BESO (C-BESO), and Classifier-Free Guided BESO (CFG-BESO).

    Environment Metric GCBC C-BeT C-BESO CFG-BESO VQ-BeT
    PushT Final IoU (⋅/1\cdot/1) 0.02 0.02 0.30 0.25 0.39
    Image PushT Final IoU (⋅/1\cdot/1) 0.02 0.01 0.02 0.01 0.10
    Kitchen Goals (⋅/4\cdot/4) 0.15 3.09 3.75 3.47 3.78
    Image Kitchen Goals (⋅/4\cdot/4) 0.64 2.41 2.00 1.59 2.60
    Multimodal Ant Goals (⋅/2\cdot/2) 0.00 1.68 1.14 0.92 1.72
    UR3 BlockPush Goals (⋅/2\cdot/2) 0.19 1.67 1.94 0.91 1.94
    BlockPush Success (⋅/1\cdot/1) 0.01 0.87 0.93 0.88 0.87

    VQ-BeT outperforms or matches all baselines in 6 out of the 7 goal-conditional environments. On tasks with visual observations (Image PushT, Image Kitchen) and high-dimensional/multimodal state trajectories (Kitchen, Ant), VQ-BeT achieves higher goal completion rates than both kk-means discretized transformers (C-BeT) and score-based diffusion policies (C-BESO, CFG-BESO).

  6. Knowl 6 — Unconditional Behavior Generation and Behavior Entropy Analysis

    data/table

    VQ-BeT was evaluated on seven unconditional simulation benchmarks against MLP-based Behavior Cloning (BC), Behavior Transformer (BeT), Convolutional Diffusion Policy (DiffPolicy-C), and Transformer Diffusion Policy (DiffPolicy-T).

    Environment Metric BC BeT DiffPolicy-C DiffPolicy-T VQ-BeT
    PushT Final IoU (⋅/1\cdot/1) 0.65 0.39 0.73 0.74 0.78
    Image PushT Final IoU (⋅/1\cdot/1) 0.13 0.01 0.66 0.45 0.68
    Kitchen Goals (⋅/4\cdot/4) 0.18 3.07 2.62 3.44 3.66
    Image Kitchen Goals (⋅/4\cdot/4) 0.75 2.48 3.11 3.01 2.98
    Multimodal Ant Goals (⋅/4\cdot/4) 0.01 2.73 3.12 2.90 3.22
    UR3 BlockPush Goals (⋅/2\cdot/2) 0.11 1.59 1.83 1.82 1.84
    BlockPush Goals (⋅/2\cdot/2) 0.01 1.67 0.47 1.93 1.79

    To quantify the diversity of multimodal rollouts rather than mode collapse to a single subtask sequence, the behavior entropy of subtask completion order was evaluated:

    • In Franka Kitchen, VQ-BeT achieves a 4-subtask entropy (p4p4-Entropy) of 4.074.07, outperforming BeT (4.014.01), DiffPolicy-C (3.623.62), and DiffPolicy-T (3.893.89).
    • In Multimodal Ant, VQ-BeT achieves a p4p4-Entropy of 4.204.20, surpassing BeT (3.553.55), DiffPolicy-C (4.184.18), and DiffPolicy-T (4.114.11).
    • In BlockPush and UR3 BlockPush, VQ-BeT achieves p2p2-Entropy values of 1.991.99 and 0.990.99, matching or exceeding all diffusion and transformer baselines.
  7. Knowl 7 — Single-Pass Inference Efficiency of VQ-BeT Versus Iterative Diffusion Policies

    empirical result

    Because VQ-BeT uses a single-pass causal transformer followed by a lightweight feedforward Residual VQ decoder, it generates action predictions substantially faster than iterative denoising diffusion models (Diffusion Policy, BESO), which require 10 to 100 sequential denoising steps per action.

    In benchmarked environments:

    • Simulation (Franka Kitchen): In single-step action prediction, VQ-BeT achieves an inference latency of 15.1 ms15.1\text{ ms} (conditional) and 22.8 ms22.8\text{ ms} (unconditional), compared to 98.6–100.5 ms98.6\text{--}100.5\text{ ms} for DiffusionPolicy-C/T (evaluated at 10 diffusion iterations) and 41.7 ms41.7\text{ ms} for CFG-BESO, providing a ≈5×\approx 5\times speedup over diffusion.
    • Real Robot (Workstation GPU - NVIDIA RTX A4000): VQ-BeT executes in 18.06 ms18.06\text{ ms} per step compared to 573.49 ms573.49\text{ ms} for DiffusionPolicy-T, an acceleration factor of 31.7×31.7\times.
    • Real Robot (Onboard CPU - 4-Core Intel): VQ-BeT executes in 207.25 ms207.25\text{ ms} per step compared to 5243.82 ms5243.82\text{ ms} for DiffusionPolicy-T, an acceleration factor of 25.3×25.3\times.

    In physical robot systems with compliant hardware or control noise (such as the Hello Robot Stretch), open-loop receding-horizon control with action execution windows (n=3n=3) leads to immediate out-of-distribution drift and 0%0\% task success. Full closed-loop control (n=1n=1) is required, making VQ-BeT's low single-step latency critical for real-time reactive execution.

  8. Knowl 8 — Autonomous Driving Trajectory Forecasting on nuScenes Dataset

    data/table

    VQ-BeT was evaluated on 6-frame trajectory planning on the nuScenes autonomous driving benchmark against end-to-end full-information models and partial-information models. Observation tokens provided to VQ-BeT include a Mission Token (go forward / turn left / turn right), an Ego-state Token, a Trajectory History Token (last 2 seconds), and up to 51 Object Tokens (position, velocity, and 15-class one-hot encoding). Lane and shoulder boundaries are excluded in the partial-information setting.

    L2L_2 Error (m) (↓\downarrow) Collision Rate (%) (↓\downarrow)
    Protocol Method 1s 2s 3s Avg 1s 2s 3s Avg
    ST-P3 ST-P3 1.33 2.11 2.90 2.11 0.23 0.62 1.27 0.71
    VAD 0.17 0.34 0.60 0.37 0.07 0.10 0.24 0.14
    Agent-Driver (Full) 0.16 0.34 0.61 0.37 0.02 0.07 0.18 0.09
    GPT-Driver (Partial) 0.20 0.40 0.70 0.44 0.04 0.12 0.36 0.17
    Diff. Traj. Pred. (Partial) 0.21 0.43 0.80 0.48 0.01 0.07 0.35 0.14
    VQ-BeT (Partial) 0.17 0.33 0.60 0.37 0.02 0.11 0.34 0.16
    UniAD FF 0.55 1.20 2.54 1.43 0.06 0.17 1.07 0.43
    EO 0.67 1.36 2.78 1.60 0.04 0.09 0.88 0.33
    UniAD (Full) 0.48 0.96 1.65 1.03 0.05 0.17 0.71 0.31
    Agent-Driver (Full) 0.22 0.65 1.34 0.74 0.02 0.13 0.48 0.21
    GPT-Driver (Partial) 0.27 0.74 1.52 0.84 0.07 0.15 1.10 0.44
    Diff. Traj. Pred. (Partial) 0.27 0.78 1.83 0.96 0.00 0.27 1.21 0.49
    VQ-BeT (Partial) 0.22 0.62 1.34 0.73 0.02 0.16 0.70 0.29

    With partial information, VQ-BeT outperforms all partial-information baselines on L2L_2 error and collision rate, and matches or outperforms full-information specialized driving systems (UniAD, VAD, Agent-Driver) on trajectory L2L_2 accuracy (0.73 m0.73\text{ m} average on UniAD protocol, 0.37 m0.37\text{ m} on ST-P3 protocol).

  9. Knowl 9 — Real-World Robotic Manipulation Performance on Single and Multi-Phase Tasks

    empirical result

    VQ-BeT was evaluated on a Hello Robot Stretch mobile manipulator equipped with an HPR visual encoder across 12 physical manipulation tasks using 45 human demonstrations per task (80% training, 20% validation):

    1. Single-Phase Tasks: Across five standalone tasks (Open Toaster, Close Toaster, Close Fridge, Can to Toaster, Can to Fridge), VQ-BeT achieved 47/5047/50 (94.0%94.0\%) cumulative success, comparable to DiffusionPolicy-T (45/5045/50, 90.0%90.0\%) and outperforming MLP-BC (29/5029/50, 58.0%58.0\%) and Depth-augmented BC (27/5027/50, 54.0%54.0\%).
    2. Two-Phase Tasks: Across three sequenced tasks (Can to Fridge →\to Close Fridge, Can to Toaster →\to Close Toaster, Close Fridge and Toaster), VQ-BeT achieved 19/3019/30 (63.3%63.3\%) success versus 11/3011/30 (36.7%36.7\%) for DiffusionPolicy-T, representing a 73%73\% relative performance gain. On "Close Fridge and Toaster", VQ-BeT successfully exhibited multimodal execution by alternating the door closure sequence across rollouts without collapsing to a single order.
    3. Long-Horizon Multi-Phase Tasks (3–4 Subtasks):
      • Task 1 (Open Drawer →\to Pick and Place Box →\to Close Drawer): VQ-BeT achieved 6/106/10 end-to-end success vs 2/102/10 for DiffusionPolicy-T.
      • Task 2 (Pick up Bread →\to Place in Bag →\to Pick up Bag →\to Place on Table): VQ-BeT achieved 3/103/10 end-to-end success vs 1/101/10 for DiffusionPolicy-T.
      • Task 3 (Can to Fridge →\to Close Fridge →\to Open Toaster): VQ-BeT achieved 7/107/10 end-to-end success vs 2/102/10 for DiffusionPolicy-T.
      • Task 4 (Can to Toaster →\to Close Toaster →\to Close Fridge): VQ-BeT achieved 6/106/10 end-to-end success vs 1/101/10 for DiffusionPolicy-T. Across all four long-horizon real-world tasks, VQ-BeT achieved at least 3×3\times the full-task completion rate of DiffusionPolicy-T.
  10. Knowl 10 — Ablation Analysis of VQ-BeT Architectural Components and Codebook Scaling

    empirical result

    Component ablations conducted across Franka Kitchen (conditional), Multimodal Ant (unconditional), and nuScenes driving establish the necessity of each design choice:

    1. Residual VQ vs. Vanilla Single-Layer VQ: Replacing Residual VQ with a single VQ layer degrades performance from 3.783.78 to 3.673.67 completed goals in Ant and from 3.223.22 to 2.932.93 completed goals in Kitchen, demonstrating that cascaded quantization is essential for representing complex continuous action spaces.
    2. Secondary Code Loss Weight (β\beta): Setting β=1.0\beta = 1.0 (equal loss weight between primary and secondary codes) causes performance drops (3.653.65 in Ant, 2.992.99 in Kitchen) compared to setting β∈[0.1,0.6]\beta \in [0.1, 0.6], showing the benefit of prioritizing coarse primary mode classification over fine-grained residual fitting during transformer training.
    3. Continuous Offset Head: Removing the continuous offset head ζoffset\zeta_{\text{offset}} causes catastrophic degradation (0.520.52 goals completed in Ant vs 3.783.78 with offset; 2.822.82 goals in Kitchen vs 3.223.22 with offset; nuScenes L2L_2 error increases from 0.73 m0.73\text{ m} to 1.41 m1.41\text{ m}), demonstrating that discrete code centroids alone lack the precision needed for closed-loop physics control.
    4. Codebook Scaling and Deadcode Masking: Increasing the number of Residual VQ code combinations by 10×10\times to 250×250\times (e.g., from 256256 to 65,53665{,}536 combinations in Kitchen Conditional) results in minimal performance decrease (from 3.783.78 to 3.613.61 goals, a 4.5%4.5\% change). The primary code classification accuracy remains at 80%80\% of original accuracy despite codebook expansion, allowing VQ-BeT to maintain task performance by relying on coarse primary codes. Masking deadcodes (combinations not present in the dataset) at sampling time does not provide consistent benefits across environments.

Coverage note — All substantial contributions—including the two-stage VQ-BeT model, the Residual VQ action discretization objective, the policy loss formulation, the conditional and unconditional task formulations, the simulation benchmarks, multimodality entropy analyses, inference latency benchmarks, nuScenes autonomous driving evaluation, real-world robotic manipulation results, and ablation studies—have been captured as self-contained knowls.

References

  1. 1.Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Bar-Tal, O., Chefer, H., Tov, O., Herrmann, C., Paiss, R., Zada, S., Ephrat, A., Hur, J., Li, Y., Michaeli, T., et al. Lumiere: A space-time diffusion model for video generation. arXiv preprint arXiv:2401.12945, 2024.
  3. 3.Bellman, R., Glicksberg, I., and Gross, O. On the “bangbang” control problem. Quarterly of Applied Mathematics, 14(1):11–18, 1956.
  4. 4.Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. Openai gym. arXiv preprint arXiv:1606.01540, 2016.
  5. 5.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020.
  6. 6.Bushaw, D. W. Differential equations with a discontinuous forcing term. PhD thesis, Princeton University, 1952.
  7. 7.Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O. nuscenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11621–11631, 2020.
  8. 8.Chen, L., Bahl, S., and Pathak, D. Playfusion: Skill acquisition via diffusion from language-annotated play. In Conference on Robot Learning, pp. 2012–2029. PMLR, 2023.
  9. 9.Chi, C., Feng, S., Du, Y., Xu, Z., Cousineau, E., Burchfiel, B., and Song, S. Diffusion policy: Visuomotor policy learning via action diffusion. arXiv preprint arXiv:2303.04137, 2023.
  10. 10.Cui, Z. J., Wang, Y., Shafiullah, N. M. M., and Pinto, L. From play to policy: Conditional behavior generation from uncurated robot data. arXiv preprint arXiv:2210.10047, 2022.
  11. 11.Dadashi, R., Hussenot, L., Vincent, D., Girgin, S., Raichuk, A., Geist, M., and Pietquin, O. Continuous control with action quantization from demonstrations. arXiv preprint arXiv:2110.10149, 2021.
  12. 12.Dhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., and Sutskever, I. Jukebox: A generative model for music. arXiv preprint arXiv:2005.00341, 2020.
  13. 13.Finn, C., Levine, S., and Abbeel, P. Guided cost learning: Deep inverse optimal control via policy optimization. In International conference on machine learning, pp. 49–58. PMLR, 2016.
  14. 14.Florence, P., Lynch, C., Zeng, A., Ramirez, O. A., Wahid, A., Downs, L., Wong, A., Lee, J., Mordatch, I., and Tompson, J. Implicit behavioral cloning. In Conference on Robot Learning, pp. 158–168. PMLR, 2022.
  15. 15.Gupta, A., Kumar, V., Lynch, C., Levine, S., and Hausman, K. Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning. arXiv preprint arXiv:1910.11956, 2019.
  16. 16.Hausknecht, M. and Stone, P. Deep reinforcement learning in parameterized action space. arXiv preprint arXiv:1511.04143, 2015.
  17. 17.Ho, J. and Ermon, S. Generative adversarial imitation learning. Advances in neural information processing systems, 29, 2016.
  18. 18.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  19. 19.Hu, P., Huang, A., Dolan, J., Held, D., and Ramanan, D. Safe local motion planning with self-supervised freespace forecasting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12732– 12741, 2021.
  20. 20.Hu, S., Chen, L., Wu, P., Li, H., Yan, J., and Tao, D. St-p3: End-to-end vision-based autonomous driving via spatialtemporal feature learning. In European Conference on Computer Vision, pp. 533–549. Springer, 2022.
  21. 21.Hu, Y., Yang, J., Chen, L., Li, K., Sima, C., Zhu, X., Chai, S., Du, S., Lin, T., Wang, W., et al. Planning-oriented autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 17853–17862, 2023.
  22. 22.Islam, R., Zang, H., Goyal, A., Lamb, A., Kawaguchi, K., Li, X., Laroche, R., Bengio, Y., and Combes, R. T. D. Discrete factorial representations as an abstraction for goal conditioned reinforcement learning. arXiv preprint arXiv:2211.00247, 2022.
  23. 23.Jiang, B., Chen, S., Xu, Q., Liao, B., Chen, J., Zhou, H., Zhang, Q., Liu, W., Huang, C., and Wang, X. Vad: Vectorized scene representation for efficient autonomous driving. arXiv preprint arXiv:2303.12077, 2023.
  24. 24.Kalakrishnan, M., Pastor, P., Righetti, L., and Schaal, S. Learning objective functions for manipulation. In 2013 IEEE International Conference on Robotics and Automation, pp. 1331–1336. IEEE, 2013.
  25. 25.Kemp, C. C., Edsinger, A., Clever, H. M., and Matulevich, B. The design of stretch: A compact, lightweight mobile manipulator for indoor human environments. In 2022 International Conference on Robotics and Automation (ICRA), pp. 3150–3157. IEEE, 2022.
  26. 26.Khurana, T., Hu, P., Dave, A., Ziglar, J., Held, D., and Ramanan, D. Differentiable raycasting for self-supervised occupancy forecasting. In European Conference on Computer Vision, pp. 353–369. Springer, 2022.
  27. 27.Kim, J., hyeon Park, J., Cho, D., and Kim, H. J. Automating reinforcement learning with example-based resets. IEEE Robotics and Automation Letters, 7(3):6606–6613, 2022.
  28. 28.Lin, T.-Y., Goyal, P., Girshick, R., He, K., and Dollár, P. Focal loss for dense object detection. In Proceedings of the IEEE international conference on computer vision, pp. 2980–2988, 2017.
  29. 29.Luo, J., Dong, P., Wu, J., Kumar, A., Geng, X., and Levine, S. Action-quantized offline reinforcement learning for robotic skill learning. In Conference on Robot Learning, pp. 1348–1361. PMLR, 2023.
  30. 30.Lynch, C., Khansari, M., Xiao, T., Kumar, V., Tompson, J., Levine, S., and Sermanet, P. Learning latent plans from play. In Conference on robot learning, pp. 1113–1132. PMLR, 2020.
  31. 31.Mandlekar, A., Zhu, Y., Garg, A., Booher, J., Spero, M., Tung, A., Gao, J., Emmons, J., Gupta, A., Orbay, E., et al. Roboturk: A crowdsourcing platform for robotic skill learning through imitation. In Conference on Robot Learning, pp. 879–893. PMLR, 2018.
  32. 32.Mao, J., Qian, Y., Zhao, H., and Wang, Y. Gpt-driver: Learning to drive with gpt. arXiv preprint arXiv:2310.01415, 2023a.
  33. 33.Mao, J., Ye, J., Qian, Y., Pavone, M., and Wang, Y. A language agent for autonomous driving. arXiv preprint arXiv:2311.10813, 2023b.
  34. 34.Mazzaglia, P., Verbelen, T., Dhoedt, B., Lacoste, A., and Rajeswar, S. Choreographer: Learning and adapting skills in imagination. arXiv preprint arXiv:2211.13350, 2022.
  35. 35.Metz, L., Ibarz, J., Jaitly, N., and Davidson, J. Discrete sequential prediction of continuous actions for deep rl. arXiv preprint arXiv:1705.05035, 2017.
  36. 36.Pearce, T., Rashid, T., Kanervisto, A., Bignell, D., Sun, M., Georgescu, R., Macua, S. V., Tan, S. Z., Momennejad, I., Hofmann, K., et al. Imitating human behaviour with diffusion models. arXiv preprint arXiv:2301.10677, 2023.
  37. 37.Pertsch, K., Lee, Y., and Lim, J. Accelerating reinforcement learning with learned skill priors. In Conference on robot learning, pp. 188–204. PMLR, 2021.
  38. 38.Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023.
  39. 39.Rajaraman, N., Yang, L., Jiao, J., and Ramchandran, K. Toward the fundamental limits of imitation learning. Advances in Neural Information Processing Systems, 33: 2914–2924, 2020.
  40. 40.Reuss, M., Li, M., Jia, X., and Lioutikov, R. Goalconditioned imitation learning using score-based diffusion policies. arXiv preprint arXiv:2304.02532, 2023.
  41. 41.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695, 2022.
  42. 42.Ross, S., Gordon, G., and Bagnell, D. A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, pp. 627–635. JMLR Workshop and Conference Proceedings, 2011.
  43. 43.Shafiullah, N. M., Cui, Z., Altanzaya, A. A., and Pinto, L. Behavior transformers: Cloning k modes with one stone. Advances in neural information processing systems, 35: 22955–22968, 2022.
  44. 44.Shafiullah, N. M. M., Rai, A., Etukuru, H., Liu, Y., Misra, I., Chintala, S., and Pinto, L. On bringing robots home. arXiv preprint arXiv:2311.16098, 2023.
  45. 45.Singh, A., Liu, H., Zhou, G., Yu, A., Rhinehart, N., and Levine, S. Parrot: Data-driven behavioral priors for reinforcement learning. arXiv preprint arXiv:2011.10024, 2020.
  46. 46.Stolle, M. and Precup, D. Learning options in reinforcement learning. In Abstraction, Reformulation, and Approximation: 5th International Symposium, SARA 2002 Kananaskis, Alberta, Canada August 2–4, 2002 Proceedings 5, pp. 212–223. Springer, 2002.
  47. 47.Sutton, R. S., Precup, D., and Singh, S. Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning. Artificial intelligence, 112(1-2): 181–211, 1999.
  48. 48.Tavakoli, A., Pardo, F., and Kormushev, P. Action branching architectures for deep reinforcement learning. In Proceedings of the aaai conference on artificial intelligence, volume 32, 2018.
  49. 49.Van Den Oord, A., Vinyals, O., et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017.
  50. 50.Vasuki, A. and Vanathi, P. A review of vector quantization techniques. IEEE Potentials, 25(4):39–47, 2006.
  51. 51.Wei, B., Ren, M., Zeng, W., Liang, M., Yang, B., and Urtasun, R. Perceive, attend, and drive: Learning spatial attention for safe self-driving. In 2021 IEEE International Conference on Robotics and Automation (ICRA), pp. 4875–4881. IEEE, 2021.
  52. 52.Wu, C., Huang, L., Zhang, Q., Li, B., Ji, L., Yang, F., Sapiro, G., and Duan, N. Godiva: Generating opendomain videos from natural descriptions. arXiv preprint arXiv:2104.14806, 2021.
  53. 53.Wulfmeier, M., Ondruska, P., and Posner, I. Maximum entropy deep inverse reinforcement learning. arXiv preprint arXiv:1507.04888, 2015.
  54. 54.Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M. Soundstream: An end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2021.
  55. 55.Zeng, W., Luo, W., Suo, S., Sadat, A., Yang, B., Casas, S., and Urtasun, R. End-to-end interpretable neural motion planner. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8660– 8669, 2019.
  56. 56.Zhao, T. Z., Kumar, V., Levine, S., and Finn, C. Learning fine-grained bimanual manipulation with low-cost hardware. arXiv preprint arXiv:2304.13705, 2023.
  57. 57.Ziv, A., Gat, I., Lan, G. L., Remez, T., Kreuk, F., Défossez, A., Copet, J., Synnaeve, G., and Adi, Y. Masked audio generation using a single non-autoregressive transformer. arXiv preprint arXiv:2401.04577, 2024.

Citation

MLA
Lee, S., et al. “Behavior Generation with Latent Actions”. PMLR 235:26991-27008, 2024, 2024, http://arxiv.org/abs/2403.03181v2.
APA
Lee, S., Wang, Y., Etukuru, H., Kim, H. J., Shafiullah, N. M. M., & Pinto, L. (2024). Behavior Generation with Latent Actions. PMLR 235:26991-27008, 2024. http://arxiv.org/abs/2403.03181v2
Chicago
Lee, S., Y. Wang, H. Etukuru, H. J. Kim, N. M. M. Shafiullah, and L. Pinto. 2024. “Behavior Generation with Latent Actions”. PMLR 235:26991-27008, 2024. http://arxiv.org/abs/2403.03181v2.
Harvard
Lee, S. et al. (2024) “Behavior Generation with Latent Actions”, PMLR 235:26991-27008, 2024 [Preprint]. Available at: http://arxiv.org/abs/2403.03181v2.
Vancouver
1. Lee S, Wang Y, Etukuru H, Kim HJ, Shafiullah NMM, Pinto L (2024) Behavior Generation with Latent Actions. PMLR 235:26991-27008, 2024

BibTeX

@article{lee2024behavior,
  title = {Behavior Generation with Latent Actions},
  author = {Lee, Seungjae and Wang, Yibin and Etukuru, Haritheja and Kim, H. Jin and Shafiullah, Nur Muhammad Mahi and Pinto, Lerrel},
  year = {2024},
  journal = {PMLR 235:26991-27008, 2024},
  url = {http://arxiv.org/abs/2403.03181v2},
  eprint = {2403.03181}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/