Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Visuomotor policy learning

Visuomotor policy learning is a robot learning paradigm in which an artificial agent learns a behavioral strategy to map visual sensory observations, such as camera images, directly into motor control commands. Rather than relying on separate, hand-crafted stages for object detection, state estimation, and trajectory planning, this framework trains neural network models using methods like imitation learning, reinforcement learning, or generative modeling on demonstration data. By establishing a direct connection between perception and physical movement, visuomotor policies enable robots to execute dexterous manipulation tasks, handle multi-modal action possibilities, and adapt dynamically to complex or changing physical environments.

4 items

DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning

DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning

Zili Lin, Wenyao Zhang, Yuyang Zhang, Zekun Qi, Junyan Lin, Hanxin Zhu, Jiaolong Yang, Zhibo Chen, Yao Mu, Xiaokang Yang, Xin Jin, Wenjun Zeng

OrganizationsEastern Institute of Technology, NingboHong Kong Polytechnic UniversityShanghai Jiao Tong UniversityTsinghua UniversityUniversity of Science and Technology of ChinaZhongguancun Academy

Why you should read this

Presents DeformGen, a dynamics-based augmentation framework that expands demonstration datasets for deformable object manipulation by forward-simulating physically plausible states and warping robot trajectories to match deformed geometries.

Demonstration augmentation is proposed for cost-efficient data acquisition, but existing methods are fundamentally limited in deformable manipulation due to two challenges: (1) the state space is high-dimensional with physics-induced constraints, making valid configurations impossible to reach via low-dimensional pose perturbations; and (2) trajectory transfer is non-equivariant, as material points no longer move rigidly together under deformation. We present DeformGen, a dynamics-based augmentation framework that achieves topological diversity for deformable objects. For the state challenge, DeformGen expands the valid state distribution by applying localized physical disturbances and forward-simulating the dynamics to obtain topology-coherent, physically plausible deformable states. For the trajectory challenge, DeformGen transfers source manipulation trajectories via deformation-field warping, which lifts per-particle displacements into a continuous spatial function to adapt the end-effector trajectory consistently with the deformed geometry. In this way, our method jointly augments the state distribution and its associated manipulation behavior. Experiments on high-fidelity deformable manipulation benchmarks show that DeformGen generally improves policy learning compared with training on the original demonstrations alone and with rigid-style augmentation baselines.

Added

2026-09-30

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Qingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu, Zhuoyang Zhang, Yecheng Wu, Zhaoshuo Li, Qianli Ma, Song Han, Chelsea Finn, Ankur Handa, Tsung-Yi Lin, Gordon Wetzstein, Ming-Yu Liu, Donglai Xiang

OrganizationsMassachusetts Institute of TechnologyNVIDIAStanford University

Why you should read this

Introduces CoT-VLA, a framework that equips vision-language-action models with explicit temporal planning by autoregressively predicting future visual sub-goals prior to action generation, outperforming state-of-the-art robotic manipulation methods in both simulation and the real world.

Vision-language-action models (VLAs) have shown potential in leveraging pretrained vision-language models and diverse robot demonstrations for learning generalizable sensorimotor control. While this paradigm effectively utilizes large-scale data from both robotic and non-robotic sources, current VLAs primarily focus on direct input--output mappings, lacking the intermediate reasoning steps crucial for complex manipulation tasks. As a result, existing VLAs lack temporal planning or reasoning capabilities. In this paper, we introduce a method that incorporates explicit visual chain-of-thought (CoT) reasoning into vision-language-action models (VLAs) by predicting future image frames autoregressively as visual goals before generating a short action sequence to achieve these goals. We introduce CoT-VLA, a state-of-the-art 7B VLA that can understand and generate visual and action tokens. Our experimental results demonstrate that CoT-VLA achieves strong performance, outperforming the state-of-the-art VLA model by 17% in real-world manipulation tasks and 6% in simulation benchmarks. Project website: this https URL

Added

2026-09-26

Diffusion policy: Visuomotor policy learning via action diffusion

Diffusion policy: Visuomotor policy learning via action diffusion

Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel, Shuran Song

Why you should read this

Introduces a diffusion-based visuomotor learning framework that models robot behavior as a conditional denoising process over action sequences, effectively handling multimodal human demonstrations and outperforming existing imitation learning methods across diverse manipulation tasks.

This paper introduces Diffusion Policy, a new way of generating robot behavior by representing a robot's visuomotor policy as a conditional denoising diffusion process. We benchmark Diffusion Policy across 15 different tasks from 4 different robot manipulation benchmarks and find that it consistently outperforms existing state-of-the-art robot learning methods with an average improvement of 46.9%. Diffusion Policy learns the gradient of the action-distribution score function and iteratively optimizes with respect to this gradient field during inference via a series of stochastic Langevin dynamics steps. We find that the diffusion formulation yields powerful advantages when used for robot policies, including gracefully handling multimodal action distributions, being suitable for high-dimensional action spaces, and exhibiting impressive training stability. To fully unlock the potential of diffusion models for visuomotor policy learning on physical robots, this paper presents a set of key technical contributions including the incorporation of receding horizon control, visual conditioning, and the time-series diffusion transformer. We hope this work will help motivate a new generation of policy learning techniques that are able to leverage the powerful generative modeling capabilities of diffusion models. Code, data, and training details is available diffusion-policy.cs.columbia.edu

Added

2026-09-11

RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Songming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan, Huayu Chen, Zhengyi Wang, Ke Xu, Hang Su, Jun Zhu

OrganizationsTsinghua-Bosch Joint ML CenterTsinghua University

Why you should read this

Presents RDT-1B, a 1.2-billion-parameter diffusion foundation model that scales bimanual manipulation through a physically interpretable unified action space and multi-robot pre-training to achieve exceptional zero-shot generalization and dexterous task performance.

Bimanual manipulation is essential in robotics, yet developing foundation models is extremely challenging due to the inherent complexity of coordinating two robot arms (leading to multi-modal action distributions) and the scarcity of training data. In this paper, we present the Robotics Diffusion Transformer (RDT), a pioneering diffusion foundation model for bimanual manipulation. RDT builds on diffusion models to effectively represent multi-modality, with innovative designs of a scalable Transformer to deal with the heterogeneity of multi-modal inputs and to capture the nonlinearity and high frequency of robotic data. To address data scarcity, we further introduce a Physically Interpretable Unified Action Space, which can unify the action representations of various robots while preserving the physical meanings of original actions, facilitating learning transferrable physical knowledge. With these designs, we managed to pre-train RDT on the largest collection of multi-robot datasets to date and scaled it up to 1.2B parameters, which is the largest diffusion-based foundation model for robotic manipulation. We finally fine-tuned RDT on a self-created multi-task bimanual dataset with over 6K+ episodes to refine its manipulation capabilities. Experiments on real robots demonstrate that RDT significantly outperforms existing methods. It exhibits zero-shot generalization to unseen objects and scenes, understands and follows language instructions, learns new skills with just 1~5 demonstrations, and effectively handles complex, dexterous tasks. We refer to https://rdt-robotics.github.io/rdt-robotics/ for the code and videos.

Added

2026-05-17