DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning

Zili LinWenyao ZhangYuyang ZhangZekun QiJunyan LinHanxin ZhuJiaolong YangZhibo ChenYao MuXiaokang Yang

article2026arXiv0 citations

Presents DeformGen, a dynamics-based augmentation framework that expands demonstration datasets for deformable object manipulation by forward-simulating physically plausible states and warping robot trajectories to match deformed geometries.

Listen

Training robotic systems to manipulate deformable items—such as garments, ropes, and flexible packages—is essential for advancing industrial automation, household robotics, and healthcare assistance. However, training these robots requires vast volumes of demonstration data, which are prohibitively costly, time-consuming, and difficult to collect manually. While traditional data augmentation techniques expand rigid-object demonstrations using simple geometric shifts, these methods fail on soft objects because material deformations alter local contact points and physical dynamics in complex ways.

To address this data bottleneck, the article introduces and evaluates DeformGen, a framework designed to automate the generation of diverse, physically realistic demonstrations for deformable object manipulation starting from just a single human example. The framework aims to expand initial object shapes and adapt corresponding robotic motion trajectories without violating physical laws.

DeformGen operates in two high-level stages. First, it generates diverse, physically plausible starting configurations by applying randomized physical disturbances via simulated robotic contact and letting physics-engine simulations naturally settle the object into stable, non-rigid states. Second, it adapts the original robot manipulation path to the newly deformed shape through spatial deformation-field warping, which adjusts gripper position and orientation to match local material shifts while preserving the broader task path. The authors evaluated this approach in high-fidelity simulation environments across three representative deformable tasks (rope routing, toy packing, and cloth folding) using four widely used robotic learning architectures.

Across the evaluated setups, policies trained on DeformGen-augmented data dramatically outperformed models trained on a single demonstration, where success rates were nearly zero (1.3% to 2.5%). DeformGen achieved average task success rates between 37.3% and 59.0% across the architectures, generally outperforming rigid-style augmentation baselines. Topological state expansion proved crucial, as policies trained on diverse physical shapes generalized far better than those trained only on rigid shifts. Furthermore, scaling the amount of synthetically generated data steadily enhanced performance, lifting average success rates from approximately 20–37% with 100 trajectories up to 61–63% with 750 trajectories.

These findings demonstrate that scalable, automated demonstration generation for soft materials is technically viable when physical dynamics and local geometry are jointly accounted for. This reduces the dependency on labor-intensive physical data collection, lowering deployment costs and development timelines for complex robotics applications without sacrificing performance on simpler, rigid object orientations.

Organizations developing robotic manipulation systems should consider integrating dynamics-based synthetic augmentation pipelines into their simulation workflows. However, before deploying these systems in production, teams should conduct real-world pilot studies to assess physical deployment risks. Future developmental efforts should focus on multi-robot or dual-arm coordination and expand testing to complex non-rigid tasks like dough manipulation or surgical tissue handling.

The findings are supported by consistent trends across multiple model architectures and thousands of simulation runs. Readers should exercise caution regarding the sim-to-real transfer gap, as all evaluations occurred within simulated environments, and trajectory synthesis success rates can drop on severe object deformations.

No sufficiently relevant recommendations were found.

Cover for DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning

Abstract

Demonstration augmentation is proposed for cost-efficient data acquisition, but existing methods are fundamentally limited in deformable manipulation due to two challenges: (1) the state space is high-dimensional with physics-induced constraints, making valid configurations impossible to reach via low-dimensional pose perturbations; and (2) trajectory transfer is non-equivariant, as material points no longer move rigidly together under deformation. We present DeformGen, a dynamics-based augmentation framework that achieves topological diversity for deformable objects. For the state challenge, DeformGen expands the valid state distribution by applying localized physical disturbances and forward-simulating the dynamics to obtain topology-coherent, physically plausible deformable states. For the trajectory challenge, DeformGen transfers source manipulation trajectories via deformation-field warping, which lifts per-particle displacements into a continuous spatial function to adapt the end-effector trajectory consistently with the deformed geometry. In this way, our method jointly augments the state distribution and its associated manipulation behavior. Experiments on high-fidelity deformable manipulation benchmarks show that DeformGen generally improves policy learning compared with training on the original demonstrations alone and with rigid-style augmentation baselines.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 2.1 Data augmentation for robot manipulation
  • 2.2 Deformable object manipulation
  • 3 Method
  • 3.1 State Augmentation
  • 3.2 Trajectory Augmentation
  • 3.3 Policy Training
  • 4 Experiments
  • 4.1 Implementation Details
  • 4.2 Experiment Results
  • 4.3 State Coverage Analysis
  • 4.4 Ablation Studies
  • 4.5 Failure Analysis
  • 5 Conclusion
  • References
  • A State Augmentation Details
  • A.1 Formal Assumption
  • A.2 Detailed Analysis of Existing Strategies
  • A.3 Advantages of Dynamics-Based Augmentation over (ii) and (iii)
  • A.4 Reachable Set Discussion
  • B Trajectory Augmentation Details
  • B.1 Decay Function
  • B.2 KNN Scope for Grasp vs. Manipulation Phases
  • B.3 Orientation Constraints
  • C Implementation Details
  • C.1 Simulation and Robot Setup
  • C.2 Task Descriptions and Success Criteria
  • C.3 State Augmentation Parameters
  • C.4 Data Splits
  • C.5 Trajectory Augmentation Hyperparameters
  • C.6 Policy Training Hyperparameters
  • D Limitations
  • E Broader Impacts
  • F Licenses

Knowls

  1. Knowl 1 — DeformGen jointly augments deformable states and their manipulation trajectories

    model/method

    DeformGen generates varied demonstrations for a deformable-manipulation task from a sparse source dataset, using one teleoperation demonstration per task in its experiments. It first creates new object configurations by applying localized physical disturbances and simulating the object dynamics from a valid source state. It then transfers the source manipulation behavior to each generated configuration with a deformation field derived from the source-to-target particle displacements. The resulting state-and-trajectory pairs are used to train visuomotor policies. The pipeline targets two shortcomings of rigid-style augmentation: rigid pose changes do not create meaningful shape or topology variation, and transferring a trajectory by one rigid transformation cannot adapt the gripper to local deformation.

  2. Knowl 2 — Deformable states occupy a physically constrained subset of particle configuration space

    assumption

    For an object represented by NN particles in three-dimensional space, a configuration is s=(p_1, ext{...},p_N)\in\mathbb{R}^{3N}, where pip_i is the position of particle ii. The physically plausible configurations form a subset Sreal⊂R3NS_{\mathrm{real}}\subset\mathbb{R}^{3N}; most mathematical configurations need not correspond to physically realizable shapes. DeformGen assumes approximate conditional closure of this plausible set under a sufficiently accurate simulator Φsim\Phi_{\mathrm{sim}}: starting at s∈Sreals\in S_{\mathrm{real}}, evolution under reasonable force field ff and duration Δt\Delta t approximately remains in SrealS_{\mathrm{real}}. The assumption does not say that simulation reliably repairs an invalid initial configuration. Plausibility is therefore relative to simulator fidelity, not guaranteed to match real-world physics exactly.

  3. Knowl 3 — Localized disturbances and forward simulation generate plausible topological variation

    model/method

    Given a valid source configuration s0∈Sreals_0\in S_{\mathrm{real}}, DeformGen creates an augmented configuration by simulating a localized disturbance:

    saug=Φsim(s0,f,Δt).s_{\mathrm{aug}}=\Phi_{\mathrm{sim}}(s_0,f,\Delta t).

    Here ff is a localized force field acting on the deformable object, and Δt\Delta t is the simulation duration. Because the disturbance is followed through the simulator's dynamics rather than applied as an arbitrary particle displacement, the resulting state is intended to retain physical and topological coherence under the simulator model. Localized interactions can produce non-rigid changes such as bending, twisting, folding, or draping. The method does not claim to cover all of SrealS_{\mathrm{real}} or to sample the real-world state distribution; it is a practical way to explore a broader set of plausible states than global rigid transformations.

  4. Knowl 4 — Position warping interpolates particle displacements with a time-varying influence

    equation

    Let piorigp_i^{\mathrm{orig}} and pidefp_i^{\mathrm{def}} be corresponding source and deformed particle positions, respectively, and let δi=pidef−piorig\delta_i=p_i^{\mathrm{def}}-p_i^{\mathrm{orig}}. For an end-effector waypoint xtx_t at trajectory time tt, let nnj(xt)\mathrm{nn}_j(x_t) denote the index of its jjth nearest source particle, for j=1,…,Kj=1,\ldots,K. DeformGen interpolates neighboring particle displacements by inverse distance:

    wt,j=1∥xt−pnnj(xt)orig∥+ε,w~t,j=wt,j∑q=1Kwt,q,d(xt)=∑j=1Kw~t,j δnnj(xt).w_{t,j}=\frac{1}{\lVert x_t-p_{\mathrm{nn}_j(x_t)}^{\mathrm{orig}}\rVert+\varepsilon},\qquad \widetilde{w}_{t,j}=\frac{w_{t,j}}{\sum_{q=1}^{K}w_{t,q}},\qquad d(x_t)=\sum_{j=1}^{K}\widetilde{w}_{t,j}\,\delta_{\mathrm{nn}_j(x_t)}.

    The warped waypoint is

    xtwarp=xt+αtd(xt).x_t^{\mathrm{warp}}=x_t+\alpha_t d(x_t).

    Here KK is the neighborhood size, ε>0\varepsilon>0 is a numerical-stability constant expressed in the same length units as the point-to-waypoint distances, and αt\alpha_t is a dimensionless time-dependent warping weight. The weight may remain constant or decay over the trajectory, allowing the path to follow the changed geometry strongly near grasp and to move back toward the source path later.

  5. Knowl 5 — A local deformation Jacobian adapts end-effector orientation

    equation

    For waypoint xtx_t, use the same local particle neighborhood to define source-relative vectors ℓt,jorig=pnnj(xt)orig−xt\ell_{t,j}^{\mathrm{orig}}=p_{\mathrm{nn}_j(x_t)}^{\mathrm{orig}}-x_t and deformed vectors ℓt,jdef=ℓt,jorig+δnnj(xt)\ell_{t,j}^{\mathrm{def}}=\ell_{t,j}^{\mathrm{orig}}+\delta_{\mathrm{nn}_j(x_t)}. Stack these three-dimensional vectors as columns of XorigX_{\mathrm{orig}} and XdefX_{\mathrm{def}}. DeformGen fits a local linear map and computes its closed-form least-squares solution:

    =X_{\mathrm{def}}X_{\mathrm{orig}}^{\mathsf T}\left(X_{\mathrm{orig}}X_{\mathrm{orig}}^{\mathsf T}\right)^+.$$ Here $J_t\in\mathbb{R}^{3\times3}$, and $^+$ denotes the Moore–Penrose pseudoinverse. The method projects $J_tR_t$ onto $SO(3)$ by singular-value decomposition to obtain an induced rotation $R'_t$, where $R_t$ is the source end-effector rotation. It interpolates between the source and induced rotations using spherical linear interpolation: $R_t^{\mathrm{warp}}=\mathrm{SLERP}(R_t,R'_t,\alpha_t)$. In the tabletop experiments, the orientation update is restricted to rotation about the axis perpendicular to the table to avoid spurious tilting or flipping from noisy Jacobian estimates.
  6. Knowl 6 — Trajectory phases and task-specific warping configurations

    experimental setup

    A synthesized trajectory has approach, grasp, and manipulation phases. The approach is interpolated from the robot reset pose to the grasp pose; the grasp and manipulation portions are transferred by deformation-field warping. The grasp uses a local particle neighborhood because grasp alignment depends on nearby geometry, while manipulation uses all NN object particles to capture global shape changes. The experiments use these settings:

    • Rope routing: grasp K=5K=5; manipulation K=NK=N; linear decay αt=max⁡(0,1−t/T)\alpha_t=\max(0,1-t/T), with TT the trajectory length.
    • Toy packing: grasp K=5K=5; manipulation K=NK=N; no decay, so αt=1\alpha_t=1.
    • Cloth folding: grasp K=10K=10; manipulation K=NK=N; exponential decay αt=e−0.02t\alpha_t=e^{-0.02t}.

    For state generation, the simulator is disturbed by randomized gripper motions in contact with the object. Rope and toy use 180 random steps, translation magnitudes from {0.012,0.006,0.003}\{0.012,0.006,0.003\} m, rotations of ±6∘\pm6^\circ, and rotation probability 0.450.45. Cloth uses 260 steps, translations from {0.018,0.009,0.0045}\{0.018,0.009,0.0045\} m, rotations of ±8∘\pm8^\circ, and rotation probability 0.550.55. Each disturbance sequence is followed by 30–40 simulation steps of stabilization.

  7. Knowl 7 — Policy evaluation shows gains over unaugmented data and rigid-style augmentation

    empirical result

    The authors evaluated success rate on rope routing, toy packing, and cloth folding for ACT, Diffusion Policy (DP), SmolVLA, and π0\pi_0. The four regimes are one source demonstration only (1 Src.), rigid state augmentation with deformation-field warping (SMG*), topological state augmentation with local rigid trajectory transfer (DG*), and full DeformGen (DG). DeformGen has the highest average success rate for three of the four architectures; DP is the exception, where DG* scores higher. The results vary across tasks, so the gains are not uniform.

    Policy Task 1 Src. SMG* DG* DG
    ACT Rope 0.00 68.00 90.00 90.50
    ACT Toy 0.00 73.00 49.00 75.50
    ACT Cloth 4.00 3.50 1.50 11.00
    ACT Average 1.33 48.17 46.83 59.00
    DP Rope 0.00 64.00 64.00 57.50
    DP Toy 0.00 49.50 56.50 54.00
    DP Cloth 7.00 0.50 2.50 0.50
    DP Average 2.33 38.00 41.00 37.33
    SmolVLA Rope 0.00 62.50 88.00 92.00
    SmolVLA Toy 0.00 42.00 49.50 53.50
    SmolVLA Cloth 7.50 16.50 27.50 24.00
    SmolVLA Average 2.50 40.33 55.00 56.50
    pi_0 Rope 0.00 56.00 98.50 99.00
    pi_0 Toy 0.00 10.00 32.50 58.00
    pi_0 Cloth 7.00 17.50 24.00 13.00
    pi_0 Average 2.33 27.83 51.67 56.67

    All entries are success rates in percent. Comparing SMG* with DG tests the change from rigid to topological state augmentation while using deformation-field trajectory transfer; comparing DG* with DG tests deformation-field warping against local rigid transfer on topologically augmented states. These comparisons generally support contributions from both broader state coverage and deformation-aware trajectory transfer, with architecture- and task-dependent effects.

  8. Knowl 8 — Generated states exhibit non-rigid coverage beyond rigid pose variation

    empirical result

    The state-coverage analysis decomposes each augmented configuration relative to its source into a rigid SE(3)SE(3) component, obtained by Procrustes alignment, and a non-rigid residual measured by root-mean-square particle displacement after alignment. Across rope, cloth, and toy, rigidly augmented samples cluster near the source and have negligible non-rigid residuals. DeformGen samples spread more broadly in the combined state representation and show larger non-rigid residuals, supporting the interpretation that the method changes object shape rather than merely generating more rigid placements. The toy has non-zero residuals even for some rigid-style samples because object interactions during stabilization deform it.

  9. Knowl 9 — Policies achieve variable but non-zero success on states where trajectory synthesis failed

    empirical result

    The authors tested whether policies trained on successful augmented demonstrations could act on hard object configurations for which trajectory synthesis had not produced a valid demonstration. The test covered toy packing and cloth folding; rope was excluded because its trajectory-synthesis success rate was nearly 100%. The policies achieved non-zero, architecture-dependent success, indicating some generalization beyond configurations with successful synthesized demonstrations, but the result was not uniformly strong across tasks or methods.

    ACT DP SmolVLA pi0\\pi_0
    Task SMG* DG* DG SMG* DG* DG SMG* DG* DG SMG* DG* DG
    Toy 45.50 37.50 55.50 37.50 55.00 47.00 15.00 15.00 18.50 7.50 18.50 36.50
    Cloth 11.00 6.00 5.50 6.00 5.00 5.00 17.50 16.50 10.00 11.50 19.00 4.00
    Average 28.25 21.75 30.50 21.75 30.00 26.00 16.25 15.75 14.25 9.50 18.75 20.25

    Entries are success rates in percent. SMG* uses rigid state augmentation with deformation-field warping; DG* uses topological state augmentation with local rigid transfer; DG is full DeformGen. The strongest average varies by policy: DG leads for ACT and π0\pi_0, DG* for DP, and SMG* for SmolVLA.

  10. Knowl 10 — Evaluation scope and known limitations

    limitation

    The evaluation is entirely in simulation, using a single xArm7 robot and three tasks: rope routing, stuffed-toy packing, and cloth folding. The authors do not establish transfer to real hardware, bimanual or multi-robot manipulation, or broader deformable-object domains such as dough shaping and surgical tissue manipulation. Trajectory synthesis is also not reliable for every target configuration: correspondence-based warping only approximately preserves task semantics, and large topological changes, contact dynamics, or kinematic constraints can cause failure. The reported synthesis success rates differ substantially by task (99.5% for rope, 60.3% for toy, and 39.5% for cloth), and policies may still fail under extreme deformation or unstable contact.

Coverage note — The rigid-only test-state ablation and the synthetic-data-quantity ablation are omitted: they are secondary checks on compatibility with simpler test distributions and scaling with dataset size, rather than load-bearing parts of the proposed method or its main evaluation.

References

  1. 1.Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al. pi0: A vision-language-action flow model for general robot control. arXiv preprint, 2024.
  2. 2.Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, et al. pi0.5: a vision-language-action model with open-world generalization. arXiv preprint, 2025.
  3. 3.Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. Openvla: An open-source vision-language-action model. arXiv preprint, 2024.
  4. 4.Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, et al. Gr00t n1: An open foundation model for generalist humanoid robots. arXiv preprint, 2025.
  5. 5.Xinyi Chen, Yilun Chen, Yanwei Fu, Ning Gao, Jiaya Jia, Weiyang Jin, Hao Li, Yao Mu, Jiangmiao Pang, Yu Qiao, et al. Internvla-m1: A spatially guided vision-language-action framework for generalist robot policy. arXiv preprint arXiv:2510.13778, 2025.
  6. 6.Wenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang, Xinqiang Yu, Jiazhao Zhang, Runpei Dong, Jiawei He, He Wang, Zhizheng Zhang, et al. Dreamvla: A vision-language-action model dreamed with comprehensive world knowledge. arXiv preprint, 2025.
  7. 7.Wenyao Zhang, Bozhou Zhang, Zekun Qi, Wenjun Zeng, Xin Jin, and Li Zhang. Disentangled robot learning via separate forward and inverse dynamics pretraining. arXiv preprint arXiv:2604.16391, 2026.
  8. 8.Jingwen Sun, Wenyao Zhang, Zekun Qi, Shaojie Ren, Zezhi Liu, Hanxin Zhu, Guangzhong Sun, Xin Jin, and Zhibo Chen. Vla-jepa: Enhancing vision-language-action model with latent world model. arXiv preprint arXiv:2602.10098, 2026.
  9. 9.Zhixuan Liang, Yizhuo Li, Tianshuo Yang, Chengyue Wu, Sitong Mao, Liuao Pei, Xiaokang Yang, Jiangmiao Pang, Yao Mu, and Ping Luo. Discrete diffusion vla: Bringing discrete diffusion to action decoding in vision-language-action policies. arXiv preprint, 2025.
  10. 10.Yao Mu, Tianxing Chen, Shijia Peng, Zanxin Chen, Zeyu Gao, Yude Zou, Lunkai Lin, Zhiqiang Xie, and Ping Luo. Robotwin: Dual-arm robot benchmark with generative digital twins (early version). In ECCV, 2025.
  11. 11.Ajay Mandlekar, Soroush Nasiriany, Bowen Wen, Iretiayo Akinola, Yashraj Narang, Linxi Fan, Yuke Zhu, and Dieter Fox. Mimicgen: A data generation system for scalable robot learning using human demonstrations. In Conference on Robot Learning, pages 1820–1864. PMLR, 2023.
  12. 12.Zhengrong Xue, Shuying Deng, Zhenyang Chen, Yixuan Wang, Zhecheng Yuan, and Huazhe Xu. Demogen: Synthetic demonstration generation for data-efficient visuomotor policy learning. arXiv preprint arXiv:2502.16932, 2025.
  13. 13.Sizhe Yang, Wenye Yu, Jia Zeng, Jun Lv, Kerui Ren, Cewu Lu, Dahua Lin, and Jiangmiao Pang. Novel demonstration generation with gaussian splatting enables robust one-shot manipulation. arXiv preprint arXiv:2504.13175, 2025.
  14. 14.Yuan Xu, Jiabing Yang, Xiaofeng Wang, Yixiang Chen, Zheng Zhu, Bowen Fang, Guan Huang, Xinze Chen, Yun Ye, Qiang Zhang, et al. Egodemogen: Novel egocentric demonstration generation enables viewpoint-robust manipulation. arXiv preprint arXiv:2509.22578, 2025.
  15. 15.Masoud Moghani, Mahdi Azizian, Animesh Garg, Yuke Zhu, Sean Huver, and Ajay Mandlekar. Softmimicgen: A data generation system for scalable robot learning in deformable object manipulation. arXiv preprint arXiv:2603.25725, 2026.
  16. 16.Yunsong Zhou, Hangxu Liu, Xuekun Jiang, Xing Shen, Yuanzhen Zhou, Hui Wang, Baole Fang, Yang Tian, Mulin Yu, Qiaojun Yu, et al. Sim1: Physics-aligned simulator as zero-shot data scaler in deformable worlds. arXiv preprint arXiv:2604.08544, 2026.
  17. 17.Jose Sanchez, Juan-Antonio Corrales, Belhassen-Chedli Bouzgarrou, and Youcef Mezouar. Robotic manipulation and sensing of deformable objects in domestic and industrial applications: a survey. The International Journal of Robotics Research, 37(7):688–716, 2018.
  18. 18.Hang Yin, Anastasia Varava, and Danica Kragic. Modeling, learning, perception, and control methods for deformable object manipulation. Science Robotics, 6(54):eabd8803, 2021.
  19. 19.Kaifeng Zhang, Shuo Sha, Hanxiao Jiang, Matthew Loper, Hyunjong Song, Guangyan Cai, Zhuo Xu, Xiaochen Hu, Changxi Zheng, and Yunzhu Li. Real-to-sim robot policy evaluation with gaussian splatting simulation of soft-body interactions. arXiv preprint arXiv:2511.04665, 2025.
  20. 20.Haoyu Zhao, Cheng Zeng, Linghao Zhuang, Yaxi Zhao, Shengke Xue, Hao Wang, Xingyue Zhao, Zhongyu Li, Kehan Li, Siteng Huang, Mingxiu Chen, Xin Li, Deli Zhao, and Hua Zou. High-fidelity simulated data generation for real-world zero-shot robotic manipulation learning with gaussian splatting. IEEE Robotics and Automation Letters, 11(5):5310–5317, 2026. doi: 10.1109/LRA.2026.3671535.
  21. 21.Stephen James, Zicong Ma, David Rovick Arrojo, and Andrew J. Davison. RLBench: The Robot Learning Benchmark & Learning Environment. arXiv preprint arXiv:1909.12271, 2019.
  22. 22.Yufei Wang, Zhou Xian, Feng Chen, Tsun-Hsuan Wang, Yian Wang, Katerina Fragkiadaki, Zackory Erickson, David Held, and Chuang Gan. RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation. In International Conference on Machine Learning, 2024.
  23. 23.Atsushi Kanehira, Naoki Wake, Kazuhiro Sasabuchi, Jun Takamatsu, and Katsushi Ikeuchi. Rl-driven data generation for robust vision-based dexterous grasping. arXiv preprint arXiv:2504.18084, 2025.
  24. 24.Zoey Chen, Zhao Mandi, Homanga Bharadhwaj, Mohit Sharma, Shuran Song, Abhishek Gupta, and Vikash Kumar. Semantically controllable augmentations for generalizable robot learning. The International Journal of Robotics Research, 44(10-11):1705–1726, 2025.
  25. 25.GigaAI. Gigabrain-0: A world model-powered vision-language-action model. 2025. URL https://arxiv.org/abs/2510.19430.
  26. 26.Zhenyu Jiang, Yuqi Xie, Kevin Lin, Zhenjia Xu, Weikang Wan, Ajay Mandlekar, Linxi Jim Fan, and Yuke Zhu. Dexmimicgen: Automated data generation for bimanual dexterous manipulation via imitation learning. In 2025 IEEE International Conference on Robotics and Automation (ICRA), pages 16923–16930. IEEE, 2025.
  27. 27.Joel Jang, Seonghyeon Ye, Zongyu Lin, Jiannan Xiang, Johan Bjorck, Yu Fang, Fengyuan Hu, Spencer Huang, Kaushil Kundalia, Yen-Chen Lin, et al. Dreamgen: Unlocking generalization in robot learning through neural trajectories. arXiv preprint, 2025.
  28. 28.Ying Li, Xiaobao Wei, Xiaowei Chi, Yuming Li, Zhongyu Zhao, Hao Wang, Ningning Ma, Ming Lu, and Sirui Han. Manipdreamer3d: Synthesizing plausible robotic manipulation video with occupancy-aware 3d trajectory. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 6644–6652, 2026.
  29. 29.Guanhua Ji, Harsha Polavaram, Lawrence Yunliang Chen, Sandeep Bajamahal, Zehan Ma, Simeon Adebola, Chenfeng Xu, and Ken Goldberg. Oxe-auge: A large-scale robot augmentation of oxe for scaling cross-embodiment policy learning. arXiv preprint arXiv:2512.13100, 2025.
  30. 30.Boyang Wang, Haoran Zhang, Shujie Zhang, Jinkun Hao, Mingda Jia, Qi Lv, Yucheng Mao, Zhaoyang Lyu, Jia Zeng, Xudong Xu, et al. Robovip: Multi-view video generation with visual identity prompting augments robot manipulation. arXiv preprint arXiv:2601.05241, 2026.
  31. 31.Chuer Pan, Litian Liang, Dominik Bauer, Eric Cousineau, Benjamin Burchfiel, Siyuan Feng, and Shuran Song. One demo is worth a thousand trajectories: Action-view augmentation for visuomotor policies. In 9th Annual Conference on Robot Learning, 2025.
  32. 32.Justin Yu, Letian Fu, Huang Huang, Karim El-Refai, Rares Andrei Ambrus, Richard Cheng, Muhammad Zubair Irshad, and Ken Goldberg. Real2render2real: Scaling robot data without dynamics simulation or robot hardware, 2025. URL https://arxiv.org/abs/2505.09601.
  33. 33.Yujie Zhao, Hongwei Fan, Di Chen, Shengcong Chen, Liliang Chen, Xiaoqi Li, Guanghui Ren, and Hao Dong. Real2edit2real: Generating robotic demonstrations via a 3d control interface. arXiv preprint arXiv:2512.19402, 2025.
  34. 34.Caelan Garrett, Ajay Mandlekar, Bowen Wen, and Dieter Fox. Skillmimicgen: Automated demonstration generation for efficient skill learning and deployment. arXiv preprint arXiv:2410.18907, 2024.
  35. 35.Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George Drettakis, et al. 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph., 42(4):139–1, 2023.
  36. 36.Xavier Provot et al. Deformation constraints in a mass-spring model to describe rigid cloth behaviour. In Graphics interface, pages 147–147. Canadian Information Processing Society, 1995.
  37. 37.Stéphane Cotin, Hervé Delingette, and Nicholas Ayache. Real-time elastic deformations of soft tissues for surgery simulation. IEEE transactions on Visualization and Computer Graphics, 5(1):62–73, 2002.
  38. 38.Yuanming Hu, Yu Fang, Ziheng Ge, Ziyin Qu, Yixin Zhu, Andre Pradhana, and Chenfanfu Jiang. A moving least squares material point method with displacement discontinuity and two-way rigid body coupling. ACM Transactions on Graphics (TOG), 37(4):1–14, 2018.
  39. 39.Matthias Müller, Bruno Heidelberger, Marcus Hennix, and John Ratcliff. Position based dynamics. Journal of Visual Communication and Image Representation, 18(2):109–118, 2007.
  40. 40.Sergio Orozco, Brandon B. May, Tushar Kusnur, George Konidaris, and Laura Herlant. Learning equivariant neural-augmented object dynamics from few interactions. In Beyond Rigid Worlds: Representing and Interacting with Non-Rigid Objects, 2025. URL https://openreview.net/forum?id=JAiJpFozaD.
  41. 41.Xingyu Lin, Zhiao Huang, Yunzhu Li, Joshua B. Tenenbaum, David Held, and Chuang Gan. Diffskill: Skill abstraction from differentiable physics for deformable object manipulations with tools. In International Conference on Learning Representations (ICLR), 2022.
  42. 42.Haochen Shi, Huazhe Xu, Samuel Clarke, Yunzhu Li, and Jiajun Wu. Robocook: Long-horizon elasto-plastic object manipulation with diverse tools. In Conference on Robot Learning (CoRL), 2023.
  43. 43.Haonan Chen, Yilong Niu, Kaiwen Hou, Shuijing Liu, Yixuan Wang, Yunzhu Li, and Katherine Driggs-Campbell. Predicting object interactions with behavior primitives: An application in stowing tasks. In Conference on Robot Learning (CoRL), 2023.
  44. 44.Isabella Huang, Yashraj Narang, Ruzena Bajcsy, Fabio Ramos, Tucker Hermans, and Dieter Fox. Defgraspsim: Physics-based simulation of grasp outcomes for 3d deformable objects. In IEEE International Conference on Robotics and Automation (ICRA), 2022.
  45. 45.Lijun Han and Hesheng Wang. Robotic manipulation of deformable objects: a comprehensive review. Robotic Intelligence and Automation, pages 1–16, 2026.
  46. 46.Ryan Paul McKennaa and John Oyekan. A perspective on open challenges in deformable object manipulation. arXiv preprint arXiv:2602.22998, 2026.
  47. 47.Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion. In The International Journal of Robotics Research, 2024.
  48. 48.Runpei Dong, Zekun Qi, Linfeng Zhang, Junbo Zhang, Jianjian Sun, Zheng Ge, Li Yi, and Kaisheng Ma. Autoencoders as cross-modal teachers: Can pretrained 2d image transformers help 3d representation learning? In ICLR, 2023.
  49. 49.Weikun Peng, Jun Lv, Yuwei Zeng, Haonan Chen, Siheng Zhao, Jichen Sun, Cewu Lu, and Lin Shao. Tiebot: Learning to knot a tie from visual demonstration through a real-to-sim-to-real approach. arXiv preprint arXiv:2407.03245, 2024.
  50. 50.Kai Wu, Rongkang Chen, Qi Chen, and Weihua Li. Robotic assembly of deformable linear objects via curriculum reinforcement learning. IEEE Robotics and Automation Letters, 2025.
  51. 51.Checheng Yu, Chonghao Sima, Gangcheng Jiang, Hai Zhang, Haoguang Mai, Hongyang Li, Huijie Wang, Jin Chen, Kaiyang Wu, Li Chen, Lirui Zhao, Modi Shi, Ping Luo, Qingwen Bu, Shijia Peng, Tianyu Li, and Yibo Yuan. χ0: Resource-aware robust manipulation via taming distributional inconsistencies. arXiv preprint arXiv:2602.09021, 2026.
  52. 52.Daniel Seita, Aditya Ganapathi, Ryan Hoque, Minho Hwang, Edward Cen, Ajay Kumar Tanwani, Ashwin Balakrishna, Brijen Thananjeyan, Jeffrey Ichnowski, Nawid Jamali, et al. Deep imitation learning of sequential fabric smoothing from an algorithmic supervisor. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2020.
  53. 53.Thomas Weng, Sujay Bajracharya, Yufei Wang, Khush Agrawal, and David Held. Fabricflownet: Bimanual cloth manipulation with a flow-based policy. In Conference on Robot Learning (CoRL), 2022.
  54. 54.Yuran Wang, Ruihai Wu, Yue Chen, Jiarui Wang, Jiaqi Liang, Ziyu Zhu, Haoran Geng, Jitendra Malik, Pieter Abbeel, and Hao Dong. Dexgarmentlab: Dexterous garment manipulation environment with generalizable policy. arXiv preprint arXiv:2505.11032, 2025.
  55. 55.Huy Ha and Shuran Song. Flingbot: The unreasonable effectiveness of dynamic manipulation for cloth unfolding. In Conference on Robot Learning (CoRL), pages 24–33. PMLR, 2022.
  56. 56.Hanxiao Jiang, Hao-Yu Hsu, Kaifeng Zhang, Hsin-Ni Yu, Shenlong Wang, and Yunzhu Li. Phystwin: Physics-informed reconstruction and simulation of deformable objects from videos. ICCV, 2025.
  57. 57.Xingyu Lin, Yufei Wang, Jake Olkin, and David Held. Softgym: Benchmarking deep reinforcement learning for deformable object manipulation. In Conference on Robot Learning, 2020.
  58. 58.Yuanming Hu, Tzu-Mao Li, Luke Anderson, Jonathan Ragan-Kelley, and Frédo Durand. Taichi: a language for high-performance computation on spatially sparse data structures. ACM Transactions on Graphics (TOG), 38(6):1–16, 2019.
  59. 59.Miles Macklin. Warp: A high-performance python framework for gpu simulation and graphics. In NVIDIA GPU Technology Conference (GTC), volume 3, 2022.
  60. 60.Yang Tian, Yuyin Yang, Yiman Xie, Zetao Cai, Xu Shi, Ning Gao, Hangxu Liu, Xuekun Jiang, Zherui Qiu, Feng Yuan, et al. Interndata-a1: Pioneering high-fidelity synthetic data for pre-training generalist policy. arXiv preprint arXiv:2511.16651, 2025.
  61. 61.John Schulman, Jonathan Ho, Cameron Lee, and Pieter Abbeel. Learning from demonstrations through the use of non-rigid registration. In Robotics Research: The 16th International Symposium ISRR, pages 339–354. Springer, 2016.
  62. 62.Tony Z Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn. Learning fine-grained bimanual manipulation with low-cost hardware. arXiv preprint, 2023.
  63. 63.Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion. The International Journal of Robotics Research, 2023.
  64. 64.Mustafa Shukor, Dana Aubakirova, Francesco Capuano, Pepijn Kooijmans, Steven Palma, Adil Zouitine, Michel Aractingi, Caroline Pascal, Martino Russi, Andres Marafioti, et al. Smolvla: A vision-language-action model for affordable and efficient robotics. arXiv preprint, 2025.

Citation

MLA
Lin, Z., et al. “DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning”. arXiv, 2026, http://arxiv.org/abs/2606.25939v2.
APA
Lin, Z., Zhang, W., Zhang, Y., Qi, Z., Lin, J., Zhu, H., Yang, J., Chen, Z., Mu, Y., Yang, X., Jin, X., & Zeng, W. (2026). DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning. arXiv. http://arxiv.org/abs/2606.25939v2
Chicago
Lin, Z., W. Zhang, Y. Zhang, et al. 2026. “DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning”. arXiv. http://arxiv.org/abs/2606.25939v2.
Harvard
Lin, Z. et al. (2026) “DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2606.25939v2.
Vancouver
1. Lin Z, Zhang W, Zhang Y, et al (2026) DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning. arXiv

BibTeX

@article{lin2026deformgen,
  title = {DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning},
  author = {Lin, Zili and Zhang, Wenyao and Zhang, Yuyang and Qi, Zekun and Lin, Junyan and Zhu, Hanxin and Yang, Jiaolong and Chen, Zhibo and Mu, Yao and Yang, Xiaokang and Jin, Xin and Zeng, Wenjun},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2606.25939v2},
  eprint = {2606.25939}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/