MetaMorph: Learning Universal Controllers with Transformers

Agrim GuptaLinxi FanSurya GanguliLi Fei-Fei

article2022ICLR154 citations

Develops MetaMorph, a Transformer architecture that conditions control policies on robot morphology to enable universal motor control with zero-shot generalization across unseen modular robot designs and tasks.

Listen

Modern robotics often relies on training individual controllers for specific hardware designs from scratch. While modular robots allow flexible physical configurations, the vast number of potential body designs makes training a dedicated controller for every design impractical. This article introduces MetaMorph, a transformer-based framework that learns a universal controller capable of operating across diverse modular robot designs. By treating physical morphology as an input modality alongside real-time sensor data, MetaMorph enables large-scale pre-training across varied robot bodies to achieve cross-body generalization.

The authors evaluated the framework in physical simulations using 100 distinct 3D robot morphologies across multiple environments, including flat ground, irregular obstacles, and unpredictable variable terrain. To train diverse robots effectively, the system employed dynamic replay buffer balancing, which prevents easier-to-control robots from dominating the training process. The pre-trained controller was tested against conventional graph neural networks and individual task-specific models, assessing its immediate zero-shot performance and its adaptation to new designs and tasks.

The findings show that MetaMorph consistently outperforms graph neural network baselines by two to three times in accumulated reward and matches the performance of separately trained individual controllers while using five times fewer training samples. The framework demonstrated strong zero-shot generalization across 400 variations in physical dynamics and dimensions without requiring randomized training. When transferred to completely new robot shapes and distinct movement tasks, the pre-trained policy reached peak performance two to three times faster than training from scratch. Analysis revealed that the transformer naturally organizes joint movements into coordinated functional groups, mirroring biological motor synergies.

These results indicate that robotics can successfully adopt the pre-training and fine-tuning paradigms that have driven major breakthroughs in language and vision AI. For organizations deploying modular or variable robotic hardware, this approach significantly lowers training costs, shortens deployment timelines, and enhances operational resilience against mechanical wear or physical damage. Future development should focus on unifying design generation with controller pre-training, expanding capabilities beyond locomotion into complex manipulation, and validating the controller on real physical hardware.

  • Paper: Multi-Game Decision Transformers, Kuang-Huei Lee et al. (2022). Demonstrates how Transformer architectures can be scaled across diverse decision-making datasets to learn multi-task control policies, establishing key foundations for MetaMorph's universal controller concept.
  • Paper: Perceiver: General Perception with Iterative Attention, Andrew Jaegle et al. (2021). Pioneers the general perception architecture for processing arbitrary input modalities as generic token sequences within Transformers, motivating the treatment of robot morphology as an input modality.
  • Paper: Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning, Tianhe Yu et al. (2019). Establishes standard benchmark protocols for multi-task and meta-reinforcement learning across distinct robotic configurations, providing foundational task contexts for modular robotic policies.
  • Paper: Sim-to-Real Transfer of Robotic Control with Dynamics Randomization, Xue Bin Peng et al. (2017). Introduces robust dynamics randomization for transferring modular simulated control policies to varied physical variations, underpinning MetaMorph's zero-shot generalization training paradigms.
  • Paper: A Survey of Transformers, Tianyang Lin et al. (2021). Provides a comprehensive architectural taxonomy of Transformer variants, self-attention mechanics, and cross-domain pretraining paradigms utilized in universal policy learning.
Cover for MetaMorph: Learning Universal Controllers with Transformers

Abstract

Multiple domains like vision, natural language, and audio are witnessing tremendous progress by leveraging Transformers for large scale pre-training followed by task specific fine tuning. In contrast, in robotics we primarily train a single robot for a single task. However, modular robot systems now allow for the flexible combination of general-purpose building blocks into task optimized morphologies. However, given the exponentially large number of possible robot morphologies, training a controller for each new design is impractical. In this work, we propose MetaMorph, a Transformer based approach to learn a universal controller over a modular robot design space. MetaMorph is based on the insight that robot morphology is just another modality on which we can condition the output of a Transformer. Through extensive experiments we demonstrate that large scale pre-training on a variety of robot morphologies results in policies with combinatorial generalization capabilities, including zero shot generalization to unseen robot morphologies. We further demonstrate that our pre-trained policy can be used for sample-efficient transfer to completely new robot morphologies and tasks.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Learning a Universal Controller
  • 3.1 The UNIMAL Design Space
  • 3.2 Joint Policy Optimization
  • 4 MetaMorph
  • 4.1 Fusing Proprioceptive States and Morphology Representations
  • 4.2 Morphology Aware Transformer
  • 4.3 Dynamic Replay Buffer Balancing
  • 5 Experiments
  • 5.1 Experimental Setup
  • 5.2 Baselines and Ablations
  • 5.3 Zero-Shot Generalization
  • 5.4 Sample Efficient Transfer Learning
  • 5.5 Emergent Motor Synergies
  • 6 Conclusion
  • 7 Acknowledgement
  • 8 Reproducibility Statement
  • References
  • A Implementation Details
  • A.1 Input Observations
  • A.2 Environments
  • A.3 Training Hyperparameters
  • A.4 Evaluation Methodology
  • B Additional Experiments
  • B.1 Baselines and Ablations
  • B.2 Position Embedding
  • B.3 Kinematic tree traversal order
  • C Limitations And Future Work

Knowls

  1. Knowl 1 — MetaMorph Architecture and State Representation

    model/method

    MetaMorph is a Transformer-based universal controller architecture for modular robots with varying kinematic structures, numbers of degrees of freedom (DoFs), and dynamic properties. A modular robot's kinematic tree (or directed acyclic graph) is converted into a 1D sequence of tokens by traversing the kinematic tree in depth-first search (DFS) order starting at the root torso module.

    At each time step tt, the state of robot kk with NN constituent modules is represented as sk=(slk,sgk)s^k = (s^k_l, s^k_g), where slk=(sl1k,…,slNk)s^k_l = (s^k_{l1}, \dots, s^k_{lN}) contains local module-level states and sgks^k_g contains global sensory information (such as a 2D terrain heightmap):

    1. Local Module State (slik=(smik,spik)s^k_{li} = (s^k_{mi}, s^k_{pi})): The morphology representation smiks^k_{mi} includes module shape parameters (radius, height), material density, local geometric orientation and position relative to the parent module, joint type, joint range, joint axis, actuator type, and motor gear ratio. The proprioceptive state spiks^k_{pi} contains the module's 3D Cartesian position, 4D quaternion orientation, 3D linear velocity, 3D angular velocity, and the joint position and velocity in generalized coordinates. Observation vectors sliks^k_{li} are zero-padded to a fixed dimensionality MM.

    2. Module Embeddings: A single-layer MLP with weights We∈RM×DW_e \in \mathbb{R}^{M \times D} linearly projects each local module vector to dimension DD, and learned 1D position embeddings Wpos∈RN×DW_{\text{pos}} \in \mathbb{R}^{N \times D} are added:

    m0=[ϕ(sl1k;We);… ;ϕ(slNk;We)]+Wposm_0 = [\phi(s^k_{l1}; W_e); \dots; \phi(s^k_{lN}; W_e)] + W_{\text{pos}}

    1. Transformer Processing: The sequence is processed through LL Transformer encoder layers with multi-head self-attention (MSA) and layer normalization (LN):

    mℓ′=MSA(LN(mℓ−1))+mℓ−1,ℓ=1,…,Lm'_\ell = \text{MSA}(\text{LN}(m_{\ell-1})) + m_{\ell-1}, \quad \ell = 1, \dots, L

    mℓ=MLP(LN(mℓ′))+mℓ′,ℓ=1,…,Lm_\ell = \text{MLP}(\text{LN}(m'_\ell)) + m'_\ell, \quad \ell = 1, \dots, L

    1. Decoding: Global sensory features g=γ(sgk;Wg)g = \gamma(s^k_g; W_g) (extracted via a 2-layer MLP γ\gamma) are concatenated with the per-module Transformer output representations mL,im_{L, i} to predict per-joint actuation distributions:

    μ(sik)=ϕ(mL,i,g;Wd),πθ(ak∣sk)=N({μ(sik)}i=1N,Σ)\mu(s^k_i) = \phi(m_{L, i}, g; W_d), \quad \pi_\theta(a^k | s^k) = \mathcal{N}\left(\{\mu(s^k_i)\}_{i=1}^N, \Sigma\right)

    where Σ\Sigma is a fixed diagonal covariance matrix. The critic network estimates the total state value by averaging scalar value predictions across all modules.

  2. Knowl 2 — Dynamic Replay Buffer Balancing

    model/method

    When jointly training a universal controller over a diverse population of modular robot morphologies, passively stable robots stay upright longer and accumulate more transitions per episode, while unstable morphologies fall early and collect fewer transitions. Under standard on-policy reinforcement learning with early termination, this causes a "rich gets richer" pathology where easier morphologies dominate policy optimization.

    To balance experience across diverse morphologies, Dynamic Replay Buffer Balancing adjusts the sampling probability PkP_k of robot kk during rollout collection at training iteration τ\tau. Let EkτE^\tau_k denote an exponentially smoothed performance metric:

    Ekτ=αEkτ+(1−α)Ek(τ−1)E^\tau_k = \alpha E^\tau_k + (1 - \alpha) E^{(\tau-1)}_k

    where α∈[0,1]\alpha \in [0, 1] is a smoothing factor (set to α=0.1\alpha = 0.1). The sampling probability distribution across all KK robots in the training pool is given by:

    Pk=(Ek)β∑i=1K(Ei)βP_k = \frac{(E_k)^\beta}{\sum_{i=1}^K (E_i)^\beta}

    where β\beta controls prioritization intensity (β=0\beta = 0 corresponds to uniform sampling, and β=1.0\beta = 1.0 is used in practice). In implementation, EiE_i is set to 1000episode_lengthi\frac{1000}{\text{episode\_length}_i}, where 10001000 is the maximum possible episode length, causing robots with shorter episode lengths to be sampled more frequently.

  3. Knowl 3 — MetaMorph Joint Training Algorithm for Modular Robots

    algorithm

    Joint optimization of the MetaMorph controller policy πθ\pi_\theta and value function VϕV_\phi across a pool of KK modular robots combines Proximal Policy Optimization (PPO) with Dynamic Replay Buffer Balancing.

    Input: Policy network πθ\pi_\theta, value network VϕV_\phi, replay buffer RR, training pool of modular robots KK, total iterations NiterN_{\text{iter}}, rollout workers NworkersN_{\text{workers}}, PPO epochs NepochsN_{\text{epochs}}, warmup iterations NwarmupN_{\text{warmup}}
    Initialize: Network parameters θ,ϕ\theta, \phi; sampling distribution PP over KK as uniform
    for i=1,2,…,Niteri = 1, 2, \dots, N_{\text{iter}} do
        for j=1,2,…,Nworkersj = 1, 2, \dots, N_{\text{workers}} in parallel do
            Sample robot k∈Kk \in K according to distribution PP
            Collect rollout trajectory τj≡{st,at,st+1,rt+1}\tau_j \equiv \{s_t, a_t, s_{t+1}, r_{t+1}\} under policy πθ\pi_\theta
            R←R∪τjR \leftarrow R \cup \tau_j
        end for
        for j=1,2,…,Nepochsj = 1, 2, \dots, N_{\text{epochs}} do
            Sample minibatch rr from buffer RR
            Update θ\theta via PPO policy objective on rr
            Update ϕ\phi via PPO value objective on rr
        end for
        if i≥Nwarmupi \ge N_{\text{warmup}} then
            Update sampling probabilities Pk∝(1000/Ek)βP_k \propto (1000 / E_k)^\beta based on smoothed inverse episode lengths in RR
        end if
        Clear buffer RR
    end for
  4. Knowl 4 — Joint Multi-Morphology Pre-Training Performance

    empirical result

    Joint pre-training of the MetaMorph controller was evaluated across 100 distinct 3D UNIMAL robot morphologies (15–20 degrees of freedom) across three locomotion environments in MuJoCo: Flat Terrain (FT), Variable Terrain (VT, featuring stochastically generated hills, steps, and rubble), and Obstacles (cuboids of varying sizes):

    1. Comparison to Single-Robot Upper Bound: MetaMorph matched the average return achieved by training 100 separate multi-layer perceptron (MLP) controllers independently for each morphology on FT and Obstacles, and came within close range on VT (without saturating at 10810^8 environment steps).
    2. Sample Efficiency: MetaMorph trained on 100 robots jointly within 10810^8 total environment interactions, achieving a 5×5\times sample efficiency improvement over independent single-robot MLP training (5×1065 \times 10^6 steps per robot, totaling 5×1085 \times 10^8 steps).
    3. Comparison to Graph Neural Networks (GNN): A GNN baseline based on NerveNet with matching encoder/decoder pipelines, parameter count (∼3.3M\sim 3.3\text{M}), and observation inputs saturated at an average reward level 2×2\times to 3×3\times lower than MetaMorph across all environments.
    4. Effect of Dynamic Buffer Balancing: Dynamic replay buffer balancing prevented training collapse on harder morphologies, increasing the proportion of robots reaching ≥75%\ge 75\% of the individual MLP baseline performance by 10%10\% to 15%15\% across all three environments.
  5. Knowl 5 — Zero-Shot Generalization to Kinematic and Dynamic Variations

    empirical result

    Zero-shot transfer performance of pre-trained MetaMorph controllers was tested across 400 newly generated robot variants for four dynamic properties (armature, density, damping, motor gear ratio) and two kinematic properties (limb radius/height shape parameters, joint angle ranges) without applying domain randomization during pre-training:

    1. Dynamic Parameter Variations: MetaMorph exhibited robust zero-shot generalization across armature [0.1,2.0][0.1, 2.0], density [0.8,1.2]×default[0.8, 1.2] \times \text{default}, damping [0.01,5.0][0.01, 5.0], and motor gear [0.8,1.2]×default[0.8, 1.2] \times \text{default}, significantly outperforming both GNN baselines and morphology-agnostic Transformer baselines (MetaMorph-NM), with the largest margins occurring on Variable Terrain and Obstacles.
    2. Kinematic Parameter Variations: MetaMorph successfully maintained locomotion across module shape perturbations (limb radius [0.03,0.05][0.03, 0.05] and height [0.15,0.45][0.15, 0.45]). For severe joint angle range alterations, performance declined relative to base morphologies due to altered gait constraints, but remained substantially higher than GNN and MetaMorph-NM models.
  6. Knowl 6 — Emergence of Motor Synergies in Self-Attention Representations

    empirical result

    To analyze how MetaMorph coordinates high-dimensional action spaces (up to ∼1600\sim 1600 total controllable DoFs across 100 morphologies), the stable rank of the Transformer attention weight matrix Aℓ∈Rm×mA_\ell \in \mathbb{R}^{m \times m} at layer ℓ\ell was tracked during locomotion:

    sr(Aℓ)=∥Aℓ∥F2∥Aℓ∥22=∑iσi2σmax⁡2\text{sr}(A_\ell) = \frac{\|A_\ell\|_F^2}{\|A_\ell\|_2^2} = \frac{\sum_i \sigma_i^2}{\sigma_{\max}^2}

    where σi\sigma_i are the singular values of AℓA_\ell and σmax⁡\sigma_{\max} is the maximum singular value.

    Across walking cycles, sr(Aℓ)\text{sr}(A_\ell) remained low (oscillating between approximately 1.251.25 and 2.02.0) and exhibited periodic transitions between distinct low-rank states. The corresponding attention heatmaps revealed synchronized co-activation patterns across groups of limbs (dense vertical attention bands). This dimensionality reduction in the attention mechanism demonstrates the unsupervised emergence of motor synergies, wherein the network dynamically couples modular joints into coordinated functional groups based on the combination of morphology inputs smks^k_m and proprioceptive states spks^k_p.

  7. Knowl 7 — Sample-Efficient Transfer to Unseen Morphologies and Novel Tasks

    empirical result

    Fine-tuning pre-trained MetaMorph policies was evaluated under two transfer learning scenarios:

    1. Transfer to Unseen Morphologies on Same Tasks: When evaluated on a test set of 100 entirely novel UNIMAL robot topologies on Flat Terrain, Variable Terrain, and Obstacles, fine-tuning pre-trained MetaMorph required 2×2\times to 3×3\times fewer training iterations to reach equivalent asymptotic reward compared to learning controllers from scratch.
    2. Transfer to Unseen Morphologies on Novel Tasks: Pre-trained controllers were transferred to 100 new morphologies evaluated on two new environments: an Escape task (navigating out of a bowl-shaped terrain surrounded by hills, transferring from Variable Terrain) and an Obstacles (Cylinders) task (navigating cylindrical obstacles with different diameter distributions, transferring from cuboid Obstacles). In both cases, pre-trained MetaMorph produced immediate non-zero zero-shot transfer and achieved the target reward level 2×2\times to 3×3\times faster than tabula rasa training.
  8. Knowl 8 — Ablation Analysis of Morphological Conditioning and Positional Encodings

    empirical result

    Ablation studies on the Flat Terrain benchmark identified the critical components of the MetaMorph architecture:

    1. Morphological Conditioning: MetaMorph-NM (excluding morphology representation smks^k_m, feeding only proprioceptive spks^k_p and global sgks^k_g) and MetaMorph-NMT (excluding smks^k_m but providing a one-hot task ID context token) completely failed to learn locomotion policies across diverse morphologies.
    2. Learned vs. Fixed Positional Embeddings: Removing learned position embeddings (MetaMorph-NPE) caused a large drop in reward. Providing hand-designed limb IDs directly in the morphology state (MetaMorph-HPE) improved upon MetaMorph-NPE but still underperformed learned position embeddings (MetaMorph).
    3. Amorpheus Feature Comparison (MetaMorph-AO): Providing limb IDs, kinematics, and joint angle ranges without full morphological specifications (module dimensions, relative spatial orientations, joint axes, motor gears) achieved substantially lower reward than MetaMorph.
    4. Multi-Task MLP: A parameter-matched 6-layer MLP receiving all observation inputs failed to learn effective multi-robot control policies.
  9. Knowl 9 — Hyperparameters for MetaMorph Policy and PPO Training

    data/table

    The following hyperparameters define the Transformer architecture and Proximal Policy Optimization (PPO) training setup across all baseline models and environments, parameterized to approximately 3.3×1063.3 \times 10^6 parameters:

    Category Hyperparameter Value
    PPO Discount factor γ\gamma 0.99
    PPO GAE parameter λ\lambda 0.95
    PPO PPO clipping parameter ϵ\epsilon 0.2
    PPO Policy epochs per update 8
    PPO Batch size 5120
    PPO Entropy coefficient 0.01
    PPO Reward / Obs normalization Yes
    PPO Reward / Obs clipping [−10,10][-10, 10]
    PPO Timesteps per rollout 2560
    PPO Parallel workers / Environments 16 / 32
    PPO Total training timesteps 1×1081 \times 10^8
    PPO Optimizer / Initial learning rate Adam / 0.0003
    PPO Learning rate schedule Linear warmup (5 iters) + Cosine decay
    PPO Gradient clipping (l2l_2 norm) 0.5
    PPO Value loss coefficient 0.5
    Transformer Number of layers (LL) 5
    Transformer Attention heads 1
    Transformer Embedding dimension (DD) 128
    Transformer Feedforward hidden dimension 1024
    Transformer Activation function ReLU
    Transformer Dropout 0.1
  10. Knowl 10 — Kinematic and Dynamic Parameter Ranges for Generalization Evaluation

    data/table

    The zero-shot generalization test set consists of 400 newly sampled variants for each parameter type, generated by perturbing the training morphologies according to the following ranges:

    Category Variation Type Sampled Value / Range
    Kinematics Limb cylinder radius [0.03,0.05][0.03, 0.05]
    Kinematics Limb cylinder height [0.15,0.45][0.15, 0.45]
    Kinematics Joint angle limits (∘^\circ) {(−30,0),(0,30),(−30,30),(−45,45),(−45,0),(0,45),\{(-30, 0), (0, 30), (-30, 30), (-45, 45), (-45, 0), (0, 45),
    (−60,0),(0,60),(−60,60),(−90,0),(0,90),(−60,30),(−30,60)}(-60, 0), (0, 60), (-60, 60), (-90, 0), (0, 90), (-60, 30), (-30, 60)\}
    Dynamics Armature [0.1,2.0][0.1, 2.0]
    Dynamics Density [0.8,1.2]×default limb density[0.8, 1.2] \times \text{default limb density}
    Dynamics Damping [0.01,5.0][0.01, 5.0]
    Dynamics Motor gear [0.8,1.2]×default motor gear[0.8, 1.2] \times \text{default motor gear}

    For joint angle perturbations, ranges are sampled from the discrete set such that each joint retains at least 50%50\% overlap with its initial range to avoid creating physically untrainable robots.

  11. Knowl 11 — Limitations and Sensitivity to DFS Traversal Ordering

    limitation

    The MetaMorph framework exhibits three primary limitations:

    1. Two-Stage Requirement: The framework assumes prior availability of a pool of diverse, functional morphologies (e.g., evolved via external morphological search), requiring a sequential pipeline of morphology creation followed by joint controller training rather than unified co-optimization.
    2. Zero-Shot Drop on Topological Shifts: While MetaMorph generalizes zero-shot to dynamic and parameter variations, direct zero-shot transfer to completely new kinematic graph topologies experiences a significant performance decrease, necessitating fine-tuning.
    3. DFS Ordering Sensitivity: Because tokens are sequenced via depth-first traversal of the kinematic tree, the policy is sensitive to traversal permutations. Reversing the visitation order of sibling nodes at the same tree depth drops zero-shot policy performance by ∼75%\sim 75\% unless data augmentation (randomizing visitation order of same-depth nodes during training) is explicitly applied.

Coverage note — None was omitted; all key architectural components, algorithms, empirical evaluations, ablation findings, parameter tables, and limitations contributed by the paper are fully covered.

References

  1. 1.Maruan Al-Shedivat, Trapit Bansal, Yuri Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel. Continuous adaptation via meta-learning in nonstationary and competitive environments. arXiv preprint arXiv:1710.03641, 2017.
  2. 2.Peter W Battaglia, Jessica B Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, et al. Relational inductive biases, deep learning, and graph networks. arXiv preprint arXiv:1806.01261, 2018.
  3. 3.Nikolai Bernstein. The co-ordination and regulation of movements. The co-ordination and regulation of movements, 1966.
  4. 4.Charlie Blake, Vitaly Kurin, Maximilian Igl, and Shimon Whiteson. Snowflake: Scaling gnns to high-dimensional continuous control via parameter freezing. arXiv preprint arXiv:2103.01009, 2021.
  5. 5.Josh Bongard. Why morphology matters. The horizons of evolutionary robotics, 6:125–152, 2014.
  6. 6.Rodney A Brooks. New approaches to robotics. Science, 253(5025):1227–1232, 1991.
  7. 7.Michaela Bruton and Nicholas O’Dwyer. Synergies in coordination: A comprehensive overview of neural, computational, and behavioral approaches. Journal of Neurophysiology, 120(6):2761–2774, 2018.
  8. 8.Tao Chen, Adithyavairavan Murali, and Abhinav Gupta. Hardware conditioned policies for multi-robot transfer learning. In NIPS, 2018.
  9. 9.Nick Cheney, Robert MacCurdy, Jeff Clune, and Hod Lipson. Unshackling evolution: Evolving soft robots with multiple materials and a powerful generative encoding. SIGEVOlution, 7(1): 11–23, August 2014. doi: 10.1145/2661735.2661737. URL https://doi.org/10.1145/2661735.2661737.
  10. 10.Nick Cheney, Josh Bongard, Vytas SunSpiral, and Hod Lipson. Scalable co-optimization of morphology and control in embodied machines. Journal of The Royal Society Interface, 15(143): 20170937, 2018.
  11. 11.Andrew M Dai and Quoc V Le. Semi-supervised sequence learning. Advances in neural information processing systems, 28:3079–3087, 2015.
  12. 12.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  13. 13.Ioannis Exarchos, Yifeng Jiang, Wenhao Yu, and C Karen Liu. Policy transfer via kinematic domain randomization and adaptation. arXiv preprint arXiv:2011.01891, 2020.
  14. 14.Toshio Fukuda and Seiya Nakagawa. Dynamically reconfigurable robotic system. In ICRA, pp. 1581–1586. IEEE, 1988.
  15. 15.Ali Ghadirzadeh, Xi Chen, Petra Poklukar, Chelsea Finn, Mårten Bj ˚ orkman, and Danica Kragic. ¨ Bayesian meta-learning for few-shot policy adaptation across robotic platforms. arXiv preprint arXiv:2103.03697, 2021.
  16. 16.Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 580–587, 2014.
  17. 17.Agrim Gupta, Silvio Savarese, Surya Ganguli, and Li Fei-Fei. Embodied intelligence via learning and evolution. Nature communications, 12(1):5721, 2021.
  18. 18.David Ha. Reinforcement learning for improving agent design. Artificial life, 25(4):352–365, 2019.
  19. 19.Leland H Hartwell, John J Hopfield, Stanislas Leibler, and Andrew W Murray. From molecular to modular cell biology. Nature, 402(6761):C47–C52, 1999.
  20. 20.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, pp. 9729–9738, 2020.
  21. 21.Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger. Deep reinforcement learning that matters. In Thirty-Second AAAI Conference on Artificial Intelligence, 2018.
  22. 22.Wenlong Huang, Igor Mordatch, and Deepak Pathak. One policy to control them all: Shared modular policies for agent-agnostic control. In ICML, pp. 4455–4464. PMLR, 2020.
  23. 23.Donald Joseph Hejna III, Pieter Abbeel, and Lerrel Pinto. Task-agnostic morphology evolution. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=CGQ6ENUMX6.
  24. 24.Nadav Kashtan and Uri Alon. Spontaneous evolution of modularity and network motifs. Proceedings of the National Academy of Sciences, 102(39):13773–13778, 2005.
  25. 25.Thomas Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. ArXiv, abs/1609.02907, 2017.
  26. 26.Ashish Kumar, Zipeng Fu, Deepak Pathak, and Jitendra Malik. Rma: Rapid motor adaptation for legged robots. arXiv preprint arXiv:2107.04034, 2021.
  27. 27.Vitaly Kurin, Maximilian Igl, Tim Rocktaschel, Wendelin Boehmer, and Shimon Whiteson. My ̈ body is a cage: the role of morphology in graph-based incompatible control. In ICLR, 2021.
  28. 28.Joonho Lee, Jemin Hwangbo, Lorenz Wellhausen, Vladlen Koltun, and Marco Hutter. Learning quadrupedal locomotion over challenging terrain. Science robotics, 5(47), 2020.
  29. 29.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer Normalization. arXiv e-prints, art. arXiv:1607.06450, July 2016.
  30. 30.T. Liao, G. Wang, B. Yang, R. Lee, K. Pister, S. Levine, and R. Calandra. Data-efficient learning of morphology and controller for a microrobot. In 2019 International Conference on Robotics and Automation (ICRA), pp. 2488–2494, 2019. doi: 10.1109/ICRA.2019.8793802.
  31. 31.Ajay Mandlekar, Danfei Xu, Josiah Wong, Soroush Nasiriany, Chen Wang, Rohun Kulkarni, Li Fei-Fei, Silvio Savarese, Yuke Zhu, and Roberto Mart´ın-Mart´ın. What matters in learning from offline human demonstrations for robot manipulation. In 5th Annual Conference on Robot Learning, 2021. URL https://openreview.net/forum?id=JrsfBJtDFdI.
  32. 32.Nithin Mathews, Anders Lyhne Christensen, Rehan O’Grady, Francesco Mondada, and Marco Dorigo. Mergeable nervous systems for robots. Nature communications, 8(1):1–7, 2017.
  33. 33.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high-performance deep learning library. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alche-Buc, E. Fox, and ´ R. Garnett (eds.), Advances in Neural Information Processing Systems 32, pp. 8024–8035. Curran Associates, Inc., 2019.
  34. 34.Deepak Pathak, Christopher Lu, Trevor Darrell, Phillip Isola, and Alexei A Efros. Learning to control self-assembling morphologies: A study of generalization via modularity. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alche-Buc, E. Fox, and R. Garnett (eds.), ´ Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019.
  35. 35.Xue Bin Peng, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel. Sim-to-real transfer of robotic control with dynamics randomization. In 2018 IEEE international conference on robotics and automation (ICRA), pp. 3803–3810. IEEE, 2018.
  36. 36.Rolf Pfeifer and Christian Scheier. Understanding intelligence. MIT press, 2001.
  37. 37.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training (2018), 2018.
  38. 38.Michael Rubenstein, Alejandro Cornejo, and Radhika Nagpal. Programmable self-assembly in a thousand-robot swarm. Science, 345(6198):795–799, 2014.
  39. 39.Alvaro Sanchez-Gonzalez, Nicolas Heess, Jost Tobias Springenberg, Josh Merel, Martin Riedmiller, Raia Hadsell, and Peter Battaglia. Graph networks as learnable physics engines for inference and control. In ICML, pp. 4470–4479. PMLR, 2018.
  40. 40.Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. The graph neural network model. IEEE transactions on neural networks, 20(1):61–80, 2008.
  41. 41.Charles Schaff, David Yunis, Ayan Chakrabarti, and Matthew R Walter. Jointly learning to construct and control agents using deep reinforcement learning. In ICRA, pp. 9798–9805. IEEE, 2019.
  42. 42.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal Policy Optimization Algorithms. arXiv e-prints, art. arXiv:1707.06347, July 2017.
  43. 43.Karl Sims. Evolving 3d morphology and behavior by competition. Artificial life, 1(4):353–372, 1994.
  44. 44.Emanuel Todorov. Optimality principles in sensorimotor control. Nature neuroscience, 7(9):907–915, 2004.
  45. 45.Emanuel Todorov and Zoubin Ghahramani. Analysis of the synergies underlying complex hand manipulation. In The 26th Annual International Conference of the IEEE Engineering in Medicine and Biology Society, volume 2, pp. 4637–4640. IEEE, 2004.
  46. 46.Emanuel Todorov, Tom Erez, and Yuval Tassa. Mujoco: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, pp. 5026–5033. IEEE, 2012.
  47. 47.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NIPS, 2017.
  48. 48.Tingwu Wang, Renjie Liao, Jimmy Ba, and Sanja Fidler. Nervenet: Learning structured policy with graph neural networks. In ICLR, 2018.
  49. 49.Tingwu Wang, Yuhao Zhou, Sanja Fidler, and Jimmy Ba. Neural graph evolution: Automatic robot design. In ICLR, 2019.
  50. 50.Ronald J Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8(3-4):229–256, 1992.
  51. 51.Mark Yim, Wei-Min Shen, Behnam Salemi, Daniela Rus, Mark Moll, Hod Lipson, Eric Klavins, and Gregory S Chirikjian. Modular self-reconfigurable robot systems [grand challenges of robotics]. IEEE Robotics & Automation Magazine, 14(1):43–52, 2007.
  52. 52.Allan Zhao, Jie Xu, Mina Konakovic-Lukovi ´ c, Josephine Hughes, Andrew Spielberg, Daniela Rus, ´ and Wojciech Matusik. Robogrammar: graph grammar for terrain-optimized robot design. ACM Transactions on Graphics (TOG), 39(6):1–16, 2020.

Citation

MLA
Gupta, A., et al. “MetaMorph: Learning Universal Controllers with Transformers”. arXiv, 2022, http://arxiv.org/abs/2203.11931v1.
APA
Gupta, A., Fan, L., Ganguli, S., & Fei-Fei, L. (2022). MetaMorph: Learning Universal Controllers with Transformers. arXiv. http://arxiv.org/abs/2203.11931v1
Chicago
Gupta, A., L. Fan, S. Ganguli, and L. Fei-Fei. 2022. “MetaMorph: Learning Universal Controllers with Transformers”. arXiv. http://arxiv.org/abs/2203.11931v1.
Harvard
Gupta, A. et al. (2022) “MetaMorph: Learning Universal Controllers with Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.11931v1.
Vancouver
1. Gupta A, Fan L, Ganguli S, Fei-Fei L (2022) MetaMorph: Learning Universal Controllers with Transformers. arXiv

BibTeX

@article{gupta2022metamorph,
  title = {MetaMorph: Learning Universal Controllers with Transformers},
  author = {Gupta, Agrim and Fan, Linxi and Ganguli, Surya and Fei-Fei, Li},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.11931v1},
  eprint = {2203.11931}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors