Seamless Human Motion Composition with Blended Positional Encodings
Germán BarqueroSergio EscaleraCristina Palmero
Presents FlowMDM, a diffusion framework that schedules absolute and relative positional encodings during denoising to generate long, continuous multi-action human animations from sequential text prompts without requiring postprocessing or specialized multi-action training data.
Generating realistic, continuous three-dimensional human motion from text descriptions is critical for emerging applications in virtual reality, video gaming, and robotics. However, current automated methods primarily synthesize isolated, short movements lasting only a few seconds under a single instruction. Extending these models to produce long sequences driven by a series of distinct actions usually leads to unnatural, abrupt transitions or requires computationally intensive postprocessing and stitching. The core difficulty stems from existing training datasets lacking long-duration sequences with annotations for transitions between consecutive actions.
The article introduces FlowMDM, a generative diffusion framework designed to evaluate and demonstrate the simultaneous synthesis of long, seamless human motion compositions directly from sequential text prompts without requiring manual stitching, postprocessing, or datasets with annotated transitions.
To achieve this, the researchers evaluated their approach on two benchmark datasets, HumanML3D and Babel, comparing it against established sequential and diffusion-based baselines. The methodology introduces two key architectural features: Blended Positional Encodings and Pose-Centric Cross-Attention. Blended Positional Encodings utilize absolute positional information early in the iterative generation process to establish overall motion coherence, gradually shifting to relative positional information in later stages to ensure fluid transitions between actions. The Pose-Centric Cross-Attention mechanism prevents conflicting conditions from tangling at action boundaries, allowing the system to train on datasets with only one text description per sequence while still executing multi-prompt sequences during inference. Additionally, the authors developed two physical metrics based on jerk—the time derivative of acceleration—to quantitatively evaluate transition smoothness.
The findings demonstrate substantial improvements over existing methods across multiple performance dimensions. First, the proposed framework achieved superior motion quality and text alignment, reducing motion distribution error scores by over 60% compared to prior diffusion models on HumanML3D. Second, it delivered noticeably smoother and more realistic transitions, minimizing sudden acceleration spikes where baseline methods exhibited severe motion glitches. Third, computational efficiency improved substantially, requiring roughly 16% to 47% fewer pose-wise computation steps than existing diffusion-based blending approaches. Finally, the model successfully generalized to cyclic, repetitive movements (such as walking or waving) over extended durations without degrading motion consistency.
These results demonstrate that long, controlled character animations can be produced efficiently in a single processing pass, lowering computational overhead and eliminating manual cleanup. This provides a direct path to reducing animation production costs and accelerating asset generation pipelines in gaming and virtual simulations. The findings also highlight that standard generative evaluation metrics fail to detect physical irregularities like abrupt jerks, underscoring the necessity of motion-derivative metrics for quality control.
Organizations developing interactive avatars or automated animation tools should consider adopting blended encoding architectures to streamline motion generation pipelines without investing in expensive, specialized transition datasets. For production deployment, teams should tune the schedule between absolute and relative encoding steps—allocating roughly 10% of initial steps to absolute positions—to optimally balance text fidelity against transition smoothness.
A recognized limitation of the current architecture is that the initial absolute generation phase models motion segments independently, leaving the system without high-level intent planning across far-apart actions. While confidence in the benchmarked smoothness and accuracy metrics is high, further validation is required before extending the architecture to multi-modal control signals (such as combining text with scene geometry or audio) or applying it to real-time physical robotics.
- Paper: Human Motion Diffusion Model, Guy Tevet et al. (2022). Provides the foundational Motion Diffusion Model (MDM) architecture that FlowMDM builds upon and extends for long-sequence composition.
- Paper: Flexible Diffusion Modeling of Long Videos, William Harvey et al. (2022). Introduces relative and flexible positional conditioning strategies in temporal diffusion models for long sequence generation.
- Paper: Flow Matching for Generative Modeling, Yaron Lipman et al. (2023). Establishes the core mathematical formulations of flow matching that underpin continuous flow-based generative modeling.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Presents the fundamental denoising diffusion probabilistic framework utilized by diffusion-based motion generation models.
- Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Formulates implicit diffusion sampling techniques that enable efficient deterministic generation in continuous denoising pipelines.
- Paper: Stochastic Interpolants: A Unifying Framework for Flows and Diffusions, Michael S. Albergo et al. (2025). Generalizes flow and diffusion mechanisms into a unified stochastic interpolant framework across finite-time generation tasks.
- Paper: Mean Flows for One-step Generative Modeling, Zhengyang Geng et al. (2025). Advances velocity-field modeling by introducing mean flows to achieve efficient single-step generation.
- Paper: Generative Modeling via Drifting, Mingyang Deng et al. (2026). Explores distribution-drifting dynamics as an alternative formulation for single-step generative modeling.
- Paper: There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation, Gabe Guo et al. (2026). Extends continuous diffusion processes to bidirectional multimodal translation bridges across heterogeneous domains.
