Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video

Xiao LiQi ChenXiulian PengKai YuXie ChenYan Lu

article2025ICCV2 citations

Proposes a self-supervised diffusion framework that uses low-bitrate vector quantization as an information bottleneck to cleanly separate dynamic motion from static video content for motion transfer and generation.

Listen

Separating dynamic motion from static content in video data is essential for video analysis, realistic animation, and content editing. However, existing methods typically rely on rigid assumptions or task-specific constraints, such as predefined optical flows, facial landmarks, or 3D parametric models. These assumptions limit representation flexibility, introduce visual artifacts, and restrict systems to narrow domains. Developing a general framework that robustly disentangles motion and content without hand-crafted domain priors remains a major practical challenge.

The article introduces and evaluates Bitrate-Controlled Diffusion (BCD), a self-supervised video representation framework. The main objective is to demonstrate that video data can be effectively separated into flexible, implicit motion and content representations by combining an information bottleneck with a generative diffusion model, requiring minimal task-specific assumptions.

The authors implemented a system that extracts global clip-level content features and per-frame motion features using a transformer encoder. To prevent information leakage between motion and content, the motion features are constrained using a low-bitrate vector quantization bottleneck. These representations then condition a latent denoising diffusion model to reconstruct the original video. The framework was evaluated on the large-scale LRS3 talking-head video dataset (over 400 hours of video) across motion transfer and motion generation tasks, and its cross-domain adaptability was validated on the LPC Sprites animated character dataset.

The evaluation produced several key findings. First, on cross-identity motion transfer, the proposed framework achieved superior visual quality and alignment over existing baselines, yielding the best image fidelity (an FID score of 86.0 compared to 98.5–106.8 for prior methods) and lower motion transfer error (3.13 versus 3.94–36.1). Second, a user study confirmed that human evaluators rated the method highest in identity preservation, motion consistency, and overall visual quality. Third, constraining the motion pathway with an optimal target bitrate (around 4 kbps for talking-head videos) proved critical: higher bitrates caused content information to leak into motion, while lower bitrates degraded visual quality. Fourth, the discrete motion space successfully enabled direct autoregressive video generation using a standard sequence model. Finally, the framework achieved 100% attribute classification accuracy on the 2D cartoon dataset, demonstrating strong generalizability across diverse visual distributions.

These findings indicate that low-bitrate information bottlenecks offer a viable alternative to complex, specialized visual priors for video decomposition. By eliminating specialized modules like keypoint detectors or 3D face models, developers can streamline video synthesis pipelines, reduce technical debt across distinct video domains, and achieve higher-fidelity editing. The authors note that the framework strictly adheres to ethical AI principles and is intended for general representation learning and forgery detection research rather than deceptive media generation.

Organizations evaluating this approach should consider piloting bitrate-controlled diffusion for video editing and generation tasks that require high visual realism across diverse styles. Implementation teams should optimize the bitrate bottleneck for their specific target datasets, as bitrate demands vary with video complexity. Further technical work is recommended to optimize inference latency—currently around 30 seconds for a 50-frame video on high-end hardware—and to introduce post-hoc controls for interactive motion editing.

Confidence in the reported outcomes is supported by extensive quantitative benchmarks, mesh-based 3D geometric error metrics, and human perceptual evaluations. Nonetheless, limitations include high computational training requirements, potential performance drops on extreme out-of-distribution static elements, and slight residual video flickering that warrants further refinement.

Cover for Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video

Abstract

We propose a novel and general framework to disentangle video data into its dynamic motion and static content components. Our proposed method is a self-supervised pipeline with less assumptions and inductive biases than previous works: it utilizes a transformer-based architecture to jointly generate flexible implicit features for frame-wise motion and clip-wise content, and incorporates a low-bitrate vector quantization as an information bottleneck to promote disentanglement and form a meaningful discrete motion space. The bitrate-controlled latent motion and content are used as conditional inputs to a denoising diffusion model to facilitate self-supervised representation learning. We validate our disentangled representation learning framework on real-world talking head videos with motion transfer and auto-regressive motion generation tasks. Furthermore, we also show that our method can generalize to other types of video data, such as pixel sprites of 2D cartoon characters. Our work presents a new perspective on self-supervised learning of disentangled video representations, contributing to the broader field of video analysis and generation.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Method
  • 3.1 Content and Motion Extraction
  • 3.2 Feature Disentanglement
  • 3.3 Conditional Diffusion Decoder
  • 3.4 Model Training
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Experiment Setup on LRS3
  • 4.3 Main Results on Talking Heads
  • 4.4 Discussions
  • 4.5 Results on Sprites dataset
  • 5 Conclusion
  • References
  • A Implementation Details
  • A.1 Video Representation Learning
  • A.2 Auto-regressive motion generation
  • B Additional Results
  • B.1 Motion Transfer and Generation
  • B.2 Additional Ablation Studies
  • C Evaluation Metrics

Knowls

  1. Knowl 1 — Bitrate-Controlled Diffusion (BCD) Video Disentanglement Architecture

    model/method

    Bitrate-Controlled Diffusion (BCD) is a self-supervised video representation learning framework designed to separate video sequences into sequence-level static content representations and frame-wise dynamic motion representations without relying on task-specific explicit priors (such as facial landmarks, optical flow, or 3D morphable face models).

    Given an input video sequence of TT frames, each frame is first tokenized into a latent feature vector ztz_t using a pre-trained image Variational Autoencoder (such as Stable Diffusion 2.0 VAE), yielding a latent sequence z={zt∣t∈[1,T]}\mathbf{z} = \{z_t \mid t \in [1, T]\}. BCD consists of three core components:

    1. A Disentangling Transformer Encoder T\mathcal{T}: Simultaneously extracts a global content embedding cc and per-frame motion latents m={mt∣t∈[1,T]}m = \{m_t \mid t \in [1, T]\} by prepending learnable query prefix tokens to the frame latent sequence z\mathbf{z}.
    2. A Bitrate-Controlled Information Bottleneck: Passes the motion latents mtm_t through a grouped vector quantizer (Group VQ) regularized by an explicit entropy constraint Htarget\mathcal{H}_{target}. This information bottleneck limits the transmission rate of the motion stream, preventing content information from leaking into motion representations while preserving sufficient capacity for motion dynamics.
    3. A Conditional Diffusion Decoder D\mathcal{D}: A Diffusion Transformer (DiT) conditioned on both the clip-level content cc (via feature concatenation) and the frame-level motion m^t\hat{m}_t (via Adaptive Group Normalization) to reconstruct the original latent sequence z\mathbf{z} via a denoising objective.
  2. Knowl 2 — Content and Motion Extraction via Learnable Query Prefix

    model/method

    BCD extracts content and motion features from video latent tokens using a unified Transformer encoder T\mathcal{T} (implemented via a 12-layer T5 architecture with relative positional encodings, hidden size 512, feedforward dimension 2048, and 8 attention heads).

    To aggregate sequence-wide content information without arbitrarily selecting fixed reference frames or relying on simple temporal pooling, a sequence of KK learnable query tokens q∈RK×d\mathbf{q} \in \mathbb{R}^{K \times d} is prepended as a prefix to the per-frame latent token sequence {zt∣t∈[1,T]}\{z_t \mid t \in [1, T]\} before passing into T\mathcal{T}. The queries q\mathbf{q} are optimized across the entire training dataset to aggregate time-invariant content information across diverse viewpoints and frames.

    The output sequence of T\mathcal{T} is split along the token length:

    • The first KK prefix tokens form the sequence-level global content representation c∈RK×Ccc \in \mathbb{R}^{K \times C_c}.
    • The remaining TT tokens form the per-frame motion sequence m={mt∣t∈[1,T]}∈RT×Cmm = \{m_t \mid t \in [1, T]\} \in \mathbb{R}^{T \times C_m}.
  3. Knowl 3 — Bitrate-Controlled Group Vector Quantization as an Information Bottleneck

    model/method

    To prevent information leakage (content leaking into motion features) and information preference (insufficient dynamic mutual information), BCD applies a low-bitrate Group Vector Quantization (Group VQ) bottleneck to the motion latents mt∈Rcm_t \in \mathbb{R}^c.

    Each motion vector mtm_t at time tt is partitioned into NN groups {mti∈Rc/N∣i∈[1,N]}\{m_t^i \in \mathbb{R}^{c/N} \mid i \in [1, N]\}. Each group ii is quantized using a dedicated codebook Ei={eji}j=1KE^i = \{e_j^i\}_{j=1}^K containing KK code entries of dimension cvq=c/Nc_{vq} = c/N. The sampling distribution μti\mu_t^i over the codebook entries is calculated via distance-based Gumbel-Softmax:

    dti=[ℓ(mti,e1i),ℓ(mti,e2i),…,ℓ(mti,eKi)]d_t^i = \left[ \ell(m_t^i, e_1^i), \ell(m_t^i, e_2^i), \dots, \ell(m_t^i, e_K^i) \right] μti=GumbelSoftmax(−α⋅dti)\mu_t^i = \text{GumbelSoftmax}(-\alpha \cdot d_t^i)

    where ℓ(⋅,⋅)\ell(\cdot, \cdot) denotes the L2L_2 distance and α\alpha is a scaling factor. During training, quantized codes m^ti\hat{m}_t^i are sampled using the Gumbel reparameterization trick, with temperature annealed over training; during inference, codes are selected deterministically via arg⁡max⁡(μti)\arg\max(\mu_t^i).

    To enforce compressibility, the entropy of the quantized motion feature is regulated. Across a training batch B\mathcal{B}, the average sampling histogram is approximated by:

    μavgi=1∣B∣∑t∈Bμti\mu_{avg}^i = \frac{1}{|\mathcal{B}|} \sum_{t \in \mathcal{B}} \mu_t^i H^(m^ti∣Ei)=−∑j=1Kμavg,jilog⁡(μavg,ji)\hat{\mathcal{H}}(\hat{m}_t^i \mid E^i) = -\sum_{j=1}^K \mu_{avg, j}^i \log(\mu_{avg, j}^i)

    The total per-frame motion bitrate of the model is Hmodel=∑i=1NH^(m^ti∣Ei)\mathcal{H}_{model} = \sum_{i=1}^N \hat{\mathcal{H}}(\hat{m}_t^i \mid E^i), and an explicit loss is applied to constrain Hmodel\mathcal{H}_{model} towards a target bitrate Htarget\mathcal{H}_{target}.

  4. Knowl 4 — BCD Training Objective and Cross-Driven Strategy

    model/method

    The BCD model is optimized end-to-end using a rate-distortion objective composed of a latent diffusion reconstruction loss Ld\mathcal{L}_d and a vector quantization bitrate loss LVQ\mathcal{L}_{VQ}:

    L=Ld+λLVQ\mathcal{L} = \mathcal{L}_d + \lambda \mathcal{L}_{VQ} Ld=MSE(z,z~)\mathcal{L}_d = \text{MSE}(z, \tilde{z}) LVQ=MSE(Hmodel,Htarget)\mathcal{L}_{VQ} = \text{MSE}(\mathcal{H}_{model}, \mathcal{H}_{target})

    where zz is the ground-truth latent, z~\tilde{z} is the reconstructed latent predicted by the conditional diffusion network, Hmodel\mathcal{H}_{model} is the estimated per-frame motion entropy, Htarget\mathcal{H}_{target} is the predefined target bitrate, and the balancing hyperparameter is set to λ=0.04\lambda = 0.04.

    To prevent the model from collapsing into trivial identity mappings where motion or content streams encode duplicate information, BCD uses a cross-driven training strategy: each 4-second training video clip is divided along the temporal axis into two equal halves with identical semantic content but distinct motion dynamics. The global content feature cc is extracted from the first half, while the motion feature sequence mm is extracted from the second half; the diffusion decoder is conditioned on this pair to reconstruct the second half of the clip.

  5. Knowl 5 — Conditional Latent Diffusion Decoder with Adaptive Group Normalization

    model/method

    BCD synthesizes video frames using a conditional latent Diffusion Transformer (DiT-B/4 for talking heads, DiT-S/8 for Sprites) trained under the Elucidating the Design Space of Diffusion Models (EDM) framework.

    To maintain temporal coherence across synthesized video latents, temporal self-attention blocks are inserted between DiT spatial transformer blocks at layer indices [0,2,4,6,8,10][0, 2, 4, 6, 8, 10].

    The conditioning signals are injected through two distinct mechanisms:

    1. Content Conditioning: The sequence-level content token cc is concatenated directly with the noisy input latent ztz_t along the spatial token dimension.
    2. Motion and Timestep Conditioning: The frame-wise quantized motion feature mtm_t modulates the intermediate layer activations hh of each DiT block via Adaptive Group Normalization (AdaGN):

    AdaGN(h,mt,t)=ms⋅(ts⋅GroupNorm(h)+tb)+mb\text{AdaGN}(h, m_t, t) = m_s \cdot \left( t_s \cdot \text{GroupNorm}(h) + t_b \right) + m_b

    where (ts,tb)(t_s, t_b) and (ms,mb)(m_s, m_b) are linear projections of the sinusoidal diffusion timestep embedding of timestep tt and the motion feature mtm_t, respectively.

  6. Knowl 6 — 3D Mesh-Based Disentanglement Evaluation Metrics for Video

    definition

    To evaluate video disentanglement in talking head videos independent of 2D alignment artifacts and viewpoint-correlated biases of 2D face recognition embeddings (CSIM), BCD utilizes dense parametric 3D face mesh reconstruction M(β,θ,ψ,t)\mathcal{M}(\beta, \theta, \psi, t), where β∈Rc\beta \in \mathbb{R}^c represents identity shape coefficients, θ∈Rn×3\theta \in \mathbb{R}^{n \times 3} is per-frame Euler head pose, ψ∈Rn×m\psi \in \mathbb{R}^{n \times m} is per-frame blendshape expression coefficients, and t∈Rn×3t \in \mathbb{R}^{n \times 3} is per-frame translation vectors. From these fitted parameters, three complementary disentanglement metrics are defined:

    1. Shape Error (ese_s): Measures identity preservation independently of pose and expression by zeroing out motion parameters: Midentity=M(β,0,0,0)M_{identity} = \mathcal{M}(\beta, 0, 0, 0) es=MSE(Midentityref,Midentitydst)e_s = \text{MSE}(M_{identity}^{ref}, M_{identity}^{dst})

    2. Motion Error (eme_m): Measures motion transfer fidelity independently of facial identity by zeroing out identity coefficients: Mmotion=M(0,θ,ψ,t)M_{motion} = \mathcal{M}(0, \theta, \psi, t) em=MSE(Mmotionref,Mmotiondst)e_m = \text{MSE}(M_{motion}^{ref}, M_{motion}^{dst})

    3. Cross Transfer Error (ece_c): Measures composite cross-identity reenactment accuracy by evaluating full composite meshes reconstructed using reference identity βref\beta^{ref} combined with driving motion (θdst,ψdst,tdst)(\theta^{dst}, \psi^{dst}, t^{dst}): Mfull=M(β,θ,ψ,t)M_{full} = \mathcal{M}(\beta, \theta, \psi, t) ec=MSE(Mfullref,Mfulldst)e_c = \text{MSE}(M_{full}^{ref}, M_{full}^{dst})

  7. Knowl 7 — Cross-Identity Motion Transfer Benchmark on LRS3 Dataset

    data/table

    BCD was evaluated on cross-identity motion transfer using the LRS3 talking-head dataset (256x256 resolution, 25 fps) against landmark/flow-based methods (FOMM, MCNET, LivePortrait), 3D parametric prior reenactment (HyperReenact), and linear motion basis decomposition (LIA).

    Method FID↓\downarrow CSIM↑\uparrow Identity error↓\downarrow (×10−1\times 10^{-1}) Motion error↓\downarrow (×10−2\times 10^{-2}) Cross error↓\downarrow (×10−2\times 10^{-2})
    FOMM 98.5 0.76 0.75 24.3 24.1
    MCNET 98.6 0.76 0.85 23.9 23.6
    HyperReenact 106.8 0.58 0.57 3.94 4.68
    LIA 104.4 0.71 0.57 36.1 34.1
    LivePortrait 100.3 0.69 0.66 24.6 23.7
    Ours (BCD) 86.0 0.69 0.41 3.13 3.67

    Despite using no face-specific priors or keypoints, BCD achieves the lowest FID (86.0), lowest identity error (0.41×10−10.41 \times 10^{-1}), lowest motion error (3.13×10−23.13 \times 10^{-2}), and lowest cross error (3.67×10−23.67 \times 10^{-2}). In a user study of 18 participants scoring 15 video sets (1-5 scale), BCD obtained top scores across Identity Preservation (4.10 vs. runner-up 3.66), Motion Consistency (4.30 vs. runner-up 4.01), and Visual Quality (4.00 vs. runner-up 3.72).

  8. Knowl 8 — Ablation on Motion Bitrate, Reference Frames, and Training Strategies

    data/table

    Ablation experiments conducted on the LRS3 dataset assess the impact of the motion target bitrate Htarget\mathcal{H}_{target}, single-frame reference conditioning, bitrate loss LVQ\mathcal{L}_{VQ}, and cross-driven training:

    Setup / Target Bitrate FID↓\downarrow CSIM↑\uparrow Shape error↓\downarrow (×10−1\times 10^{-1}) Motion error↓\downarrow (×10−2\times 10^{-2}) Cross error↓\downarrow (×10−2\times 10^{-2})
    2.0 kbps (80 bits/frame) 88.5 0.71 0.34 5.26 5.68
    4.0 kbps (160 bits/frame) 86.0 0.69 0.41 3.13 3.67
    6.0 kbps (240 bits/frame) 87.6 0.68 0.56 3.23 4.13
    8.0 kbps (320 bits/frame) 89.3 0.66 0.49 3.04 3.74
    4.0 kbps (single ref. frame) 87.9 0.69 0.47 3.13 3.81
    4.0 kbps (w/o cross-driven) 120.1 0.64 0.58 41.5 40.4
    4.0 kbps (w/o bitrate loss) 92.7 0.70 0.59 4.81 6.21

    A motion target bitrate of 4.0 kbps achieves the optimal trade-off between image fidelity and disentanglement (minimal cross error of 3.67×10−23.67 \times 10^{-2}). Removing the cross-driven training strategy causes severe disentanglement collapse, increasing cross error from 3.67×10−23.67 \times 10^{-2} to 40.4×10−240.4 \times 10^{-2} and FID to 120.1. Conditioning on only a single content frame maintains robust performance (FID 87.9, Cross error 3.81×10−23.81 \times 10^{-2}).

  9. Knowl 9 — Autoregressive Video Motion Generation from Discrete Motion Tokens

    model/method

    The discrete motion space learned by BCD enables generative autoregressive video modeling. A GPT-2 architecture (12 layers, hidden size 1024, 8 attention heads) is trained directly on sequences of quantized motion tokens extracted by a pre-trained BCD model.

    During training:

    1. The sequence-level content embedding from a single reference frame is prepended to the motion token sequence as an initial prompt.
    2. The multi-codebook discrete tokens are modeled by modifying GPT-2's output head to produce multi-group logits corresponding to each codebook group.
    3. The model is optimized using average cross-entropy loss over all codebook groups to predict the motion tokens at timestep t+1t+1 given tokens up to timestep tt.
    4. Video temporal downsampling (by a factor of 2 with 50% probability) is applied as data augmentation to improve generation across subtle and large motion dynamics.

    At inference time, motion token sequences sampled autoregressively from GPT-2 are passed alongside the reference content embedding into BCD's conditional diffusion decoder to generate video sequences.

  10. Knowl 10 — Disentanglement Accuracy on LPC Sprites Synthetic Dataset

    empirical result

    To evaluate generalization beyond natural talking heads, BCD was applied to the 2D synthetic LPC Sprites dataset (64x64 resolution, 8-frame clips, 3 actions, 3 directions, 4 content attributes with 6 variations each). For Sprites, BCD operated directly in pixel space with a target motion bitrate of Htarget=6\mathcal{H}_{target} = 6 (150 bps) and a single codebook of 64 entries.

    Using a pre-trained attribute classification network, cross-driven motion transfer evaluated on the test set achieved 100.00% attribute classification accuracy across all categories:

    • Action (Motion): 100.00%
    • Skin Color (Content): 100.00%
    • Pants Color (Content): 100.00%
    • Top Color (Content): 100.00%
    • Hair Color (Content): 100.00%

    This matches the state-of-the-art performance of C-DSVAE while validating that bitrate-controlled diffusion transfers effectively to discrete pixel-art domains.

  11. Knowl 11 — Limitations of Bitrate-Controlled Diffusion Disentanglement

    limitation

    BCD exhibits four key limitations identified by the authors:

    1. Data Dependency: Because the framework minimizes inductive biases and domain-specific priors, it requires substantial volumes of training data to learn effective disentanglement.
    2. Dataset-Specific Bitrate Tuning: The optimal target bitrate Htarget\mathcal{H}_{target} acting as the information bottleneck varies across data domains and resolutions (e.g., 4 kbps for LRS3 vs. 150 bps for Sprites) and must be tuned empirically per dataset.
    3. Temporal Flickering: A minor degree of frame-to-frame video flickering remains detectable even after fine-tuning with temporal attention layers in the diffusion decoder.
    4. Sensitivity to Out-of-Distribution Inputs: Real-world training datasets exhibit high dynamic variation relative to static variation; consequently, the model's ability to isolate static content can degrade when provided with completely out-of-distribution video inputs.

Coverage note — None omitted; all primary architectural components, training formulations, evaluation metrics, empirical benchmarks, ablations, autoregressive generation extensions, and stated limitations are covered.

References

  1. 1.Alessandro Achille and Stefano Soatto. Emergence of invariance and disentanglement in deep representations. The Journal of Machine Learning Research, 19(1):1947–1980, 2018.
  2. 2.Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman. Lrs3-ted: a large-scale dataset for visual speech recognition. arXiv preprint arXiv:1809.00496, 2018.
  3. 3.Alexander A Alemi, Ian Fischer, Joshua V Dillon, and Kevin Murphy. Deep variational information bottleneck. arXiv preprint arXiv:1612.00410, 2016.
  4. 4.Alexei Baevski, Steffen Schneider, and Michael Auli. vqwav2vec: Self-supervised learning of discrete speech representations. arXiv preprint arXiv:1910.05453, 2019.
  5. 5.Junwen Bai, Weiran Wang, and Carla P Gomes. Contrastively disentangled sequential variational autoencoder. Advances in Neural Information Processing Systems, 34: 10105–10118, 2021.
  6. 6.Nimrod Berman, Ilan Naiman, and Omri Azencot. Multifactor sequential disentanglement via structured koopman autoencoders. arXiv preprint arXiv:2303.17264, 2023.
  7. 7.Nimrod Berman, Ilan Naiman, Idan Arbiv, Gal Fadlon, and Omri Azencot. Sequential disentanglement by extracting static information from a single sequence element. arXiv preprint arXiv:2406.18131, 2024.
  8. 8.Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22563–22575, 2023.
  9. 9.Stella Bounareli, Christos Tzelepis, Vasileios Argyriou, Ioannis Patras, and Georgios Tzimiropoulos. Hyperreenact: One-shot reenactment via jointly learning to refine and retarget faces. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023.
  10. 10.Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators. 2024.
  11. 11.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020.
  12. 12.Thomas M Cover. Elements of information theory. John Wiley & Sons, 1999.
  13. 13.Jiankang Deng, Jia Guo, Xue Niannan, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In CVPR, 2019.
  14. 14.Yu Deng, Jiaolong Yang, Sicheng Xu, Dong Chen, Yunde Jia, and Xin Tong. Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, pages 0–0, 2019.
  15. 15.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  16. 16.Michail Christos Doukas, Stefanos Zafeiriou, and Viktoriia Sharmanska. Headgan: One-shot neural head synthesis and editing. In Proceedings of the IEEE/CVF International conference on Computer Vision, pages 14398–14407, 2021.
  17. 17.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12873–12883, 2021.
  18. 18.Yue Gao, Yuan Zhou, Jinglu Wang, Xiao Li, Xiang Ming, and Yan Lu. High-fidelity and freely controllable talking head video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5609–5619, 2023.
  19. 19.Yue Gao, Jiahao Li, Lei Chu, and Yan Lu. Implicit motion function. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19278–19289, 2024.
  20. 20.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. Advances in neural information processing systems, 27, 2014.
  21. 21.Philip-William Grassal, Malte Prinzler, Titus Leistner, Carsten Rother, Matthias Nießner, and Justus Thies. Neural head avatars from monocular rgb videos. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 18653–18664, 2022.
  22. 22.Jean-Bastien Grill, Florian Strub, Florent Altche, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. Advances in neural information processing systems, 33:21271–21284, 2020.
  23. 23.Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. Vector quantized diffusion model for text-to-image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10696–10706, 2022.
  24. 24.Jianzhu Guo, Dingyun Zhang, Xiaoqiang Liu, Zhizhou Zhong, Yuan Zhang, Pengfei Wan, and Di Zhang. Liveportrait: Efficient portrait animation with stitching and retargeting control. arXiv preprint arXiv:2407.03168, 2024.
  25. 25.Jun Han, Martin Renqiang Min, Ligong Han, Li Erran Li, and Xuan Zhang. Disentangled recurrent wasserstein autoencoder. arXiv preprint arXiv:2101.07496, 2021.
  26. 26.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9729–9738, 2020.
  27. 27.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017.
  28. 28.Charlie Hewitt, Fatemeh Saleh, Sadegh Aliakbarian, Lohit Petikam, Shideh Rezaeifar, Louis Florentin, Zafiirah Hosenie, Thomas J Cashman, Julien Valentin, Darren Cosker, and Tadas Baltruevsaitis. Look ma, no markers: holistic performance capture without the hassle. ACM Transactions on Graphics (TOG), 43(6), 2024.
  29. 29.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  30. 30.Fa-Ting Hong and Dan Xu. Implicit identity representation conditioned memory compensation network for talking head video generation. In ICCV, 2023.
  31. 31.Cong Huang, Jiahao Li, Bin Li, Dong Liu, and Yan Lu. Neural compression-based feature learning for video restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5872–5881, 2022.
  32. 32.Drew A Hudson, Daniel Zoran, Mateusz Malinowski, Andrew K Lampinen, Andrew Jaegle, James L McClelland, Loic Matthey, Felix Hill, and Alexander Lerchner. Soda: Bottleneck diffusion models for representation learning. arXiv preprint arXiv:2311.17901, 2023.
  33. 33.Xue Jiang, Xiulian Peng, Yuan Zhang, and Yan Lu. Disentangled feature learning for real-time neural speech coding. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE, 2023.
  34. 34.Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. Advances in Neural Information Processing Systems, 35:26565–26577, 2022.
  35. 35.Jiahao Li, Bin Li, and Yan Lu. Deep contextual video compression. Advances in Neural Information Processing Systems, 34:18114–18125, 2021.
  36. 36.Jiahao Li, Bin Li, and Yan Lu. Hybrid spatial-temporal entropy modelling for neural video compression. In Proceedings of the 30th ACM International Conference on Multimedia, 2022.
  37. 37.Jiahao Li, Bin Li, and Yan Lu. Neural video compression with diverse contexts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22616–22626, 2023.
  38. 38.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023.
  39. 39.Rui Li, Yiheng Zhang, Zhaofan Qiu, Ting Yao, Dong Liu, and Tao Mei. Motion-focused contrastive learning of video representations. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2105–2114, 2021.
  40. 40.Tianhong Li, Dina Katabi, and Kaiming He. Selfconditioned image generation via generating representations. arXiv preprint arXiv:2312.03701, 2023.
  41. 41.Yingzhen Li and Stephan Mandt. Disentangled sequential autoencoder. arXiv preprint arXiv:1803.02991, 2018.
  42. 42.Tao Liu, Feilong Chen, Shuai Fan, Chenpeng Du, Qi Chen, Xie Chen, and Kai Yu. Anitalker: animate vivid and diverse talking faces through identity-decoupled facial motion encoding. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 6696–6705, 2024.
  43. 43.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  44. 44.Arun Mallya, Ting-Chun Wang, and Ming-Yu Liu. Implicit warping for animation with image sets. Advances in Neural Information Processing Systems, 35:22438–22450, 2022.
  45. 45.Ilan Naiman, Nimrod Berman, and Omri Azencot. Sample and predict your latent: modality-free sequential disentanglement via contrastive estimation. In International Conference on Machine Learning, pages 25694–25717. PMLR, 2023.
  46. 46.William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4195–4205, 2023.
  47. 47.Zhiliang Peng, Li Dong, Hangbo Bao, Qixiang Ye, and Furu Wei. Beit v2: Masked image modeling with vector-quantized visual tokenizers. arXiv preprint arXiv:2208.06366, 2022.
  48. 48.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  49. 49.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020.
  50. 50.Scott E Reed, Yi Zhang, Yuting Zhang, and Honglak Lee. Deep visual analogy-making. Advances in neural information processing systems, 28, 2015.
  51. 51.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022.
  52. 52.Xihua Sheng, Jiahao Li, Bin Li, Li Li, Dong Liu, and Yan Lu. Temporal context mining for learned video compression. IEEE Transactions on Multimedia, 2022.
  53. 53.Aliaksandr Siarohin, Stephane Lathuiliere, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. First order motion model for image animation. Advances in neural information processing systems, 32, 2019.
  54. 54.Mathieu Cyrille Simon, Pascal Frossard, and Christophe De Vleeschouwer. Sequential representation learning via staticdynamic conditional disentanglement. In European Conference on Computer Vision, pages 110–126. Springer, 2024.
  55. 55.Naftali Tishby, Fernando C Pereira, and William Bialek. The information bottleneck method. arXiv preprint physics/0004057, 2000.
  56. 56.Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017.
  57. 57.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  58. 58.Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu. One-shot free-view neural talking-head synthesis for video conferencing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10039–10049, 2021.
  59. 59.Yaohui Wang, Di Yang, Francois Bremond, and Antitza Dantcheva. Latent image animator: Learning to animate images via latent space navigation. In International Conference on Learning Representations, 2021.
  60. 60.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, et al. Huggingface’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771, 2019.
  61. 61.Erroll Wood, Tadas Baltruevsaitis, Charlie Hewitt, Sebastian Dziadzio, Thomas J Cashman, and Jamie Shotton. Fake it till you make it: face analysis in the wild using synthetic data alone. In Proceedings of the IEEE/CVF international conference on computer vision, pages 3681–3691, 2021.
  62. 62.Erroll Wood, Tadas Baltruevsaitis, Charlie Hewitt, Matthew Johnson, Jingjing Shen, Nikola Milosavljevic, Daniel Wilde, Stephan Garbin, Toby Sharp, Ivan Stojiljkovic, et al. 3d face reconstruction with dense landmarks. In European Conference on Computer Vision, pages 160–177. Springer, 2022.
  63. 63.Chenfei Wu, Jian Liang, Lei Ji, Fan Yang, Yuejian Fang, Daxin Jiang, and Nan Duan. Nuwa: Visual synthesis pretraining for neural visual world creation. In European conference on computer vision, pages 720–736. Springer, 2022.
  64. 64.Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu. Simmim: A simple framework for masked image modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9653–9663, 2022.
  65. 65.Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas. Videogpt: Video generation using vq-vae and transformers. arXiv preprint arXiv:2104.10157, 2021.
  66. 66.Yizhe Zhu, Martin Renqiang Min, Asim Kadav, and Hans Peter Graf. S3vae: Self-supervised sequential vae for representation disentanglement and data generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6538–6547, 2020.

Citation

MLA
Li, X., et al. “Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video”. 2025 IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 12904–14, https://doi.org/10.1109/iccv51701.2025.01199.
APA
Li, X., Chen, Q., Peng, X., Yu, K., Chen, X., & Lu, Y. (2025). Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video. 2025 IEEE/CVF International Conference on Computer Vision (ICCV), 12904–12914. https://doi.org/10.1109/iccv51701.2025.01199
Chicago
Li, X., Q. Chen, X. Peng, K. Yu, X. Chen, and Y. Lu. 2025. “Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video”. 2025 IEEE/CVF International Conference on Computer Vision (ICCV), 12904–14. https://doi.org/10.1109/iccv51701.2025.01199.
Harvard
Li, X. et al. (2025) “Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video”, 2025 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, pp. 12904–12914. Available at: https://doi.org/10.1109/iccv51701.2025.01199.
Vancouver
1. Li X, Chen Q, Peng X, Yu K, Chen X, Lu Y (2025) Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video. In: 2025 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, pp 12904–12914

BibTeX

@inproceedings{Li_2025, title={Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video}, url={http://dx.doi.org/10.1109/iccv51701.2025.01199}, DOI={10.1109/iccv51701.2025.01199}, booktitle={2025 IEEE/CVF International Conference on Computer Vision (ICCV)}, publisher={IEEE}, author={Li, Xiao and Chen, Qi and Peng, Xiulian and Yu, Kai and Chen, Xie and Lu, Yan}, year={2025}, month=Oct, pages={12904–12914} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/