A2RD: Agentic Autoregressive Diffusion for Long Video Consistency

Xuan Long DoYale SongMin-Yen KanTomas PfisterLong Le

article2026arXiv7 citations

Develops an agentic autoregressive diffusion framework that prevents semantic drift and narrative collapse in multi-minute video generation by coupling multimodal memory with closed-loop test-time refinement.

Listen

Generating consistent and narrative-driven long-form video has emerged as a major goal for digital storytelling, marketing, and education. However, current video diffusion models struggle when extending beyond short clips. Existing passive methods suffer from error accumulation, semantic drift, and narrative collapse, causing characters, objects, and backgrounds to abruptly alter appearance or repeat actions over time. The article introduces and evaluates an agentic framework designed to maintain temporal consistency and story coherence across extended multi-minute video generation without requiring model retraining.

The proposed framework, called Agentic Auto-Regressive Diffusion (A2RD), decouples creative synthesis from consistency enforcement through a closed-loop Retrieve-Synthesize-Refine-Update cycle. The system utilizes three primary mechanisms: a multimodal video memory that tracks visual arcs, spatial layouts, and camera paths across modalities; an adaptive generation engine that dynamically chooses between extrapolating forward from a single frame or interpolating between keyframes; and a hierarchical test-time self-improvement routine that inspects and refines frames and video segments using rubric-based feedback. To stress-test these capabilities under realistic narrative conditions, the authors developed LVbench-C, a benchmark comprising 120 multi-scene scenarios spanning three- to ten-minute durations featuring non-linear and recurring entity appearances.

The evaluation demonstrates that A2RD sets a new performance benchmark for long video generation. Across automated evaluations on standard benchmarks and LVbench-C, the architecture improved visual consistency by up to 30% and narrative coherence by approximately 20% compared to state-of-the-art baselines. In human evaluation studies on a 1-to-5 scale, A2RD achieved an average score of 4.68, significantly outperforming top baselines (3.93) in character identity preservation, transition smoothness, and prompt fidelity. Furthermore, testing across multiple open-source diffusion models confirmed that the framework generalizes effectively across different underlying generative engines.

These findings suggest that active, agentic self-refinement and structured multimodal memory provide a practical, training-free path toward scalable long-form video generation. Although the self-refinement cycle introduces additional inference latency—requiring several additional model queries and minutes of processing per segment—this automated overhead is substantially faster and cheaper than manual human inspection and prompt re-engineering. It also acts as an early gating mechanism that reduces wasted compute on defective end-to-end video renders.

Organizations developing automated video generation pipelines should consider adopting closed-loop agentic architectures and multimodal state tracking rather than relying solely on open-loop prompt conditioning. Future development should focus on strengthening physical layout reasoning in underlying foundation models, expanding automated evaluation judges to catch subtle visual errors, and adapting quality verification rubrics for specialized creative domains.

No sufficiently relevant recommendations were found.

Cover for A2RD: Agentic Autoregressive Diffusion for Long Video Consistency

Abstract

Synthesizing consistent and coherent long video remains a fundamental challenge. Existing methods suffer from semantic drift and narrative collapse over long horizons. We present A2^2RD, an Agentic Auto-Regressive Diffusion architecture that decouples creative synthesis from consistency enforcement. A2^2RD formulates long video synthesis as a closed-loop process that synthesizes and self-improves video segment-by-segment through a Retrieve--Synthesize--Refine--Update cycle. It comprises three core components: (i) Multimodal Video Memory that tracks video progression across modalities; (ii) Adaptive Segment Generation that switches among generation modes for natural progression and visual consistency; and (iii) Hierarchical Test-Time Self-Improvement that self-improves each segment at frame and video levels to prevent error propagation. We further introduce LVBench-C, a challenging benchmark with non-linear entity and environment transitions to stress-test long-horizon consistency. Across public and LVBench-C benchmarks spanning one- to ten-minute videos, A2^2RD outperforms state-of-the-art baselines by up to 30% in consistency and 20% in narrative coherence. Human evaluations corroborate these gains while also highlighting notable improvements in motion and transition smoothness.

Citation

MLA
Long, D. X., et al. “A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistency”. arXiv, 2026, http://arxiv.org/abs/2605.06924v1.
APA
Long, D. X., Song, Y., Kan, M.-Y., Pfister, T., & Le, L. T. (2026). A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistency. arXiv. http://arxiv.org/abs/2605.06924v1
Chicago
Long, D. X., Y. Song, M.-Y. Kan, T. Pfister, and L. T. Le. 2026. “A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistency”. arXiv. http://arxiv.org/abs/2605.06924v1.
Harvard
Long, D.X. et al. (2026) “A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistency”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2605.06924v1.
Vancouver
1. Long DX, Song Y, Kan M-Y, Pfister T, Le LT (2026) A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistency. arXiv

BibTeX

@article{long2026agentic,
  title = {A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistency},
  author = {Long, Do Xuan and Song, Yale and Kan, Min-Yen and Pfister, Tomas and Le, Long T.},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2605.06924v1},
  eprint = {2605.06924}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/