A2RD: Agentic Autoregressive Diffusion for Long Video Consistency
Xuan Long DoYale SongMin-Yen KanTomas PfisterLong Le
Develops an agentic autoregressive diffusion framework that prevents semantic drift and narrative collapse in multi-minute video generation by coupling multimodal memory with closed-loop test-time refinement.
Generating consistent and narrative-driven long-form video has emerged as a major goal for digital storytelling, marketing, and education. However, current video diffusion models struggle when extending beyond short clips. Existing passive methods suffer from error accumulation, semantic drift, and narrative collapse, causing characters, objects, and backgrounds to abruptly alter appearance or repeat actions over time. The article introduces and evaluates an agentic framework designed to maintain temporal consistency and story coherence across extended multi-minute video generation without requiring model retraining.
The proposed framework, called Agentic Auto-Regressive Diffusion (A2RD), decouples creative synthesis from consistency enforcement through a closed-loop Retrieve-Synthesize-Refine-Update cycle. The system utilizes three primary mechanisms: a multimodal video memory that tracks visual arcs, spatial layouts, and camera paths across modalities; an adaptive generation engine that dynamically chooses between extrapolating forward from a single frame or interpolating between keyframes; and a hierarchical test-time self-improvement routine that inspects and refines frames and video segments using rubric-based feedback. To stress-test these capabilities under realistic narrative conditions, the authors developed LVbench-C, a benchmark comprising 120 multi-scene scenarios spanning three- to ten-minute durations featuring non-linear and recurring entity appearances.
The evaluation demonstrates that A2RD sets a new performance benchmark for long video generation. Across automated evaluations on standard benchmarks and LVbench-C, the architecture improved visual consistency by up to 30% and narrative coherence by approximately 20% compared to state-of-the-art baselines. In human evaluation studies on a 1-to-5 scale, A2RD achieved an average score of 4.68, significantly outperforming top baselines (3.93) in character identity preservation, transition smoothness, and prompt fidelity. Furthermore, testing across multiple open-source diffusion models confirmed that the framework generalizes effectively across different underlying generative engines.
These findings suggest that active, agentic self-refinement and structured multimodal memory provide a practical, training-free path toward scalable long-form video generation. Although the self-refinement cycle introduces additional inference latency—requiring several additional model queries and minutes of processing per segment—this automated overhead is substantially faster and cheaper than manual human inspection and prompt re-engineering. It also acts as an early gating mechanism that reduces wasted compute on defective end-to-end video renders.
Organizations developing automated video generation pipelines should consider adopting closed-loop agentic architectures and multimodal state tracking rather than relying solely on open-loop prompt conditioning. Future development should focus on strengthening physical layout reasoning in underlying foundation models, expanding automated evaluation judges to catch subtle visual errors, and adapting quality verification rubrics for specialized creative domains.
- Paper: StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text, Roberto Henschel et al. (2025). Its segment-by-segment long-video generation and short- and long-term visual memory provide direct context for A²RD’s autoregressive consistency cycle.
- Paper: History-Guided Video Diffusion, Kiwhan Song et al. (2025). Its flexible conditioning on video history clarifies the long-range context problem that A²RD addresses with multimodal memory and refinement.
- Paper: FIFO-Diffusion: Generating Infinite Videos from Text without Training, Jihwan Kim et al. (2024). Its queue-based strategy for extending diffusion generation shows how inference-time autoregression can sustain long videos while controlling error and memory.
- Paper: Phenaki: Variable Length Video Generation From Open Domain Textual Description, Ruben Villegas et al. (2023). Its variable-length, autoregressive text-to-video generation offers foundational context for A²RD’s long-horizon segment generation.
No sufficiently relevant recommendations were found.
