keyword
video stylization
Video stylization is a computer vision and computer graphics process that alters the visual appearance of a video sequence to match a target artistic style, texture, or aesthetic while preserving the underlying motion, structure, and semantic content of the original footage. Unlike static image stylization, video stylization requires maintaining temporal consistency across successive frames to prevent visual artifacts such as flickering, jitter, and unnatural texture drift. Modern techniques achieve this using generative models, neural networks, or optimization algorithms guided by reference images, text prompts, or learned style distributions, often incorporating motion estimation and temporal loss functions to ensure smooth, coherent transitions throughout the sequence.
2 items

VideoPoet: A Large Language Model for Zero-Shot Video Generation
Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng, Joshua V. Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso Martinez, David Minnen, Mikhail Sirotenko, Kihyuk Sohn, Xuan Yang, Hartwig Adam, Ming-Hsuan Yang, Irfan Essa, Huisheng Wang, David A. Ross, Bryan Seybold, Lu Jiang
Why you should read this
Demonstrates that a unified decoder-only language model trained on discrete multimodal tokens can match or outperform diffusion approaches across diverse video generation and editing tasks without task-specific architectural changes.
We present VideoPoet, a model for synthesizing high-quality videos from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs – including images, videos, text, and audio. The training protocol follows that of Large Language Models (LLMs), consisting of two stages: pretraining and task-specific adaptation. During pretraining, VideoPoet incorporates a mixture of multimodal generative objectives within an autoregressive Transformer framework. The pretrained LLM serves as a foundation that is adapted to a range of video generation tasks. We present results demonstrating the model’s state-of-the-art capabilities in zero-shot video generation, specifically highlighting the generation of high-fidelity motions. Project page: https://sites.research.google/videopoet/.
Added
2026-09-28

Precomputed Real-Time Texture Synthesis with Markovian Generative Adversarial Networks
Chuan Li, Michael Wand
Why you should read this
Introduces Markovian Generative Adversarial Networks, a feed-forward approach that eliminates test-time optimization to synthesize textures and stylize videos in real time at speeds hundreds of times faster than previous neural methods.
This paper proposes Markovian Generative Adversarial Networks (MGANs), a method for training generative neural networks for efficient texture synthesis. While deep neural network approaches have recently demonstrated remarkable results in terms of synthesis quality, they still come at considerable computational costs (minutes of run-time for low-res images). Our paper addresses this efficiency issue. Instead of a numerical deconvolution in previous work, we precompute a feed-forward, strided convolutional network that captures the feature statistics of Markovian patches and is able to directly generate outputs of arbitrary dimensions. Such network can directly decode brown noise to realistic texture, or photos to artistic paintings. With adversarial training, we obtain quality comparable to recent neural texture synthesis methods. As no optimization is required any longer at generation time, our run-time performance (0.25M pixel images at 25Hz) surpasses previous neural texture synthesizers by a significant margin (at least 500 times faster). We apply this idea to texture synthesis, style transfer, and video stylization.
Added
2026-09-25
