Built independently by an author, for readers. Read the story and support ChapterPal

keyword

image extrapolation

Image extrapolation is the process of extending an image beyond its original boundaries by generating plausible visual content that continues the scene, often using visible context to keep the added areas coherent with the original.

3 items

Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single Image

Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single Image

Xuanchi Ren, Xiaolong Wang

OrganizationsThe Hong Kong University of Science and TechnologyUniversity of California, San Diego

Why you should read this

Develops an autoregressive Transformer framework that generates geometrically consistent long-term 3D scene videos from a single input image and a large camera trajectory by using a camera-aware bias to guide space-time attention.

Novel view synthesis from a single image has recently attracted a lot of attention, and it has been primarily advanced by 3D deep learning and rendering techniques. However, most work is still limited by synthesizing new views within relatively small camera motions. In this paper, we propose a novel approach to synthesize a consistent long-term video given a single scene image and a trajectory of large camera motions. Our approach utilizes an autoregressive Transformer to perform sequential modeling of multiple frames, which reasons the relations between multiple frames and the corresponding cameras to predict the next frame. To facilitate learning and ensure consistency among generated frames, we introduce a locality constraint based on the input cameras to guide self-attention among a large number of patches across space and time. Our method outperforms state-of-the-art view synthesis approaches by a large margin, especially when synthesizing long-term future in indoor 3D scenes. Project page at https://xrenaa.github.io/look-outside-room/.

Added

2026-09-26

Globally and locally consistent image completion

Globally and locally consistent image completion

SATOSHI IIZUKA, EDGAR SIMO-SERRA, HIROSHI ISHIKAWA

OrganizationsWaseda University

Why you should read this

Proposes a fully convolutional image completion framework trained with dual global and local context discriminators, enabling realistic synthesis of arbitrary-shaped missing regions while preserving both fine local details and overall semantic coherence across diverse scenes.

We present a novel approach for image completion that results in images that are both locally and globally consistent. With a fully-convolutional neural network, we can complete images of arbitrary resolutions by filling-in missing regions of any shape. To train this image completion network to be consistent, we use global and local context discriminators that are trained to distinguish real images from completed ones. The global discriminator looks at the entire image to assess if it is coherent as a whole, while the local discriminator looks only at a small area centered at the completed region to ensure the local consistency of the generated patches. The image completion network is then trained to fool the both context discriminator networks, which requires it to generate images that are indistinguishable from real ones with regard to overall consistency as well as in details. We show that our approach can be used to complete a wide variety of scenes. Furthermore, in contrast with the patch-based approaches such as PatchMatch, our approach can generate fragments that do not appear elsewhere in the image, which allows us to naturally complete the images of objects with familiar and highly specific structures, such as faces.

Added

2026-09-16