EscherNet: A Generative Model for Scalable View Synthesis
Xin KongShikun LiuXiaoyang LyuMarwan TaherXiaojuan QiAndrew J. Davison
Introduces a multi-view conditioned diffusion model with relative camera positional encoding that synthesizes over 100 consistent target views simultaneously on a single GPU from arbitrary reference angles, unifying novel view synthesis and 3D reconstruction without expensive 3D convolutions or volumetric rendering.
Synthesizing new camera viewpoints and generating 3D models from limited 2D imagery are critical capabilities for spatial computing, robotics, and digital content creation. Existing neural rendering methods typically require expensive, scene-specific optimization tied to explicit coordinate systems, which prevents them from generalizing across different scenes and causes them to fail when only a few reference images are available. Meanwhile, recent 3D generative diffusion models remain rigid, typically supporting only a single reference viewpoint or producing a fixed number of target camera angles.
The article introduces and evaluates EscherNet, a conditional diffusion framework designed to enable scalable, coordinate-free view synthesis and 3D reconstruction. The primary objective is to demonstrate a generative model capable of producing an arbitrary number of consistent novel target views from an arbitrary number of reference views with continuous relative camera control.
To achieve this, the authors adapt a standard 2D latent diffusion model by incorporating a lightweight visual encoder to capture image details and introducing a specialized camera positional encoding mechanism. This encoding allows internal transformer attention layers to compute cross-view relationships purely based on relative camera transformations rather than fixed global coordinates. The model was trained on roughly 800,000 synthetic 3D objects using sets of three reference and three target views. The authors evaluated the system across standard benchmarks for novel view synthesis, 3D object reconstruction, and two-stage text-to-3D generation, comparing performance against leading diffusion baselines and neural rendering techniques.
The experimental findings demonstrate significant technical and qualitative improvements over existing methods. First, the framework substantially outperformed competing 3D diffusion baselines in novel view synthesis quality, surpassing models trained on ten times more data. Second, when tested on 3D shape reconstruction, the model achieved approximately a 25% improvement in geometric accuracy (measured by Chamfer distance) over the best baseline when starting from a single image, widening to a 60% improvement when conditioned on ten images. Third, in few-view scenarios with fewer than five input images, the zero-shot framework synthesized plausible images where state-of-the-art scene-specific optimization methods failed to construct meaningful geometry. Finally, rendering quality scaled consistently as more reference images were provided, enabling the simultaneous generation of over 100 mutually consistent target views on a single consumer-grade graphics processing unit.
These results demonstrate that 3D synthesis can be effectively achieved through 2D generative priors without requiring explicit 3D volumetric operations or ground-truth 3D meshes. By unifying novel view synthesis, few-shot 3D reconstruction, and multi-view generation into one system, the approach lowers the computational and data barriers needed to generate high-quality 3D assets. However, while the model excels in sparse-data settings, scene-specific optimization methods still provide superior photorealism when rich data (more than ten reference images) is available.
Organizations developing spatial computing and 3D asset generation workflows should consider adopting relative-pose generative models to automate few-image reconstruction pipelines. Future technical exploration should prioritize bridging the rendering fidelity gap in dense-image scenarios and extending model training from bounded object-centric settings to unconstrained, complex real-world environments with full six-degree-of-freedom camera movement.
Readers should interpret these findings within the context of current system boundaries. The primary evaluations rely on object-centric synthetic datasets, and generating views autoregressively across very long sequences leads to accumulated quality degradation. Overall, there is high confidence that relative camera conditioning provides an efficient and scalable foundation for generalized 3D view synthesis.
- Paper: Zero-1-to-3: Zero-shot One Image to 3D Object, Ruoshi Liu et al. (2023). Zero-1-to-3 establishes the fundamental paradigm of viewpoint-conditioned 2D diffusion models for novel view synthesis that EscherNet generalizes to multi-view conditions.
- Paper: pixelNeRF: Neural Radiance Fields from One or Few Images, Alex Yu et al. (2021). pixelNeRF introduces the conditioning of neural radiance and view synthesis architectures directly on sparse input views and relative camera poses.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). NeRF provides the essential foundation of continuous coordinate representations and positional encodings for novel view synthesis from posed images.
- Paper: Scene Representation Networks: Continuous 3D-Structure-Aware Neural Scene Representations, V. Sitzmann et al. (2019). Scene Representation Networks formulate continuous, structure-aware implicit scene representations conditioned on camera poses without requiring explicit 3D geometry.
- Paper: Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single Image, Xuanchi Ren et al. (2022). This work introduces relative camera transformation conditioning in attention mechanisms for consistent novel view synthesis.
- Paper: 3D-R2N2: A Unified Approach for Single and Multi-view 3D Object Reconstruction, Christopher B. Choy et al. (2016). 3D-R2N2 provides the foundational architecture for unifying single- and multi-view 3D synthesis within a cohesive learned neural network.
- Paper: Free3D: Consistent Novel View Synthesis Without 3D Representation, Chuanxia Zheng et al. (2024). Free3D develops camera-ray conditioned normalization and multi-view attention mechanisms to synthesize consistent 360-degree views without explicit 3D representations.
- Paper: SyncDreamer: Generating Multiview-consistent Images from a Single-view Image, Yuan Liu et al. (2024). SyncDreamer explores synchronized multi-view generation with 3D-aware cross-attention to address consistency bottlenecks in novel view diffusion.
- Paper: Wonder3D: Single Image to 3D Using Cross-Domain Diffusion, Xiaoxiao Long et al. (2024). Wonder3D extends generative multi-view diffusion models by cross-domain coupling of normal maps and color views for downstream reconstruction.
- Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). VGGT scales transformer-based feed-forward multi-view geometry estimation across dense sets of images, building on scalable neural camera conditioning concepts.
- Paper: DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision, Lu Ling et al. (2024). DL3DV-10K provides a large-scale real-world multi-view dataset and benchmark to evaluate generalizable novel view synthesis frameworks like EscherNet.
