Multi-View Consistent Generative Adversarial Networks for 3D-aware Image Synthesis

Xuanmeng ZhangZhedong ZhengDaiheng GaoBang ZhangPan PanYi Yang

article2022CVPR62 citations

Proposes a 3D-aware generative adversarial network that enforces explicit multi-view geometric and photometric consistency constraints during training to prevent visual artifacts and synthesize high-resolution, view-consistent images across wide pose variations.

Listen

Generating realistic 3D images with precise camera viewpoint control is essential for modern computer vision and graphics applications. Traditional approaches typically model 2D image collections without internal 3D structural awareness, while existing 3D-aware generative models often optimize camera views independently. This lack of explicit multi-view geometric constraints frequently leads to collapsed shapes or noticeable visual inconsistencies, such as shifting facial features, when viewing angles change significantly.

The article introduces and evaluates Multi-View Consistent Generative Adversarial Networks (MVCGAN), a framework designed to synthesize high-resolution, photorealistic images from unposed 2D datasets while strictly preserving 3D structural consistency across various viewpoints.

The researchers established geometric correspondence between different camera views using projective warping based on rendered depth maps. They optimized pairs of views jointly during training by enforcing photometric consistency and applying a stereo mixing technique to maintain realistic image properties. To overcome the high computational cost of full-resolution 3D rendering, the authors developed a two-stage hybrid architecture: the model first establishes underlying 3D geometry at a lower resolution and subsequently refines fine 2D visual details at higher resolutions up to 512x512 pixels. The approach was evaluated on three standard benchmark datasets: CELEBA-HQ, FFHQ, and AFHQv2.

The experimental findings show substantial improvements over previous state-of-the-art methods across all tested benchmarks. For instance, at 512x512 resolution on the FFHQ dataset, the proposed method reduced the standard image quality error metric (Frechet Inception Distance) to 13.4, compared to scores ranging between 37.7 and 71.2 for competing models, representing an improvement of roughly 64% to 81%. On CELEBA-HQ at the same resolution, the error decreased from over 36.0 in existing methods to 12.9. Qualitative evaluations confirmed that synthesized images maintained coherent identity and structural integrity under wide camera pose variations, successfully separating 3D shape control from 2D texture details.

These results demonstrate that enforcing explicit multi-view geometric constraints resolves the trade-off between 3D viewpoint consistency and computational rendering efficiency. For technical stakeholders, this offers a practical method to generate controllable, high-fidelity visual assets from standard 2D image libraries without requiring expensive 3D scans or multi-camera setups.

Organizations developing 3D generative pipelines should consider incorporating multi-view warping constraints and hybrid rendering strategies into their model architectures. However, the current model is specifically designed for single-object subjects with uncluttered backgrounds and struggles with complex multi-object scenes. Future work should focus on developing compositional radiance fields to segment complex backgrounds and foreground objects before applying this framework in unconstrained real-world environments.

arXiv: 2204.06307
Cover for Multi-View Consistent Generative Adversarial Networks for 3D-aware Image Synthesis

Abstract

3D-aware image synthesis aims to generate images of objects from multiple views by learning a 3D representation. However, one key challenge remains: existing approaches lack geometry constraints, hence usually fail to generate multi-view consistent images. To address this challenge, we propose Multi-View Consistent Generative Adversarial Networks (MVCGAN) for high-quality 3D-aware image synthesis with geometry constraints. By leveraging the underlying 3D geometry information of generated images, i.e., depth and camera transformation matrix, we explicitly establish stereo correspondence between views to perform multi-view joint optimization. In particular, we enforce the photometric consistency between pairs of views and integrate a stereo mixup mechanism into the training process, encouraging the model to reason about the correct 3D shape. Besides, we design a two-stage training strategy with feature-level multi-view joint optimization to improve the image quality. Extensive experiments on three datasets demonstrate that MVCGAN achieves the state-of-the-art performance for 3D-aware image synthesis.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Preliminaries
  • 3.2. Multi-view Joint Optimization
  • 3.2.1 Image-level Multi-view Joint Optimization
  • 3.2.2 Feature-level Multi-view Joint Optimization
  • 3.3. Training Details
  • 4. Experiments
  • 4.1. Experimental Settings
  • 4.2. Comparison with SOTA
  • 4.3. Ablation Studies
  • 5. Conclusion and Discussion
  • References

Knowls

  1. Knowl 1 — MVCGAN geometry-constrained generator

    model/method

    MVCGAN is a 3D-aware generative model designed to synthesize photorealistic images under explicit camera-pose control while maintaining consistency across views. A latent code z∈R256z\in\mathbb{R}^{256} is mapped by an MLP GmG_m to an intermediate latent code w∈Ww\in\mathcal{W}. The intermediate code conditions both a SIREN-based generative radiance-field network GsG_s and a progressive 2D decoder GdG_d.

    During training, MVCGAN samples a primary camera pose ξpri\xi_{\mathrm{pri}} and an auxiliary camera pose ξaux\xi_{\mathrm{aux}} from a pose distribution pξp_\xi, and uses the same latent code to generate both views. Depth and the relative camera transformation establish pixel correspondences between the views, allowing the two generated views to be jointly optimized. During inference, the auxiliary pose is not needed: MVCGAN generates an image from only zz and the requested primary pose.

  2. Knowl 2 — Geometry constraints resolve shape-radiance ambiguity

    theoretical result

    The paper identifies shape-radiance ambiguity as the source of multi-view inconsistency in generative radiance fields trained only with single-view image losses. Without cross-view geometry constraints, a radiance field can reproduce each training image using an inaccurate 3D shape together with view-dependent radiance choices. Consequently, a generated object may look plausible from some viewpoints but exhibit inconsistent details or collapse under large pose changes.

    MVCGAN addresses this degeneracy by jointly optimizing two views of the same generated object. The paper diagnoses the ambiguity by warping one rendered view into another using the predicted depth and relative camera transformation; incorrect geometry produces a warped image with incorrect appearance. The resulting cross-view photometric constraint encourages the learned radiance field to represent a geometrically coherent 3D shape.

  3. Knowl 3 — Depth-based projective warping between generated views

    equation

    For a primary-view pixel vpriv_{\mathrm{pri}}, MVCGAN computes its rendered depth from the volume densities along the corresponding camera ray. Let NN be the number of sampled points, did_i the depth of sample ii, δi=di+1−di\delta_i=d_{i+1}-d_i the distance between adjacent samples, and σi≥0\sigma_i\ge 0 the volume density at sample ii. The accumulated transmittance is Ti=exp⁡(−∑j=1iσjδj)T_i=\exp\left(-\sum_{j=1}^{i}\sigma_j\delta_j\right), and the depth is

    D(vpri)=∑i=1NTi(1−exp⁡(−σiδi))di.\mathcal{D}(v_{\mathrm{pri}})=\sum_{i=1}^{N}T_i\left(1-\exp(-\sigma_i\delta_i)\right)d_i.

    Let hprih_{\mathrm{pri}} be the homogeneous image coordinate of the pixel in the primary camera, KK be the known camera-intrinsic matrix, and [R,t][R,t] be the rigid transformation from the primary camera to the auxiliary camera, obtained from ξpri\xi_{\mathrm{pri}} and ξaux\xi_{\mathrm{aux}}. The corresponding auxiliary-view coordinate is computed by

    haux=K[R,t]D(vpri)K−1hpri.h_{\mathrm{aux}}=K[R,t]\mathcal{D}(v_{\mathrm{pri}})K^{-1}h_{\mathrm{pri}}.

    MVCGAN samples the auxiliary image or feature map at the location represented by hauxh_{\mathrm{aux}} and uses that value to reconstruct the corresponding pixel in the warped primary-view representation. This produces Iwarp\mathcal{I}_{\mathrm{warp}} from the auxiliary image Iaux\mathcal{I}_{\mathrm{aux}} while using the primary-view depth and the relative camera geometry.

  4. Knowl 4 — Image-level reprojection and stereo mixup

    model/method

    MVCGAN's first-stage multi-view optimization enforces photometric agreement between the directly rendered primary image Ipri\mathcal{I}_{\mathrm{pri}} and the auxiliary image warped into the primary view, Iwarp\mathcal{I}_{\mathrm{warp}}. The image-level reprojection loss combines pixelwise L1L_1 error with structural similarity:

    Lir=(1−μ)∥Ipri−Iwarp∥1+μ2(1−SSIM⁡(Ipri,Iwarp)),L_{\mathrm{ir}}=(1-\mu)\left\|\mathcal{I}_{\mathrm{pri}}-\mathcal{I}_{\mathrm{warp}}\right\|_1+\frac{\mu}{2}\left(1-\operatorname{SSIM}(\mathcal{I}_{\mathrm{pri}},\mathcal{I}_{\mathrm{warp}})\right),

    where SSIM⁡\operatorname{SSIM} is the structural-similarity metric and MVCGAN uses μ=0.85\mu=0.85.

    To make the warped view resemble a realistic image without adding a separate discriminator for it, MVCGAN forms a stereo-mixed image

    Imix=ηIpri+(1−η)Iwarp,\mathcal{I}_{\mathrm{mix}}=\eta\mathcal{I}_{\mathrm{pri}}+(1-\eta)\mathcal{I}_{\mathrm{warp}},

    where η\eta is randomly sampled from [0,1][0,1] at every training iteration. The discriminator receives this virtual image, while the reprojection loss directly encourages geometric and photometric consistency between the two views.

  5. Knowl 5 — Feature-level multi-view optimization

    equation

    In MVCGAN's second training stage, volume rendering is applied to a learned feature vector rather than directly to high-resolution RGB values. Let fi∈R256f_i\in\mathbb{R}^{256} be the feature at sample ii before the final color-producing layer of the generative radiance-field network, and let TiT_i, σi\sigma_i, and δi\delta_i denote transmittance, density, and adjacent-sample distance as defined by the volume-rendering process. The primary feature map is

    Fpri=∑i=1NTi(1−exp⁡(−σiδi))fi.\mathcal{F}_{\mathrm{pri}}=\sum_{i=1}^{N}T_i\left(1-\exp(-\sigma_i\delta_i)\right)f_i.

    The auxiliary feature map Faux\mathcal{F}_{\mathrm{aux}} is rendered analogously, and projective warping produces Fwarp\mathcal{F}_{\mathrm{warp}} in the primary view. Feature-level geometric consistency is imposed with an implicit diversified Markov random field loss:

    Lfr=Lmrf(Fpri,Fwarp).L_{\mathrm{fr}}=L_{\mathrm{mrf}}(\mathcal{F}_{\mathrm{pri}},\mathcal{F}_{\mathrm{warp}}).

    The feature maps are also combined through stereo mixup, Fmix=ηFpri+(1−η)Fwarp\mathcal{F}_{\mathrm{mix}}=\eta\mathcal{F}_{\mathrm{pri}}+(1-\eta)\mathcal{F}_{\mathrm{warp}}, before being passed to the 2D decoder. This allows the radiance field to learn geometry at low resolution while preserving high-frequency geometric information in the feature representation.

  6. Knowl 6 — Two-stage hybrid MLP-CNN and progressive rendering strategy

    model/method

    MVCGAN separates 3D geometry from high-resolution 2D appearance with a hybrid MLP-CNN design. The MLP-based radiance-field network GsG_s predicts volumetric geometry and low-resolution features, while the CNN-based progressive decoder GdG_d synthesizes fine image details. This avoids the computational cost of directly rendering high-resolution RGB images through a NeRF-style network.

    Stage I trains at 64×6464\times64 resolution. The radiance field directly volume-renders primary and auxiliary RGB images, and image-level reprojection optimization is applied. Stage II renders primary and auxiliary feature maps at 64×6464\times64, performs feature-level warping, reprojection, and stereo mixup, and then uses GdG_d to progressively upsample the mixed feature map to 128×128128\times128, 256×256256\times256, and 512×512512\times512 RGB images.

    The decoder is conditioned by ww through adaptive instance normalization. Each progressive stage uses bilinear upsampling and a 1×11\times1 convolution to produce RGB outputs; skip connections fade newly added layers into the network to stabilize progressive training. The resulting architecture assigns shape and pose changes primarily to the radiance-field branch and fine appearance details to the 2D decoder.

  7. Knowl 7 — Geometry-constrained adversarial training objective

    equation

    MVCGAN trains a progressive convolutional discriminator DϕD_\phi jointly with the generator GθG_\theta. Let z∼Zz\sim\mathcal{Z} be a latent code, ξpri,ξaux∼pξ\xi_{\mathrm{pri}},\xi_{\mathrm{aux}}\sim p_\xi be sampled camera poses, and I∼pD\mathcal{I}\sim p_\mathcal{D} be a real training image. The paper uses the non-saturating GAN objective with an R1R_1 gradient penalty and a geometry-reprojection term:

    V(θ,ϕ)=Ez,ξpri,ξaux[f(Dϕ(Gθ(z,ξpri,ξaux)))]+EI∼pD[f(−Dϕ(I))−λ∥∇Dϕ(I)∥2]+Lre,\begin{aligned} \mathcal{V}(\theta,\phi)= {} & \mathbb{E}_{z,\xi_{\mathrm{pri}},\xi_{\mathrm{aux}}}\left[f\left(D_\phi\left(G_\theta(z,\xi_{\mathrm{pri}},\xi_{\mathrm{aux}})\right)\right)\right] \\ &+\mathbb{E}_{\mathcal{I}\sim p_\mathcal{D}}\left[f\left(-D_\phi(\mathcal{I})\right)-\lambda\left\|\nabla D_\phi(\mathcal{I})\right\|^2\right]+L_{\mathrm{re}}, \end{aligned}

    where f(t)=−log⁡(1+exp⁡(−t))f(t)=-\log(1+\exp(-t)), ∇Dϕ(I)\nabla D_\phi(\mathcal{I}) is the discriminator gradient with respect to the real image, and MVCGAN uses λ=10\lambda=10. The geometry term is Lre=LirL_{\mathrm{re}}=L_{\mathrm{ir}} during Stage I and Lre=LfrL_{\mathrm{re}}=L_{\mathrm{fr}} during Stage II.

  8. Knowl 8 — Quantitative image-quality results across three datasets

    data/table

    MVCGAN was evaluated on CELEBA-HQ, FFHQ, and AFHQv2 using Fréchet Inception Distance (FID), computed between 20,000 generated images and real images. AFHQv2 experiments used cat faces to match the comparison setting. Lower FID indicates better image quality. MVCGAN achieved the lowest FID at every tested dataset and resolution, substantially outperforming GRAF, pi-GAN, and GIRAFFE.

    Method CELEBA-HQ FFHQ AFHQv2
    2562256^2 5122512^2 2562256^2 5122512^2 2562256^2
    GRAF 47.5 57.7 67.2 71.2 75.8
    pi-GAN 39.7 41.8 38.1 39.9 42.0
    GIRAFFE 36.0 36.2 34.6 37.7 29.2
    MVCGAN 11.8 12.9 13.7 13.4 17.1
  9. Knowl 9 — Ablation confirms complementary roles of image and feature optimization

    empirical result

    An ablation on FFHQ at 256×256256\times256 resolution shows that image-level multi-view optimization alone produces pose-consistent images but gives an FID of 22.522.5. Adding feature-level multi-view optimization reduces the FID to 13.713.7 and improves fine 2D details.

    The result supports the division of labor in MVCGAN: image-level reprojection provides geometric guidance for learning a coherent 3D shape, while feature-level reprojection and progressive decoding improve high-resolution visual quality without requiring direct high-resolution NeRF rendering.

  10. Knowl 10 — Limitation to simple single-object scenes

    limitation

    MVCGAN is primarily designed for single-object scenes with relatively simple backgrounds. The method does not perform well when scenes contain multiple objects or complex background-attached structure; the paper reports unsuccessful high-quality synthesis on the University-1652 dataset as an example.

    The authors suggest extending MVCGAN with a compositional radiance field that separately models foreground objects and background. Such an extension would need to establish geometry relationships between foreground and background using depth and occlusion maps, and the authors also identify additional image annotations as a possible way to address more complex real-world scenes.

Coverage note — The style-mixing and latent/camera interpolation analyses were omitted because they are secondary qualitative diagnostics rather than load-bearing components of the method or primary performance results.

References

  1. 1.Hassan Abu Alhaija, Siva Karthik Mustikovela, Andreas Geiger, and Carsten Rother. Geometric image synthesis. In ACCV, 2018. 1, 3
  2. 2.Alex M Andrew. Multiple view geometry in computer vision. Kybernetes, 2001. 2, 4
  3. 3.Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale gan training for high fidelity natural image synthesis. In ICLR, 2018. 1
  4. 4.Eric R Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu, and Gordon Wetzstein. pi-gan: Periodic implicit generative adversarial networks for 3d-aware image synthesis. In CVPR, 2021. 1, 2, 3, 4, 5, 6, 7
  5. 5.Anpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang, Fanbo Xiang, Jingyi Yu, and Hao Su. Mvsnerf: Fast generalizable radiance field reconstruction from multi-view stereo. arXiv preprint arXiv:2103.15595, 2021. 3
  6. 6.Shenchang Eric Chen and Lance Williams. View interpolation for image synthesis. In Conference on Computer graphics and interactive techniques, 1993. 2, 4
  7. 7.Julian Chibane, Aayush Bansal, Verica Lazova, and Gerard Pons-Moll. Stereo radiance fields (srf): Learning view synthesis for sparse views of novel scenes. In CVPR, 2021. 3
  8. 8.Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo. Stargan: Unified generative adversarial networks for multi-domain image-to-image translation. In CVPR, 2018. 1
  9. 9.Yunjey Choi, Youngjung Uh, Jaejun Yoo, and Jung-Woo Ha. Stargan v2: Diverse image synthesis for multiple domains. In CVPR, 2020. 1, 2, 6, 7
  10. 10.Robert T Collins. A space-sweep approach to true multi-image matching. In CVPR, 1996. 2
  11. 11.Paul E Debevec, Camillo J Taylor, and Jitendra Malik. Modeling and rendering architecture from photographs: A hybrid geometry-and image-based approach. In Conference on Computer graphics and interactive techniques, 1996. 2, 4
  12. 12.Kangle Deng, Andrew Liua, Jun-Yan Zhu, and Deva Ramanan. Depth-supervised nerf: Fewer views and faster training for free. arXiv preprint arXiv:2107.02791, 2021. 5
  13. 13.Terrance DeVries, Miguel Angel Bautista, Nitish Srivastava, Graham W. Taylor, and Joshua M. Susskind. Unconstrained scene generation with locally conditioned radiance fields. In ICCV, 2021. 1, 3
  14. 14.Vincent Dumoulin, Jonathon Shlens, and Manjunath Kudlur. In A learned representation for artistic style, 2020. 6
  15. 15.Yao Feng, Haiwen Feng, Michael J Black, and Timo Bolkart. Learning an animatable detailed 3d face model from in-the-wild images. ACM Transactions on Graphics (TOG), 40(4):1–13, 2021. 6
  16. 16.Chen Gao, Ayush Saraf, Johannes Kopf, and Jia-Bin Huang. Dynamic view synthesis from dynamic monocular video. arXiv preprint arXiv:2105.06468, 2021. 3
  17. 17.Stephan J Garbin, Marek Kowalski, Matthew Johnson, Jamie Shotton, and Julien Valentin. Fastnerf: High-fidelity neural rendering at 200fps. arXiv preprint arXiv:2103.10380, 2021. 3
  18. 18.Clement Godard, Oisin Mac Aodha, Michael Firman, and Gabriel J Brostow. Digging into self-supervised monocular depth estimation. In ICCV, 2019. 2, 4, 5
  19. 19.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In NeurIPS, 2014. 1
  20. 20.Paul Henderson, Vagia Tsiminaki, and Christoph H Lampert. Leveraging 2d data to learn textured 3d mesh generation. In CVPR, 2020. 1, 3
  21. 21.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In NeurIPS, 2017. 6
  22. 22.Xun Huang and Serge Belongie. Arbitrary style transfer in real-time with adaptive instance normalization. In ICCV, 2017. 1, 6
  23. 23.Yoonwoo Jeong, Seokjun Ahn, Christopehr Choy, Animashree Anandkumar, Minsu Cho, and Jaesik Park. Self-calibrating neural radiance fields. In ICCV, 2021. 3
  24. 24.Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of gans for improved quality, stability, and variation. In ICLR, 2018. 1, 2, 6
  25. 25.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In CVPR, 2019. 1, 2, 3, 6, 7, 8
  26. 26.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In CVPR, 2020. 1, 6
  27. 27.Tianye Li, Mira Slavcheva, Michael Zollhoefer, Simon Green, Christoph Lassner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, and Zhaoyang Lv. Neural 3d video synthesis, 2021. 3
  28. 28.Zhengqi Li, Simon Niklaus, Noah Snavely, and Oliver Wang. Neural scene flow fields for space-time view synthesis of dynamic scenes. https://arxiv.org/abs/2011.13084, 2020. 3
  29. 29.Yiyi Liao, Katja Schwarz, Lars Mescheder, and Andreas Geiger. Towards unsupervised learning of generative models for 3d controllable image synthesis. In CVPR, 2020. 1, 3
  30. 30.Chen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, and Simon Lucey. Barf: Bundle-adjusting neural radiance fields. In ICCV, 2021. 3
  31. 31.David B Lindell, Julien NP Martel, and Gordon Wetzstein. Autoint: Automatic integration for fast neural volume rendering. In CVPR, 2021. 3
  32. 32.Xiaoyang Lyu, Liang Liu, Mengmeng Wang, Xin Kong, Lina Liu, Yong Liu, Xinxin Chen, and Yi Yuan. Hr-depth: High resolution self-supervised monocular depth estimation. In AAAI, 2021. 5
  33. 33.Quan Meng, Anpei Chen, Haimin Luo, Minye Wu, Hao Su, Lan Xu, Xuming He, and Jingyi Yu. Gnerf: Gan-based neural radiance field without posed camera. arXiv preprint arXiv:2103.15606, 2021. 3
  34. 34.Lars Mescheder, Andreas Geiger, and Sebastian Nowozin. Which training methods for gans do actually converge? In ICML, 2018. 6
  35. 35.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020. 2, 3, 4, 5
  36. 36.Raul Mur-Artal, Jose Maria Martinez Montiel, and Juan D Tardos. Orb-slam: a versatile and accurate monocular slam system. IEEE Transactions on Robotics, 31(5):1147–1163, 2015. 2
  37. 37.Thu Nguyen-Phuoc, Chuan Li, Lucas Theis, Christian Richardt, and Yong-Liang Yang. Hologan: Unsupervised learning of 3d representations from natural images. In CVPR, 2019. 1, 3
  38. 38.Thu Nguyen-Phuoc, Christian Richardt, Long Mai, Yong-Liang Yang, and Niloy J Mitra. Blockgan: Learning 3d object-aware scene representations from unlabelled images. In NeurIPS, 2020. 1, 3
  39. 39.Michael Niemeyer and Andreas Geiger. Giraffe: Representing scenes as compositional generative neural feature fields. In CVPR, 2021. 1, 2, 3, 5, 6, 7
  40. 40.Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. In CVPR, 2021. 3
  41. 41.Sudeep Pillai, Rareș Ambruș, and Adrien Gaidon. Superdepth: Self-supervised, super-resolved monocular depth estimation. In ICRA, 2019. 5
  42. 42.Daniel Rebain, Wei Jiang, Soroosh Yazdani, Ke Li, Kwang Moo Yi, and Andrea Tagliasacchi. Derf: Decomposed radiance fields. In CVPR, 2021. 3
  43. 43.Christian Reiser, Songyou Peng, Yiyi Liao, and Andreas Geiger. Kilonerf: Speeding up neural radiance fields with thousands of tiny mlps. arXiv preprint arXiv:2103.13744, 2021. 3
  44. 44.Johannes L Schonberger and Jan-Michael Frahm. Structure-from-motion revisited. In CVPR, 2016. 2
  45. 45.Katja Schwarz, Yiyi Liao, Michael Niemeyer, and Andreas Geiger. Graf: Generative radiance fields for 3d-aware image synthesis. In NeurIPS, 2020. 1, 2, 3, 5, 6, 7
  46. 46.Steven M Seitz and Charles R Dyer. View morphing. In Conference on Computer graphics and interactive techniques, 1996. 2, 4
  47. 47.Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. Implicit neural representations with periodic activation functions. In NeurIPS, 2020. 4
  48. 48.Richard Szeliski and Polina Golland. Stereo matching with transparency and matting. In ICCV, 1998. 2
  49. 49.Alex Trevithick and Bo Yang. Grf: Learning a general radiance field for 3d scene representation and rendering. In ICCV, 2021. 3
  50. 50.Yi Wang, Xin Tao, Xiaojuan Qi, Xiaoyong Shen, and Jiaya Jia. Image inpainting via generative multi-column convolutional neural networks. In NeurIPS, 2018. 6
  51. 51.Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4):600–612, 2004. 5
  52. 52.Zirui Wang, Shangzhe Wu, Weidi Xie, Min Chen, and Victor Adrian Prisacariu. NeRF–: Neural radiance fields without known camera parameters. https://arxiv.org/abs/2102.07064, 2021. 3
  53. 53.Yi Wei, Shaohui Liu, Yongming Rao, Wang Zhao, Jiwen Lu, and Jie Zhou. Nerfingmvs: Guided optimization of neural radiance fields for indoor multi-view stereo. In ICCV, 2021. 3
  54. 54.Wenqi Xian, Jia-Bin Huang, Johannes Kopf, and Changil Kim. Space-time neural irradiance fields for free-viewpoint video. https://arxiv.org/abs/2011.12950, 2020. 3
  55. 55.Bangbang Yang, Yinda Zhang, Yinghao Xu, Yijin Li, Han Zhou, Hujun Bao, Guofeng Zhang, and Zhaopeng Cui. Learning object-compositional neural radiance field for editable scene rendering. In ICCV, 2021. 8
  56. 56.Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan. Mvsnet: Depth inference for unstructured multi-view stereo. In ECCV, 2018. 2
  57. 57.Lin Yen-Chen, Pete Florence, Jonathan T. Barron, Alberto Rodriguez, Phillip Isola, and Tsung-Yi Lin. iNeRF: Inverting neural radiance fields for pose estimation. https://arxiv.org/abs/2012.05877, 2020. 3
  58. 58.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In CVPR, 2021. 3
  59. 59.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In ICLR, 2018. 5
  60. 60.Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. Nerf++: Analyzing and improving neural radiance fields. arXiv preprint arXiv:2010.07492, 2020. 4
  61. 61.Hang Zhao, Orazio Gallo, Iuri Frosio, and Jan Kautz. Loss functions for image restoration with neural networks. IEEE Transactions on Computational Imaging, 3(1):47–57, 2016. 5
  62. 62.Zhedong Zheng, Yunchao Wei, and Yi Yang. University-1652: A multi-view multi-source benchmark for drone-based geo-localization. In ACM MM, 2020. 8
  63. 63.Zhedong Zheng, Xiaodong Yang, Zhiding Yu, Liang Zheng, Yi Yang, and Jan Kautz. Joint discriminative and generative learning for person re-identification. In CVPR, 2019. 1
  64. 64.Tinghui Zhou, Matthew Brown, Noah Snavely, and David G Lowe. Unsupervised learning of depth and ego-motion from video. In CVPR, 2017. 2, 4, 5
  65. 65.Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In ICCV, 2017. 1
  66. 66.Jun-Yan Zhu, Zhoutong Zhang, Chengkai Zhang, Jiajun Wu, Antonio Torralba, Josh Tenenbaum, and Bill Freeman. Visual object networks: Image generation with disentangled 3d representations. 2018. 1, 3

Citation

MLA
Zhang, X., et al. “Multi-View Consistent Generative Adversarial Networks for 3D-aware Image Synthesis”. arXiv, 2022, http://arxiv.org/abs/2204.06307v1.
APA
Zhang, X., Zheng, Z., Gao, D., Zhang, B., Pan, P., & Yang, Y. (2022). Multi-View Consistent Generative Adversarial Networks for 3D-aware Image Synthesis. arXiv. http://arxiv.org/abs/2204.06307v1
Chicago
Zhang, X., Z. Zheng, D. Gao, B. Zhang, P. Pan, and Y. Yang. 2022. “Multi-View Consistent Generative Adversarial Networks for 3D-aware Image Synthesis”. arXiv. http://arxiv.org/abs/2204.06307v1.
Harvard
Zhang, X. et al. (2022) “Multi-View Consistent Generative Adversarial Networks for 3D-aware Image Synthesis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2204.06307v1.
Vancouver
1. Zhang X, Zheng Z, Gao D, Zhang B, Pan P, Yang Y (2022) Multi-View Consistent Generative Adversarial Networks for 3D-aware Image Synthesis. arXiv

BibTeX

@article{zhang2022multi,
  title = {Multi-View Consistent Generative Adversarial Networks for 3D-aware Image Synthesis},
  author = {Zhang, Xuanmeng and Zheng, Zhedong and Gao, Daiheng and Zhang, Bang and Pan, Pan and Yang, Yi},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2204.06307v1},
  eprint = {2204.06307}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE