MET3R: Measuring Multi-View Consistency in Generated Images

Mohammad AsimChristopher WewerThomas WimmerBernt SchieleJan Eric Lenssen

article2025CVPR114 citations

Introduces a pose-free metric that leverages feed-forward 3D reconstructions from DUSt3R and semantic feature warping to reliably evaluate geometric consistency across generated multi-view images without requiring ground truth.

Listen

Generative artificial intelligence models are increasingly used to synthesize multiple views of 3D objects and scenes from limited 2D images. However, conventional evaluation metrics are poorly suited for this task: distribution-based visual quality metrics ignore whether generated viewpoints physically align in three dimensions, while traditional 3D metrics depend heavily on known camera angles, computationally heavy scene reconstructions, or brittle feature matching. Without an accurate way to evaluate multi-view consistency, researchers and technology leaders face substantial uncertainty when benchmarking novel 3D and video generation systems.

To address this gap, the article introduces MEt3R, a lightweight and automated evaluation metric designed to measure the 3D physical consistency between pairs of generated images. The primary objective is to reliably quantify multi-view consistency without requiring ground-truth 3D reference data or known camera poses, while remaining resilient to changes in scene lighting and independent of general image resolution or blur.

The authors develop MEt3R by combining a feed-forward 3D reconstruction model with robust semantic feature extractors. Given an image pair, the tool uses DUSt3R to estimate pixel-aligned 3D point clouds in a shared space, reprojects the images onto a common viewing plane, and evaluates alignment by computing cosine similarity across upsampled high-resolution semantic features. To benchmark the metric, the authors evaluate a broad suite of leading generative models across 100 multi-view trajectories from the RealEstate10K dataset, full video sequences, and 30 object instances from the Google Scanned Objects benchmark. Additionally, the authors release an open-source multi-view latent diffusion model (MV-LDM) that employs an anchored sampling strategy to generate coherent multi-view sequences.

The findings establish that MEt3R successfully overcomes the core failure modes of existing metrics. First, MEt3R accurately captures subtle, frame-by-frame structural drift and distinguishes perfectly consistent sequences from degraded ones, whereas legacy metrics like TSED classify inconsistent views as consistent or fail to register gradual drift. Second, evaluating generative architectures reveals a clear performance trade-off: single-view models such as GenWarp produce high single-image fidelity but fail to maintain 3D structure (scoring a poor 0.120 on MEt3R), whereas rigid 3D representations like DFM achieve superior consistency (0.026) at the expense of severe image blur. Third, the authors' open-source MV-LDM model achieves the most favorable balance between visual fidelity and spatial coherence, recording a strong consistency score of 0.036 while avoiding the visual degradation seen in 3D-bound baselines. Finally, because MEt3R does not require camera poses, it effectively measures geometric stability across standard video generation models, identifying smooth, highly consistent camera trajectories in models like Stable Video Diffusion.

These results provide a practical path forward for teams developing 3D generative vision pipelines. By decoupling geometric consistency from standard visual appeal, MEt3R lowers benchmarking costs, accelerates development timelines, and mitigates the risk of deploying generative models that produce physically implausible geometry. Organizations evaluating or training multi-view, video, or 3D generative frameworks should integrate MEt3R alongside standard distribution-based image quality metrics to monitor the trade-off between geometric stability and visual fidelity. Teams pursuing multi-view synthesis should prioritize multi-view diffusion architectures with anchored generation over purely sequential single-frame generators.

Confidence in the metric is supported by extensive comparative benchmarks against existing metrics across scenes, objects, and video sequences. However, users should note minor boundary limitations: the metric exhibits a baseline lower bound slightly above zero even on real video due to minor residual errors in point-map reconstruction and feature extraction. As feature extraction backbones continue to mature, adopting more robust foundation models can further refine MEt3R's absolute sensitivity.

  • Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). VGGT extends feed-forward visual geometry from pairwise reconstruction methods like DUSt3R to an all-in-one transformer estimating multi-view camera poses, depths, and tracks across entire sequences.
  • Paper: Structured 3D Latents for Scalable and Versatile 3D Generation, Jianfeng Xiang et al. (2025). TRELLIS builds upon multi-view visual features and rectified-flow transformers to produce high-fidelity, view-consistent 3D assets decodable into various representations.
Cover for MET3R: Measuring Multi-View Consistency in Generated Images

Abstract

We introduce MEt3R, a metric for multi-view consistency in generated images. Large-scale generative models for multi-view image generation are rapidly advancing the field of 3D inference from sparse observations. However, due to the nature of generative modeling, traditional reconstruction metrics are not suitable to measure the quality of generated outputs and metrics that are independent of the sampling procedure are desperately needed. In this work, we specifically address the aspect of consistency between generated multi-view images, which can be evaluated independently of the specific scene. Our approach uses DUST3R to obtain dense 3D reconstructions from image pairs in a feed-forward manner, which are used to warp image contents from one view into the other. Then, feature maps of these images are compared to obtain a similarity score that is invariant to view-dependent effects. Using MEt3R, we evaluate the consistency of a large set of previous methods for novel view and video generation, including our open, multi-view latent diffusion model. Code is available online: geometric-rl.mpi-inf.mpg.de/met3r/.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. MEt3R: Measuring Consistency
  • 3.1. Stereo Reconstruction with DUSt3R
  • 3.2. High-Resolution Feature Similarity
  • 4. Multi-View Latent Diffusion Model
  • 5. Experiments
  • 5.1. Experimental Setup
  • 5.2. Validating MEt3R
  • 5.3. Evaluations of Models
  • 5.3.1 Multi-View Generation
  • 5.3.2 Video Generation
  • 5.3.3 Object-Level Generation
  • 5.4. Analyzing Alternative Similarities
  • 6. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — MEt3R pose-free multi-view consistency metric

    equation

    MEt3R measures the geometric consistency of a pair of generated images without requiring ground-truth images, camera poses, or a particular scene-generation model. For images I1I_1 and I2I_2, let S(I1,I2)S(I_1,I_2) and S(I2,I1)S(I_2,I_1) be directed feature similarities obtained after reconstructing and reprojecting the two images into a shared view. The metric is

    MEt3R⁡(I1,I2)=1−12(S(I1,I2)+S(I2,I1)).\operatorname{MEt3R}(I_1,I_2)=1-\frac{1}{2}\left(S(I_1,I_2)+S(I_2,I_1)\right).

    Each directed similarity lies in [−1,1][-1,1], so MEt3R⁡(I1,I2)∈[0,2]\operatorname{MEt3R}(I_1,I_2)\in[0,2]; lower values indicate better consistency. The score is symmetric by construction and is intended to penalize incompatible 3D content while not penalizing images merely because they differ from a ground-truth view, image-quality distribution, or pixel-level appearance. The two directed similarities are approximately symmetric in practice, so one direction can be used as a runtime approximation.

  2. Knowl 2 — Pose-free dense reconstruction with DUSt3R

    model/method

    The first stage of MEt3R uses the feed-forward DUSt3R model to reconstruct a dense, pixel-aligned 3D point map from an image pair. For images I1I_1 and I2I_2 of height HH and width WW, DUSt3R predicts

    X1,X2=Ψ(I1,I2),X1,X2∈RH×W×3,X_1,X_2=\Psi(I_1,I_2),\qquad X_1,X_2\in\mathbb{R}^{H\times W\times 3},

    where Ψ\Psi is the DUSt3R network and each point in XkX_k is associated with a pixel in IkI_k. Both point maps are expressed in the camera coordinate frame of I1I_1, so the two images can be compared without explicitly estimating or supplying camera poses. DUSt3R uses a shared vision-transformer backbone, cross-view transformer decoders, and regression heads that output the two point maps. MEt3R uses the point maps only for geometric unprojection and reprojection; it does not use the additional feature correspondences provided by MASt3R.

  3. Knowl 3 — Feature-space reprojection and similarity computation

    equation

    MEt3R compares reprojected high-resolution features rather than RGB pixels. DINO extracts semantic feature maps from I1I_1 and I2I_2, and FeatUp upsamples them using image-adaptive joint bilateral upsampling. Let F1,F2∈RH×W×dF_1,F_2\in\mathbb{R}^{H\times W\times d} be the resulting feature maps, where dd is the feature dimension. The operator P(Fk,Xk)P(F_k,X_k) assigns each 3D point in XkX_k the feature vector of its corresponding pixel in FkF_k and rasterizes the feature point cloud into the image plane of I1I_1:

    F^1=P(F1,X1),F^2=P(F2,X2).\widehat F_1=P(F_1,X_1),\qquad \widehat F_2=P(F_2,X_2).

    For pixel coordinates (i,j)(i,j), let f^1ij,f^2ij∈Rd\widehat f_1^{ij},\widehat f_2^{ij}\in\mathbb{R}^{d} be the corresponding feature vectors, let mij∈{0,1}m_{ij}\in\{0,1\} indicate whether both reprojected maps overlap at that pixel, and let ∣M∣=∑i=1W∑j=1Hmij|\mathcal M|=\sum_{i=1}^{W}\sum_{j=1}^{H}m_{ij} be the number of overlapping pixels. The directed similarity is the mean cosine similarity over the overlap:

    S(I1,I2)=1∣M∣∑i=1W∑j=1Hmijf^1ij⋅f^2ij∥f^1ij∥2 ∥f^2ij∥2.S(I_1,I_2)=\frac{1}{|\mathcal M|}\sum_{i=1}^{W}\sum_{j=1}^{H}m_{ij}\frac{\widehat f_1^{ij}\cdot\widehat f_2^{ij}}{\|\widehat f_1^{ij}\|_2\,\|\widehat f_2^{ij}\|_2}.

    The reverse similarity is computed analogously after expressing the projections in the coordinate frame of I2I_2. DINO features are used because they preserve semantic and image-level structure while being less sensitive than RGB comparisons to view-dependent lighting and reflections.

  4. Knowl 4 — Open-source multi-view latent diffusion model

    model/method

    The paper introduces MV-LDM, an open-source multi-view latent diffusion model used as both a benchmark system and a generated-view baseline. MV-LDM encodes images into Stable Diffusion 2.1's latent space with its pretrained VAE, concatenates ray maps to the input latents to provide camera-pose information, and adapts a 2D UNet by adding attention between views at every UNet block. The full model is fine-tuned on RealEstate10K for 1.651.65 million iterations. The architecture is designed to impose a stronger joint multi-view prior than models that generate one target view independently at a time.

  5. Knowl 5 — Anchored generation for long multi-view sequences

    algorithm

    MV-LDM generates large sets of views with an anchored sampling strategy. Given one conditioning image and a set of widely distributed target cameras, the procedure is:

    Input: One conditioning image, target camera views, and the trained MV-LDM
    Output: Generated images for all target views
    1. Sample four anchor images for widely distributed cameras, conditioning each on the initial input image.
    2. For every remaining target view, identify the closest anchor view.
    3. Generate the target view conditioned on both the closest anchor image and the initial input image.
    4. Return the anchor and non-anchor images as the multi-view sequence.

    The anchors are intended to prevent error accumulation from autoregressively conditioning each new view on the immediately preceding generated view. In the experiments, switching between anchors produces periodic transition artifacts that are visible as spikes in several consistency metrics, including MEt3R.

  6. Knowl 6 — Benchmark protocol for generated views, videos, and objects

    experimental setup

    The evaluation uses three generation settings. For scene-level multi-view generation, the authors collect 100100 sequences from the RealEstate10K test set, use the first image as the conditioning image, and generate 8080 target views. MEt3R is evaluated on consecutive generated-image pairs in a sliding window, using 256×256256\times256 inputs. DFM outputs are upsampled from 128×128128\times128, and GenWarp outputs are bilinearly downsampled from 512×512512\times512. The sliding-window protocol maximizes overlap, includes extrapolated regions not visible in the conditioning image, and measures how consistency changes as views move farther from the input.

    The scene-level methods are GenWarp, PhotoNVS, DFM, and MV-LDM. For image-to-video evaluation, the authors use Stable Video Diffusion, Ruyi-Mini-7B, and I2VGen-XL on the same test sequences, limit videos to 4848 frames for memory reasons, and resize them to approximately 256×256256\times256 while preserving aspect ratio. These video models do not provide explicit camera control, so their trajectories are not directly equivalent.

    For object-level evaluation, the authors use Google Scanned Objects. Each example contains 1616 views uniformly covering a closed 360∘360^\circ rotation, with the 0∘0^\circ view used as the conditioning image. EpiDiff and SyncDreamer directly generate 1616 views, while VideoMV generates 3232 views that are uniformly downsampled to 1616.

  7. Knowl 7 — Quantitative consistency and quality comparison

    data/table

    The reported averages compare consistency, image quality, and video-distribution metrics across scene-level multi-view and image-to-video generators. MEt3R is lower-is-better; TSED is higher-is-better in the reported table; SED, FVD, and FID are lower-is-better; and FWS measured by PSNR is higher-is-better. A dash indicates that the metric was not computed or is not applicable to that video model.

    Could not parse LaTeX table

    Among multi-view methods, DFM has the best MEt3R score, but its 73.0273.02 FID indicates poor agreement with the image distribution, consistent with its blurry renderings. GenWarp has the best multi-view FID of 29.8029.80 but the worst MEt3R score of 0.1200.120, showing the distinction between image quality and 3D consistency. MV-LDM obtains a favorable compromise, with MEt3R 0.0360.036, FID 37.2937.29, and FWS PSNR 28.4628.46. Among video models, SVD has the best MEt3R score at 0.0320.032, followed by Ruyi-Mini-7B at 0.0470.047 and I2VGen-XL at 0.0500.050.

  8. Knowl 8 — MEt3R reveals gradual and localized consistency failures

    empirical result

    Across the 100100 RealEstate10K sequences, MEt3R generally increases as generated views move farther from the conditioning image, indicating progressively worse pairwise consistency. It separates the behavior of GenWarp, PhotoNVS, MV-LDM, DFM, and real videos more clearly than TSED: TSED gives nearly indistinguishable scores to several methods even when their visual consistency differs substantially. MEt3R also detects periodic spikes in MV-LDM, which correspond to transitions between anchored generation branches.

    Real videos, treated as approximately 3D-consistent, produce a small nonzero MEt3R lower bound rather than exactly zero. The paper attributes this residual to DUSt3R point-map alignment errors and small 3D inconsistencies in the DINO features. The metric therefore provides a gradual score with a practical lower bound instead of a binary consistent/inconsistent decision.

  9. Knowl 9 — Cross-domain evaluation of generation models

    empirical result

    For scene-level multi-view generation, GenWarp has the poorest consistency because it generates one view at a time and changes scene content substantially across views. PhotoNVS is somewhat more consistent but produces lower-quality images. MV-LDM improves consistency by diffusing multiple views jointly, while DFM is the most consistent because it uses an internal 3D representation; that inductive bias also produces blur and worsens image-distribution metrics.

    For video generation, SVD is the most consistent of the evaluated video models, but it produces smoother and shorter camera trajectories. Ruyi-Mini-7B and I2VGen-XL generate larger camera motions at the cost of consistency; Ruyi-Mini-7B exhibits abrupt MEt3R spikes associated with unstable motion. I2VGen-XL starts with a relatively high error and gradually improves as generation becomes more in-distribution while preserving global structure.

    For object-level 360∘360^\circ generation, SyncDreamer has the best MEt3R behavior, followed by VideoMV and EpiDiff. MEt3R detects a gradual increase in EpiDiff's inconsistency as the view moves farther from the conditioning image, demonstrating sensitivity to conditioning distance.

  10. Knowl 10 — Feature-space design is more robust than RGB similarity

    empirical result

    Replacing DINO-feature cosine similarity with RGB-space PSNR or SSIM produces a weaker consistency signal. DFM can score better than real videos under PSNR and SSIM because its low-resolution, blurry renderings are favored by blur-sensitive pixel metrics, whereas real videos contain brightness changes, reflections, and other view-dependent effects that increase pixel error. MEt3R's feature-space comparison reduces sensitivity to these effects and better separates structural 3D consistency from image sharpness.

    The feature backbone also affects the metric. DINOv2 and MaskCLIP compress scores into a narrower range, reducing the separation between highly inconsistent and consistent generations. In the reported ablation, the original DINO features provide more informative separation among methods and capture random-noise generations as strongly inconsistent. The metric remains extensible: a feature backbone with stronger 3D consistency could potentially lower the practical nonzero lower bound.

Coverage note — No substantial contributed material was omitted; appendix-only implementation details and additional qualitative examples were not promoted to separate knowls.

References

  1. 1.Mikołaj Binkowski, Danica J Sutherland, Michael Arbel, and Arthur Gretton. Demystifying mmd gans. arXiv preprint arXiv:1801.01401, 2018. 1, 3
  2. 2.Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, Varun Jampani, and Robin Rombach. Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023. 5, 6, 7, 14
  3. 3.Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), page 22563–22575. IEEE, 2023. 1
  4. 4.Mathilde Caron, Hugo Touvron, Ishan Misra, Herve Jegou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, 2021. 2, 3, 5, 8
  5. 5.Eric R. Chan, Koki Nagano, Matthew A. Chan, Alexander W. Bergman, Jeong Joon Park, Axel Levy, Miika Aittala, Shalini De Mello, Tero Karras, and Gordon Wetzstein. GeNVS: Generative novel view synthesis with 3D-aware diffusion models. In ICCV, 2023. 2
  6. 6.Hansheng Chen, Jiatao Gu, Anpei Chen, Wei Tian, Zhuowen Tu, Lingjie Liu, and Hao Su. Single-stage diffusion NeRF: A unified approach to 3d generation and reconstruction. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, 2023. 1
  7. 7.Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In arxiv, 2022. 2
  8. 8.Alexey Dosovitskiy. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 3
  9. 9.Laura Downs, Anthony Francis, Nate Koenig, Brandon Kinman, Ryan Hickman, Krista Reymann, Thomas B. McHugh, and Vincent Vanhoucke. Google scanned objects: A high-quality dataset of 3d scanned household items, 2022. 5, 7, 8
  10. 10.Stephanie Fu, Mark Hamilton, Laura Brandt, Axel Feldman, Zhoutong Zhang, and William T Freeman. Featup: A model-agnostic framework for features at any resolution. arXiv preprint arXiv:2403.10516, 2024. 2, 3, 4
  11. 11.Ruiqi Gao*, Aleksander Holynski*, Philipp Henzler, Arthur Brussee, Ricardo Martin-Brualla, Pratul P. Srinivasan, Jonathan T. Barron, and Ben Poole*. Cat3d: Create anything in 3d with multi-view diffusion models. NeurIPS, 2024. 1, 2, 4, 12
  12. 12.Junlin Han, Filippos Kokkinos, and Philip Torr. Vfusion3d: Learning scalable 3d generative models from video diffusion models. In European Conference on Computer Vision, pages 333–350. Springer, 2025. 2
  13. 13.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. GANs trained by a two time-scale update rule converge to a local nash equilibrium. In NeurIPS, 2017. 1, 3, 6, 7
  14. 14.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In NeurIPS, 2020. 1, 12
  15. 15.Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. Advances in Neural Information Processing Systems, 35:8633–8646, 2022. 1
  16. 16.Zehuan Huang, Hao Wen, Junting Dong, Yaohui Wang, Yangguang Li, Xinyuan Chen, Yan-Pei Cao, Ding Liang, Yu Qiao, Bo Dai, et al. Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9784–9794, 2024. 5, 7, 8
  17. 17.Sadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner, Ayan Chakrabarti, and Sanjiv Kumar. Rethinking FID: Towards a better evaluation metric for image generation. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), page 9307–9315. IEEE, 2024. 1, 3
  18. 18.Justin Johnson, Nikhila Ravi, Jeremy Reizenstein, David Novotny, Shubham Tulsiani, Christoph Lassner, and Steve Branson. Accelerating 3d deep learning with pytorch3d. In SIGGRAPH Asia 2020 Courses, page 1–1. ACM, 2020. 4
  19. 19.Vincent Leroy, Yohann Cabon, and Jerome Revaud. Grounding image matching in 3d with mast3r, 2024. 3
  20. 20.Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot one image to 3d object. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), page 9264–9275. IEEE, 2023. 2
  21. 21.Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. In ICLR, 2023. 1
  22. 22.Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. SyncDreamer: Generating multiview-consistent images from a single-view image. In ICLR, 2024. 2, 5, 7, 8
  23. 23.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization, 2019. 12
  24. 24.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, page 405–421. Springer International Publishing, 2020. 3
  25. 25.Norman Müller, Yawar Siddiqui, Lorenzo Porzi, Samuel Rota Bulo, Peter Kontschieder, and Matthias Nießner. DiffRF: Rendering-guided 3d radiance field diffusion. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), page 4328–4338. IEEE, 2023. 1
  26. 26.Norman Müller, Katja Schwarz, Barbara Rössle, Lorenzo Porzi, Samuel Rota Bulo, Matthias Nießner, and Peter Kontschieder. MultiDiff: Consistent novel view synthesis from a single image. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), page 10258–10268. IEEE, 2024. 2
  27. 27.Maxime Oquab, Timothee Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 8
  28. 28.Robin Rombach, Patrick Esser, and Bjorn Ommer. Geometry-free view synthesis: Transformers and no 3d priors. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), page 14336–14346. IEEE, 2021. 2, 3
  29. 29.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), page 10674–10685. IEEE, 2022. 1, 2, 4, 12, 13
  30. 30.Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. Advances in neural information processing systems, 29, 2016. 3
  31. 31.Philipp Schröppel, Christopher Wewer, Jan Eric Lenssen, Eddy Ilg, and Thomas Brox. Neural point cloud diffusion for disentangled 3d shape and appearance generation. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), page 8785–8794. IEEE, 2024. 1
  32. 32.Junyoung Seo, Kazumi Fukuda, Takashi Shibuya, Takuya Narihira, Naoki Murata, Shoukang Hu, Chieh-Hsin Lai, Seungryong Kim, and Yuki Mitsufuji. GenWarp: Single image to novel views with semantic-preserving generative warping. arXiv preprint arXiv:2405.17251, 2024. 2, 3, 5, 6, 7, 14, 15, 16
  33. 33.Yichun Shi, Peng Wang, Jianglong Ye, Long Mai, Kejie Li, and Xiao Yang. MVDream: Multi-view diffusion for 3d generation. In The Twelfth International Conference on Learning Representations, 2024. 2, 12
  34. 34.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In ICLR, 2021. 1, 12
  35. 35.Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020. 1
  36. 36.CreateAI Team. Ruyi-mini-7b. https://github.com/IamCreateAI/Ruyi-Models, 2024. 5, 6, 7, 14
  37. 37.Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow, 2020. 6, 13
  38. 38.Ayush Tewari, Tianwei Yin, George Cazenavette, Semon Rezchikov, Joshua B. Tenenbaum, Fredo Durand, William T. Freeman, and Vincent Sitzmann. Diffusion with forward models: Solving stochastic inverse problems without direct supervision. In arXiv, 2023. 1, 2, 3, 5, 6, 7, 8, 14, 15, 16
  39. 39.Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717, 2018. 3, 6, 7
  40. 40.Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. FVD: A new metric for video generation, 2019. 6, 7
  41. 41.Vikram Voleti, Chun-Han Yao, Mark Boss, Adam Letts, David Pankratz, Dmitry Tochilkin, Christian Laforte, Robin Rombach, and Varun Jampani. Sv3d: Novel multi-view synthesis and 3d generation from a single image using latent video diffusion. In European Conference on Computer Vision, pages 439–457. Springer, 2025. 1, 2
  42. 42.Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vision made easy. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), page 20697–20709. IEEE, 2024. 2, 3, 5, 14, 15
  43. 43.Daniel Watson, William Chan, Ricardo Martin-Brualla, Jonathan Ho, Andrea Tagliasacchi, and Mohammad Norouzi. Novel view synthesis with diffusion models. arXiv preprint arXiv:2210.04628, 2022. 2, 3
  44. 44.Christopher Wewer, Kevin Raj, Eddy Ilg, Bernt Schiele, and Jan Eric Lenssen. latentSplat: Autoencoding variational gaussians for fast generalizable 3d reconstruction. In arxiv, 2024. 2
  45. 45.Rundi Wu, Ben Mildenhall, Philipp Henzler, Keunhong Park, Ruiqi Gao, Daniel Watson, Pratul P. Srinivasan, Dor Verbin, Jonathan T. Barron, Ben Poole, and Aleksander Hołynski. ReconFusion: 3d reconstruction with diffusion priors. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), page 21551–21561. IEEE, 2024. 1, 2
  46. 46.Yiming Xie, Chun-Han Yao, Vikram Voleti, Huaizu Jiang, and Varun Jampani. Sv4d: Dynamic 3d content generation with multi-frame and multi-view consistency. arXiv preprint arXiv:2407.17470, 2024. 3
  47. 47.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelNeRF: Neural radiance fields from one or few images. In CVPR, 2021. 8, 16
  48. 48.Jason J. Yu, Fereshteh Forghani, Konstantinos G. Derpanis, and Marcus A. Brubaker. Long-term photometric consistent novel view synthesis with diffusion models. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), page 7071–7081. IEEE, 2023. 2, 3, 5, 6, 7, 13, 14, 15, 16
  49. 49.Shiwei Zhang, Jiayu Wang, Yingya Zhang, Kang Zhao, Hangjie Yuan, Zhiwu Qin, Xiang Wang, Deli Zhao, and Jingren Zhou. I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models, 2023. 5, 6, 7, 14
  50. 50.Zhuowen Tu Zheng Ding, Jieke Wang. Open-vocabulary universal image segmentation with MaskCLIP. In International Conference on Machine Learning, 2023. 8
  51. 51.Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: Learning view synthesis using multiplane images. In SIGGRAPH, 2018. 5, 7, 12
  52. 52.Zhizhuo Zhou and Shubham Tulsiani. SparseFusion: Distilling view-conditioned diffusion for 3d reconstruction. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), page 12588–12597. IEEE, 2023. 1, 2
  53. 53.Qi Zuo, Xiaodong Gu, Lingteng Qiu, Yuan Dong, Zhengyi Zhao, Weihao Yuan, Rui Peng, Siyu Zhu, Zilong Dong, Liefeng Bo, and Qixing Huang. Videomv: Consistent multi-view generation based on large video generative model, 2024. 5, 7, 8

Citation

MLA
Asim, M., et al. “MEt3R: Measuring Multi-View Consistency in Generated Images”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 6034–44, https://doi.org/10.1109/CVPR52734.2025.00566.
APA
Asim, M., Wewer, C., Wimmer, T., Schiele, B., & Lenssen, J. E. (2025). MEt3R: Measuring Multi-View Consistency in Generated Images. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 6034–6044. https://doi.org/10.1109/CVPR52734.2025.00566
Chicago
Asim, M., C. Wewer, T. Wimmer, B. Schiele, and J. E. Lenssen. 2025. “MEt3R: Measuring Multi-View Consistency in Generated Images”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 6034–44. https://doi.org/10.1109/CVPR52734.2025.00566.
Harvard
Asim, M. et al. (2025) “MEt3R: Measuring Multi-View Consistency in Generated Images”, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 6034–6044. Available at: https://doi.org/10.1109/CVPR52734.2025.00566.
Vancouver
1. Asim M, Wewer C, Wimmer T, Schiele B, Lenssen JE (2025) MEt3R: Measuring Multi-View Consistency in Generated Images. In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 6034–6044

BibTeX

@inproceedings{Asim_2025, title={MEt3R: Measuring Multi-View Consistency in Generated Images}, url={http://dx.doi.org/10.1109/CVPR52734.2025.00566}, DOI={10.1109/cvpr52734.2025.00566}, booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Asim, Mohammad and Wewer, Christopher and Wimmer, Thomas and Schiele, Bernt and Lenssen, Jan Eric}, year={2025}, month=June, pages={6034–6044} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE