MET3R: Measuring Multi-View Consistency in Generated Images
Mohammad AsimChristopher WewerThomas WimmerBernt SchieleJan Eric Lenssen
Introduces a pose-free metric that leverages feed-forward 3D reconstructions from DUSt3R and semantic feature warping to reliably evaluate geometric consistency across generated multi-view images without requiring ground truth.
Generative artificial intelligence models are increasingly used to synthesize multiple views of 3D objects and scenes from limited 2D images. However, conventional evaluation metrics are poorly suited for this task: distribution-based visual quality metrics ignore whether generated viewpoints physically align in three dimensions, while traditional 3D metrics depend heavily on known camera angles, computationally heavy scene reconstructions, or brittle feature matching. Without an accurate way to evaluate multi-view consistency, researchers and technology leaders face substantial uncertainty when benchmarking novel 3D and video generation systems.
To address this gap, the article introduces MEt3R, a lightweight and automated evaluation metric designed to measure the 3D physical consistency between pairs of generated images. The primary objective is to reliably quantify multi-view consistency without requiring ground-truth 3D reference data or known camera poses, while remaining resilient to changes in scene lighting and independent of general image resolution or blur.
The authors develop MEt3R by combining a feed-forward 3D reconstruction model with robust semantic feature extractors. Given an image pair, the tool uses DUSt3R to estimate pixel-aligned 3D point clouds in a shared space, reprojects the images onto a common viewing plane, and evaluates alignment by computing cosine similarity across upsampled high-resolution semantic features. To benchmark the metric, the authors evaluate a broad suite of leading generative models across 100 multi-view trajectories from the RealEstate10K dataset, full video sequences, and 30 object instances from the Google Scanned Objects benchmark. Additionally, the authors release an open-source multi-view latent diffusion model (MV-LDM) that employs an anchored sampling strategy to generate coherent multi-view sequences.
The findings establish that MEt3R successfully overcomes the core failure modes of existing metrics. First, MEt3R accurately captures subtle, frame-by-frame structural drift and distinguishes perfectly consistent sequences from degraded ones, whereas legacy metrics like TSED classify inconsistent views as consistent or fail to register gradual drift. Second, evaluating generative architectures reveals a clear performance trade-off: single-view models such as GenWarp produce high single-image fidelity but fail to maintain 3D structure (scoring a poor 0.120 on MEt3R), whereas rigid 3D representations like DFM achieve superior consistency (0.026) at the expense of severe image blur. Third, the authors' open-source MV-LDM model achieves the most favorable balance between visual fidelity and spatial coherence, recording a strong consistency score of 0.036 while avoiding the visual degradation seen in 3D-bound baselines. Finally, because MEt3R does not require camera poses, it effectively measures geometric stability across standard video generation models, identifying smooth, highly consistent camera trajectories in models like Stable Video Diffusion.
These results provide a practical path forward for teams developing 3D generative vision pipelines. By decoupling geometric consistency from standard visual appeal, MEt3R lowers benchmarking costs, accelerates development timelines, and mitigates the risk of deploying generative models that produce physically implausible geometry. Organizations evaluating or training multi-view, video, or 3D generative frameworks should integrate MEt3R alongside standard distribution-based image quality metrics to monitor the trade-off between geometric stability and visual fidelity. Teams pursuing multi-view synthesis should prioritize multi-view diffusion architectures with anchored generation over purely sequential single-frame generators.
Confidence in the metric is supported by extensive comparative benchmarks against existing metrics across scenes, objects, and video sequences. However, users should note minor boundary limitations: the metric exhibits a baseline lower bound slightly above zero even on real video due to minor residual errors in point-map reconstruction and feature extraction. As feature extraction backbones continue to mature, adopting more robust foundation models can further refine MEt3R's absolute sensitivity.
- Paper: DUSt3R: Geometric 3D Vision Made Easy, Shuzhe Wang et al. (2023). DUSt3R introduces the feed-forward, uncalibrated dense stereo pointmap regression framework that MET3R directly adapts to perform 3D reconstruction and cross-view image warping.
- Paper: CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion, Philippe Weinzaepfel et al. (2022). CroCo establishes the cross-view visual completion transformer architecture that forms the algorithmic foundation for DUSt3R's multi-view geometry models.
- Paper: Zero-1-to-3: Zero-shot One Image to 3D Object, Ruoshi Liu et al. (2023). Zero-1-to-3 pioneered viewpoint-conditioned diffusion models for novel view synthesis, providing essential context on the multi-view generative paradigms evaluated by MET3R.
- Paper: SyncDreamer: Generating Multiview-consistent Images from a Single-view Image, Yuan Liu et al. (2024). SyncDreamer addresses multi-view consistency in single-image 3D generation through synchronized attention, illustrating the generative challenges and consistency shortcomings that MET3R seeks to benchmark.
- Paper: GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping, Junyoung Seo et al. (2024). GenWarp explores geometry-guided image warping and generative inpainting between viewpoints, establishing key concepts for feature-level cross-view warping.
- Paper: Image Quality Assessment: Unifying Structure and Texture Similarity, Keyan Ding et al. (2020). This work formulates perceptual metrics that separate structural correspondence from texture variations, motivating MET3R's design of feature-based similarity invariant to view-dependent effects.
- Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). VGGT extends feed-forward visual geometry from pairwise reconstruction methods like DUSt3R to an all-in-one transformer estimating multi-view camera poses, depths, and tracks across entire sequences.
- Paper: Structured 3D Latents for Scalable and Versatile 3D Generation, Jianfeng Xiang et al. (2025). TRELLIS builds upon multi-view visual features and rectified-flow transformers to produce high-fidelity, view-consistent 3D assets decodable into various representations.
