Improved Precision and Recall Metric for Assessing Generative Models

Tuomas KynkäänniemiTero KarrasSamuli LaineJaakko LehtinenTimo Aila

article2019NeurIPS1,365 citations

Introduces an improved precision and recall metric that reliably separates sample quality from distribution coverage in generative models, exposing failure modes in standard evaluation metrics and guiding practical architectural improvements in state-of-the-art image generators.

Listen

Generative artificial intelligence models have advanced rapidly in synthesizing photorealistic images, yet evaluating their performance remains a critical bottleneck. Standard metrics, notably the Fréchet Inception Distance (FID), collapse two distinct operational goals—individual image fidelity (precision) and dataset diversity (recall)—into a single aggregate score. This single-value approach obscures crucial performance trade-offs, making it difficult for researchers and decision-makers to diagnose failure modes such as distorted outputs or lack of sample variety.

The article develops and validates an improved evaluation framework that separately and reliably quantifies precision (the fraction of generated images that look realistic) and recall (the proportion of real-world training data variation the model successfully reproduces). In addition, it establishes a continuous individual sample "realism score" to assess single-image fidelity and latent space interpolations.

To accomplish this, the authors construct non-parametric representations of the real and generated data distributions in a high-dimensional feature space derived from a standard image classifier. By enclosing data points within hyperspheres determined by their k-nearest neighbors (setting k=3), the method checks whether a generated image falls within the estimated support volume of real data, and vice versa. The framework was evaluated across standard benchmarks using prominent models (StyleGAN and BigGAN) with 50,000 samples, avoiding the estimation errors that plague earlier relative-density methods.

The analysis yields four key findings. First, the metric successfully disentangles fidelity from diversity: conventional FID heavily penalizes restricted diversity while often masking severe visual artifacts, whereas the proposed method accurately captures both dimensions. Second, the metric enables Pareto-frontier analysis across architectural choices; for example, removing instance normalization in StyleGAN's core blocks improved recall and achieved a new state-of-the-art FID of 4.16. Third, a principled comparison of post-processing truncation techniques revealed that clamping low-density latent vectors to a high-density boundary provides a superior balance of quality and coverage compared to traditional linear interpolation or rejection sampling. Fourth, latent space interpolations between realistic endpoints stayed within realistic bounds in 97.6% of tested paths, demonstrating that high-quality regions within StyleGAN's intermediate latent space are highly convex.

These findings have direct practical implications for model deployment and development budgets. In production environments where output defects carry high brand or operational risk, teams can intentionally optimize for precision over recall, rather than relying on blunt aggregate metrics. Conversely, applications requiring broad creative variety can benchmark coverage directly. The results also clarify that fine-tuning training configurations can achieve the benefits of truncation techniques without requiring artificial post-hoc data restrictions.

The article recommends that practitioners adopt separate precision and recall evaluations alongside aggregate metrics when selecting generative models, and use clamping methods when truncation is required. Further research should explore whether training configurations can entirely eliminate the need for post-processing truncation. While the metric relies on feature embeddings that require adequate sample counts (typically 50,000 images) and slightly overestimates volume at sparse dataset boundaries, it demonstrates high reliability and stability across standard image benchmarks.

arXiv: 1904.06991
Cover for Improved Precision and Recall Metric for Assessing Generative Models

Abstract

The ability to automatically estimate the quality and coverage of the samples produced by a generative model is a vital requirement for driving algorithm research. We present an evaluation metric that can separately and reliably measure both of these aspects in image generation tasks by forming explicit, non-parametric representations of the manifolds of real and generated data. We demonstrate the effectiveness of our metric in StyleGAN and BigGAN by providing several illustrative examples where existing metrics yield uninformative or contradictory results. Furthermore, we analyze multiple design variants of StyleGAN to better understand the relationships between the model architecture, training methods, and the properties of the resulting sample distribution. In the process, we identify new variants that improve the state-of-the-art. We also perform the first principled analysis of truncation methods and identify an improved method. Finally, we extend our metric to estimate the perceptual quality of individual samples, and use this to study latent space interpolations.

Table of Contents

  • 1 Introduction
  • 1.1 Background
  • 2 Improved precision and recall metric using 𝒌\boldsymbol{k}-nearest neighbors
  • 3 Precision and recall of state-of-the-art generative models
  • 4 Using precision and recall to analyze and improve StyleGAN
  • 4.1 Network architectures and training configurations
  • 5 Estimating the quality of individual samples
  • 5.1 Quality of interpolations
  • 6 Conclusion
  • 7 Acknowledgements
  • References
  • A Pseudocode and implementation details
  • B Precision and recall with synthetic dataset
  • C Analysis of truncation methods
  • D Quality of samples and interpolations

Knowls

  1. Knowl 1 — k-NN Manifold Precision and Recall Metric for Generative Models

    definition

    Precision and recall for generative models assess two distinct aspects of sample generation: precision quantifies fidelity (the fraction of generated images that look realistic), while recall quantifies coverage (the fraction of the training data manifold covered by the generator).

    Let Φr={ϕr(i)}i=1N\Phi_r = \{\phi_r^{(i)}\}_{i=1}^N and Φg={ϕg(j)}j=1N\Phi_g = \{\phi_g^{(j)}\}_{j=1}^N denote sets of NN high-dimensional feature representations of real images Xr∼PrX_r \sim P_r and generated images Xg∼PgX_g \sim P_g, extracted using a pre-trained feature network F\mathcal{F} (such as VGG-16). For any feature set Φ\Phi, the data manifold is approximated non-parametrically as the union of Euclidean hyperspheres centered at each sample ϕ′∈Φ\phi' \in \Phi, where each sphere's radius is the Euclidean distance from ϕ′\phi' to its kk-th nearest neighbor in Φ\Phi, denoted NNk(ϕ′,Φ)\text{NN}_k(\phi', \Phi).

    The binary membership query determining whether a feature vector ϕ\phi lies within the estimated manifold of Φ\Phi is: f(ϕ,Φ)={1,if ∥ϕ−ϕ′∥2≤∥ϕ′−NNk(ϕ′,Φ)∥2 for at least one ϕ′∈Φ,0,otherwise.f(\phi, \Phi) = \begin{cases} 1, & \text{if } \|\phi - \phi'\|_2 \le \|\phi' - \text{NN}_k(\phi', \Phi)\|_2 \text{ for at least one } \phi' \in \Phi, \\ 0, & \text{otherwise.} \end{cases}

    Precision and recall are then defined as: precision(Φr,Φg)=1∣Φg∣∑ϕg∈Φgf(ϕg,Φr)\text{precision}(\Phi_r, \Phi_g) = \frac{1}{|\Phi_g|} \sum_{\phi_g \in \Phi_g} f(\phi_g, \Phi_r) recall(Φr,Φg)=1∣Φr∣∑ϕr∈Φrf(ϕr,Φg)\text{recall}(\Phi_r, \Phi_g) = \frac{1}{|\Phi_r|} \sum_{\phi_r \in \Phi_r} f(\phi_r, \Phi_g)

    Precision measures the probability that a generated image falls within the estimated support of the real distribution, while recall measures the probability that a real image falls within the estimated support of the generated distribution.

  2. Knowl 2 — Algorithm for k-NN Precision and Recall

    algorithm

    The kk-nearest neighbors precision and recall algorithm computes fidelity and coverage between sets of real and generated images by estimating non-parametric manifold boundaries in a deep feature space.

    function PRECISION-RECALL(XrX_r, XgX_g, F\mathcal{F}, kk)
        Inputs:
            XrX_r: Set of real images
            XgX_g: Set of generated images
            F\mathcal{F}: Pre-trained feature extractor network
            kk: Integer neighborhood size for manifold boundary estimation
        Output:
            precision: Float in [0,1][0, 1] estimating sample quality
            recall: Float in [0,1][0, 1] estimating dataset coverage
        Φr=F(Xr)\Phi_r = \mathcal{F}(X_r)
        Φg=F(Xg)\Phi_g = \mathcal{F}(X_g)
        precision = MANIFOLD-ESTIMATE(Φr\Phi_r, Φg\Phi_g, kk)
        recall = MANIFOLD-ESTIMATE(Φg\Phi_g, Φr\Phi_r, kk)
        return precision, recall
    function MANIFOLD-ESTIMATE(Φa\Phi_a, Φb\Phi_b, kk)
        Inputs:
            Φa\Phi_a: Reference feature set defining the manifold
            Φb\Phi_b: Query feature set evaluated against the manifold
            kk: Neighborhood size parameter
        Output:
            fraction: Fraction of points in Φb\Phi_b falling inside Φa\Phi_a's manifold
        for ϕ∈Φa\phi \in \Phi_a:
            d={∥ϕ−ϕ′∥2 for all ϕ′∈Φa}d = \{ \|\phi - \phi'\|_2 \text{ for all } \phi' \in \Phi_a \}
            rϕ=mink+1(d)r_\phi = \text{mink}_{+1}(d) // (k+1)(k+1)-th smallest distance to exclude ϕ\phi itself
        n=0n = 0
        for ϕ∈Φb\phi \in \Phi_b:
            if ∥ϕ−ϕ′∥2≤rϕ′\|\phi - \phi'\|_2 \le r_{\phi'} for any ϕ′∈Φa\phi' \in \Phi_a:
                n=n+1n = n + 1
        return n/∣Φb∣n / |\Phi_b|

    To optimize throughput, pairwise distance calculations and manifold containment tests can be performed in mini-batches. Evaluating 50,000 real and 50,000 generated images takes approximately 8 minutes on a single NVIDIA Tesla V100 GPU.

  3. Knowl 3 — Continuous Realism Score for Individual Generated Samples

    model/method

    While binary precision evaluates the aggregate quality of a distribution, the continuous realism score R(ϕg,Φr)R(\phi_g, \Phi_r) quantifies how close an individual generated feature vector ϕg\phi_g is to the manifold of real image feature vectors Φr\Phi_r:

    R(ϕg,Φr)=max⁡ϕr∈Φr∗(∥ϕr−NNk(ϕr,Φr)∥2∥ϕg−ϕr∥2)R(\phi_g, \Phi_r) = \max_{\phi_r \in \Phi_r^*} \left( \frac{\|\phi_r - \text{NN}_k(\phi_r, \Phi_r)\|_2}{\|\phi_g - \phi_r\|_2} \right)

    where NNk(ϕr,Φr)\text{NN}_k(\phi_r, \Phi_r) denotes the kk-th nearest neighbor of ϕr\phi_r in Φr\Phi_r.

    Key characteristics:

    • Relation to Binary Precision: When evaluated over all ϕr∈Φr\phi_r \in \Phi_r, R(ϕg,Φr)≥1R(\phi_g, \Phi_r) \ge 1 if and only if the binary indicator f(ϕg,Φr)=1f(\phi_g, \Phi_r) = 1. A score R≥1R \ge 1 indicates that the sample falls inside at least one kk-NN hypersphere of the real data manifold.
    • Hypersphere Pruning: In finite datasets, sparsely populated fringe regions can form very large kk-NN hyperspheres, potentially assigning artificially high scores to low-quality outliers that land in these sparse areas. To prevent this, the set Φr∗\Phi_r^* is restricted to only those real feature vectors whose kk-NN hypersphere radius is smaller than the median radius across Φr\Phi_r (discarding the 50% largest hyperspheres). This pruning yields a conservative manifold boundary and ensures consistent realism scoring.
  4. Knowl 4 — Latent Space Truncation Strategies and Density Clamping

    empirical result

    Analysis of latent space truncation methods in StyleGAN demonstrates how different truncation formulations affect the precision-recall tradeoff:

    • Intermediate vs. Input Latent Space: Performing truncation or interpolation in the intermediate latent space W\mathcal{W} yields a substantially better precision-recall tradeoff than operating in the initial Gaussian latent space Z\mathcal{Z}. Sampling density in Z\mathcal{Z} does not reliably predict image realism.
    • Density-based Rejection vs. Distance-based Rejection: Rejecting latent vectors in W\mathcal{W} based on low probability density under a fitted multivariate Gaussian outperforms rejecting vectors based on Euclidean distance from the mean μw\mu_w.
    • Clamping by Density: Clamping low-density outlier vectors to the boundary of a high-density hyperellipsoid in W\mathcal{W} (by projecting them to their closest points on the hyperellipsoid boundary) achieves the best overall performance. It minimizes the Wasserstein-2 distance between the truncated distribution and the original distribution in W\mathcal{W}, providing strong coverage at the distribution extremes while discarding visual artifacts.
    • Interpolation toward Mean vs. Clamping: Linearly interpolating all latents toward the mean (w′=μw+ψ(w−μw)w' = \mu_w + \psi(w - \mu_w)) alters all latent vectors indiscriminately. This artificially increases central density and inflates precision at the cost of global distribution distortion and degraded Fréchet Inception Distance (FID).
  5. Knowl 5 — Impact of StyleGAN Architectural and Training Variants on Precision-Recall Tradeoffs

    empirical result

    Evaluating StyleGAN training snapshots on FFHQ at 1024×10241024 \times 1024 resolution along the Pareto frontier of precision versus recall reveals the distinct effects of architectural components and regularizers:

    Configuration Best FID (↓\downarrow)
    A) StyleGAN baseline 4.43
    B) No minibatch standard deviation 8.58
    C) Configuration B + low R1R_1 regularization (γ/100\gamma / 100) 10.34
    D) No progressive growing 6.14
    E) Random translation (discriminator inputs translated by −16…16-16 \dots 16 px) 4.27
    F) No instance normalization in AdaIN 4.16

    Key observations:

    • Minibatch Standard Deviation Layer: Removing this layer (Configuration B) severely decreases recall while increasing precision, shifting the model toward lower output variation and worsening FID to 8.58.
    • R1R_1 Regularization Weight (γ\gamma): Reducing γ\gamma by 100×100\times (Configuration C) shifts the balance even further toward high precision and low recall (FID 10.34).
    • Progressive Growing: Disabling progressive growing (Configuration D) reduces both precision and recall simultaneously (FID 6.14).
    • Discriminator Input Translation: Translating discriminator input images randomly by [−16,16][-16, 16] pixels (Configuration E) improves precision at a slight cost to recall, improving FID to 4.27.
    • Instance Normalization in AdaIN: Disabling instance normalization within Adaptive Instance Normalization layers (Configuration F) improves recall and achieves a state-of-the-art FID of 4.16.
    • FID Bias: Standard FID strongly favors configurations with high recall (Configurations A and F) and penalizes models that favor high individual sample precision at the expense of diversity (Configurations B and C).
  6. Knowl 6 — BigGAN Precision and Recall Characteristics Across ImageNet Classes

    empirical result

    Applying kk-NN precision and recall to BigGAN across ImageNet classes reveals distinct structural behaviors:

    • Textural / Simple Classes ("Easy"): Classes such as "Lemon" and "Broccoli" exhibit high precision (>0.85>0.85) but very low recall (<0.15<0.15). The generator produces realistic images but misses substantial intra-class variation. Despite missing significant variation, these classes achieve low FID values (FID=46.4\text{FID} = 46.4 for Lemon, FID=40.2\text{FID} = 40.2 for Broccoli) because low intrinsic variance in feature space results in low Wasserstein-2 distance.
    • Complex / Structured Classes ("Difficult"): Classes requiring global geometry or unaligned human features (such as "Trumpet", "Park bench", and "Baseball player") exhibit higher recall (0.25−0.400.25 - 0.40) but noticeably lower precision (0.55−0.750.55 - 0.75). BigGAN covers a larger portion of their diverse training manifolds, but generated samples often contain visible geometric distortions (FID=100.4\text{FID} = 100.4 for Trumpet, FID=80.3\text{FID} = 80.3 for Park bench).
    • Well-Represented Classes (Dogs and Cats): Classes with large aggregate representation in ImageNet (e.g., "Great Pyrenees" among ≈100\approx 100 dog breeds; "Egyptian cat" among ≈10\approx 10 cat classes) attain both high precision and high recall, benefiting from the extensive shared training data.
  7. Knowl 7 — Convexity Analysis of Realistic Subsets in StyleGAN Intermediate Latent Space

    empirical result

    Using the continuous realism score RR, linear interpolations in StyleGAN's intermediate latent space W\mathcal{W} provide insights into the geometry of realistic latent representations:

    • Interpolation Quality Tracking: Along linear interpolation paths w(t)=(1−t)w0+tw1w(t) = (1-t)w_0 + t w_1 in W\mathcal{W}, the realism score R(w(t),Φr)R(w(t), \Phi_r) corresponds closely with visual image quality. Paths connecting endpoints within the real manifold maintain high realism scores and consistent image backgrounds, whereas paths crossing non-manifold regions show sharp drops in RR corresponding to distorted intermediate images.
    • Convexity of Realistic W\mathcal{W} Subsets: In a test of 500,000 linear interpolation paths between randomly sampled latent vectors w0,w1∈Ww_0, w_1 \in \mathcal{W} that both have R≥1R \ge 1, an interpolation path was defined as straying outside the manifold if more than 25%25\% of its intermediate points had R<0.9R < 0.9. Under this criterion, only 2.4%2.4\% of paths crossed unrealistic regions, demonstrating that the subset of W\mathcal{W} corresponding to realistic images is highly convex.
  8. Knowl 8 — Hyperparameter Configuration and Feature Space for k-NN Manifold Estimation

    experimental setup

    The kk-NN precision and recall metric relies on the following standard configuration:

    • Feature Space: Real and generated images are embedded using a pre-trained VGG-16 classifier, extracting the activation vector after the second fully connected layer (FC2). This feature representation emphasizes semantic image content over exact spatial pixel alignment. Inception-v3 features yield qualitatively similar metric trajectories.
    • Neighborhood Size (kk): The hyperparameter kk controls the volume of the estimated manifold hyperspheres. Higher values of kk increase both precision and recall estimates, while lower values decrease them. A setting of k=3k = 3 provides a robust balance that avoids saturation at 0.00.0 or 1.01.0 across diverse GAN architectures and datasets.
    • Sample Count (NN): Evaluations use N=∣Φr∣=∣Φg∣=50,000N = |\Phi_r| = |\Phi_g| = 50,000 samples, matching standard practice for FID and providing low variance in the estimated manifold volumes.

Coverage note — All core contributions—the k-NN precision and recall metrics, the algorithm, the continuous realism score, the StyleGAN architecture ablation experiments, the truncation method comparison, the BigGAN ImageNet analysis, the latent space convexity study, and the hyperparameter analysis—are fully covered. Synthetic Gaussian mixture toy experiments (Appendix B) were omitted as they only replicate the validation setup of prior work.

References

  1. 1.S. Arora and Y. Zhang. Do GANs actually learn the distribution? An empirical study. CoRR, abs/1706.08224, 2017.
  2. 2.M. Bińkowski, D. J. Sutherland, M. Arbel, and A. Gretton. Demystifying MMD GANs. CoRR, abs/1801.01401, 2018.
  3. 3.J. Branke, J. Branke, K. Deb, K. Miettinen, and R. Słowiński. Multiobjective optimization: Interactive and evolutionary approaches, volume 5252. Springer Science & Business Media, 2008.
  4. 4.A. Brock, J. Donahue, and K. Simonyan. Large scale GAN training for high fidelity natural image synthesis. In Proc. ICLR, 2019.
  5. 5.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In Proc. CVPR, 2009.
  6. 6.L. Dinh, J. Sohl-Dickstein, and S. Bengio. Density estimation using Real NVP. CoRR, abs/1605.08803, 2016.
  7. 7.D. Eberly. Distance from a point to an ellipse, an ellipsoid, or a hyperellipsoid. Geometric Tools, LLC, 2011.
  8. 8.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative Adversarial Networks. In NIPS, 2014.
  9. 9.A. Grover, M. Dhar, and S. Ermon. Flow-GAN: Combining maximum likelihood and adversarial learning in generative models. In Proc. AAAI, 2018.
  10. 10.M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In NIPS, pages 6626–6637, 2017.
  11. 11.X. Huang and S. J. Belongie. Arbitrary style transfer in real-time with adaptive instance normalization. CoRR, abs/1703.06868, 2017.
  12. 12.T. Karras, T. Aila, S. Laine, and J. Lehtinen. Progressive growing of GANs for improved quality, stability, and variation. CoRR, abs/1710.10196, 2017.
  13. 13.T. Karras, S. Laine, and T. Aila. A style-based generator architecture for generative adversarial networks. In Proc. CVPR, 2019.
  14. 14.D. P. Kingma and P. Dhariwal. Glow: Generative flow with invertible 1x1 convolutions. CoRR, abs/1807.03039, 2018.
  15. 15.D. P. Kingma, D. J. Rezende, S. Mohamed, and M. Welling. Semi-supervised learning with deep generative models. In Proc. NIPS, 2014.
  16. 16.Z. Lin, A. Khetan, G. Fanti, and S. Oh. PacGAN: The power of two samples in generative adversarial networks. CoRR, abs/1712.04086, 2017.
  17. 17.D. Lopez-Paz and M. Oquab. Revisiting classifier two-sample tests. In Proc. ICLR, 2017.
  18. 18.M. Marchesi. Megapixel size image creation using generative adversarial networks. CoRR, abs/1706.00082, 2017.
  19. 19.L. Mescheder, A. Geiger, and S. Nowozin. Which training methods for GANs do actually converge? CoRR, abs/1801.04406, 2018.
  20. 20.L. Metz, B. Poole, D. Pfau, and J. Sohl-Dickstein. Unrolled generative adversarial networks. CoRR, abs/1611.02163, 2016.
  21. 21.T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida. Spectral normalization for generative adversarial networks. CoRR, abs/1802.05957, 2018.
  22. 22.T. Miyato and M. Koyama. cGANs with projection discriminator. CoRR, abs/1802.05637, 2018.
  23. 23.E. Nalisnick, A. Matsukawa, Y. W. Teh, D. Gorur, and B. Lakshminarayanan. Do deep generative models know what they don’t know? In Proc. ICLR, 2019.
  24. 24.A. Odena, C. Olah, and J. Shlens. Conditional image synthesis with auxiliary classifier GANs. In ICML, 2017.
  25. 25.M. S. M. Sajjadi, O. Bachem, M. Lucic, O. Bousquet, and S. Gelly. Assessing generative models via precision and recall. CoRR, abs/1806.00035, 2018.
  26. 26.T. Salimans, I. J. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen. Improved techniques for training GANs. In NIPS, 2016.
  27. 27.L. Simon, R. Webster, and J. Rabin. Revisiting precision and recall definition for generative model evaluation. CoRR, abs/1905.05441, 2019.
  28. 28.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. CoRR, abs/1409.1556, 2014.
  29. 29.I. Tolstikhin, O. Bousquet, S. Gelly, and B. Schoelkopf. Wasserstein auto-encoders. In Proc. ICLR, 2018.
  30. 30.A. van den Oord, N. Kalchbrenner, and K. Kavukcuoglu. Pixel recurrent neural networks. In ICML, pages 1747–1756, 2016.
  31. 31.A. van den Oord, N. Kalchbrenner, O. Vinyals, L. Espeholt, A. Graves, and K. Kavukcuoglu. Conditional image generation with PixelCNN decoders. CoRR, abs/1606.05328, 2016.
  32. 32.H. Zhang, I. Goodfellow, D. Metaxas, and A. Odena. Self-attention generative adversarial networks. CoRR, abs/1805.08318, 2018.
  33. 33.R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proc. CVPR, 2018.
  34. 34.J. Zhu, T. Park, P. Isola, and A. A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. CoRR, abs/1703.10593, 2017.

Citation

MLA
Kynkäänniemi, T., et al. “Improved Precision and Recall Metric for Assessing Generative Models”. arXiv, 2019, http://arxiv.org/abs/1904.06991v3.
APA
Kynkäänniemi, T., Karras, T., Laine, S., Lehtinen, J., & Aila, T. (2019). Improved Precision and Recall Metric for Assessing Generative Models. arXiv. http://arxiv.org/abs/1904.06991v3
Chicago
Kynkäänniemi, T., T. Karras, S. Laine, J. Lehtinen, and T. Aila. 2019. “Improved Precision and Recall Metric for Assessing Generative Models”. arXiv. http://arxiv.org/abs/1904.06991v3.
Harvard
Kynkäänniemi, T. et al. (2019) “Improved Precision and Recall Metric for Assessing Generative Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1904.06991v3.
Vancouver
1. Kynkäänniemi T, Karras T, Laine S, Lehtinen J, Aila T (2019) Improved Precision and Recall Metric for Assessing Generative Models. arXiv

BibTeX

@article{kynkaanniemi2019improved,
  title = {Improved Precision and Recall Metric for Assessing Generative Models},
  author = {Kynkäänniemi, Tuomas and Karras, Tero and Laine, Samuli and Lehtinen, Jaakko and Aila, Timo},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1904.06991v3},
  eprint = {1904.06991}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors