Diffusion Models Encode the Intrinsic Dimension of Data Manifolds

Jan StanczukGeorgios BatzolisTeo DeveneyCarola-Bibiane Schönlieb

article2024ICML74 citations

Proves that diffusion models approximate the normal bundles of data distributions at low noise levels and presents the first diffusion-based method to estimate the intrinsic dimensionality of high-dimensional datasets.

Listen

Modern data analysis and artificial intelligence frequently handle high-dimensional observations, such as complex imagery, where the number of measurable features far exceeds the actual degrees of freedom. Under the manifold hypothesis, such data primarily concentrates along a lower-dimensional underlying structure known as the data manifold. Correctly identifying this intrinsic dimension is essential for determining model capacity, optimizing compression architectures, understanding sample efficiency, and reducing computational costs. However, classical statistical estimators often fail in high-dimensional settings, while recent invertible deep learning approaches suffer from severe architectural and numerical stability trade-offs.

The article demonstrates that trained diffusion models implicitly capture data manifolds and introduces a novel framework to accurately estimate their intrinsic dimension. The core approach leverages the theoretical insight that diffusion models learn the score function—the gradient of the perturbed data distribution's log-density—which points perpendicularly toward the underlying manifold at low noise levels. By perturbing test points with slight noise, sampling the resulting model vectors, and performing singular value decomposition, the intrinsic dimension is extracted by identifying the transition point where singular values drop sharply, separating the perpendicular normal space from the tangential space. The method was evaluated across synthetic Euclidean manifolds, synthetic image datasets with known dimensions up to 100, and the standard MNIST handwritten digit dataset.

The evaluation revealed several key findings. First, the diffusion-based estimator reliably recovered the exact or near-exact intrinsic dimension across synthetic benchmarks, maintaining high accuracy on complex, non-linear structures such as a 100-dimensional Gaussian blob manifold (estimating 98 dimensions against a ground truth of 100) where normalizing flow methods degraded significantly (estimating 56.3) and linear methods failed completely (estimating 985). Second, traditional estimators like maximum likelihood and local principal component analysis consistently and severely underestimated manifold dimensionality, capturing only around 35–40 dimensions on a 100-dimensional target. Third, the method accurately separated composite datasets, identifying the distinct dimensions of a union of spheres (10 and 31 dimensions) based on local evaluation points. Finally, on the MNIST dataset, the method estimated an intrinsic dimension of 152, varying by digit complexity from 66 for digit '1' to 152 for digit '9'. These estimates closely align with the point of diminishing returns in auto-encoder reconstruction error, confirming that historical estimates below 15 dimensions derived from classical tools significantly understate data complexity.

These findings indicate that diffusion architectures can be deployed as robust, dual-purpose tools for generative modeling and geometric diagnostics without requiring specialized, fragile architectures. For machine learning deployment and resource planning, understanding the true intrinsic dimension prevents underfitting from undersized latent representations and explains why high-capacity models require specific sample sizes for generalization. Furthermore, the systematic underestimation by traditional methods suggests that benchmark complexity figures across current research literature should be re-evaluated.

Organizations should adopt diffusion-based dimension estimation when calibrating latent dimensions for generative pipelines, auto-encoders, and representation learning models. When applying this approach, practitioners should evaluate multiple localized points and select the maximum estimated dimension to mitigate potential score approximation noise. Before relying heavily on these metrics for new operational domains, further empirical testing is warranted on diverse data types, such as audio, text embeddings, and scientific data, alongside monitoring for extreme non-uniformity or high ambient noise where structural boundaries blur.

No sufficiently relevant recommendations were found.

Cover for Diffusion Models Encode the Intrinsic Dimension of Data Manifolds

Abstract

In this work, we provide a mathematical proof that diffusion models encode data manifolds by approximating their normal bundles. Based on this observation we propose a novel method for extracting the intrinsic dimension of the data manifold from a trained diffusion model. Our insights are based on the fact that a diffusion model approximates the score function i.e. the gradient of the log density of a noise-corrupted version of the target distribution for varying levels of corruption. We prove that as the level of corruption decreases, the score function points towards the manifold, as this direction becomes the direction of maximal likelihood increase. Therefore, at low noise levels, the diffusion model provides us with an approximation of the manifold's normal bundle, allowing for an estimation of the manifold's intrinsic dimension. To the best of our knowledge our method is the first estimator of intrinsic dimension based on diffusion models and it outperforms well established estimators in controlled experiments on both Euclidean and image data. The code is available at https://github.com/GBATZOLIS/ID-diff.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Proposed Method for Estimation of Intrinsic Dimension
  • 4. Theoretical Analysis
  • 5. Limitations
  • 6. Experiments
  • 6.1. Experiments on Euclidean datasets
  • 6.2. Experiments on image datasets
  • 7. Conclusions and further directions
  • Impact Statement
  • Acknowledgements
  • References
  • A. Extended background on diffusion models
  • B. Training details
  • B.1. Euclidean data
  • B.2. Image data
  • B.3. Auto-encoder
  • C. Benchmarking
  • D. Proofs
  • Assumptions
  • Illustrative simple case
  • Deriving the formula for the density of Pt
  • Tubular Neighbourhoods
  • Definition D.3. Tubular Neighbourhood
  • Preliminary lemmas and Morse theory
  • Proof of Theorem D.1
  • Proof of Theorem 4.1
  • Proof of Corollary 4.2
  • E. Details on the design of synthetic image manifolds
  • F. Additional Experimental Results
  • F.1. Euclidean Data
  • F.2. Synthetic Image Data
  • F.3. MNIST
  • G. Robustness analysis
  • G.1. Robustness to score approximation error
  • G.2. Robustness to non-uniform distribution on the manifold
  • G.3. Relaxing the strict manifold assumption

Knowls

  1. Knowl 1 — Low-noise diffusion scores align with the manifold normal

    theoretical result

    Let M⊂RdM\subset\mathbb{R}^d be a compact, smoothly embedded manifold, and let the data distribution have a smooth density on MM that is bounded away from zero. Let ptp_t be the density obtained by adding isotropic Gaussian noise with variance σt2\sigma_t^2 to data from MM, where σt→0\sigma_t\to 0 as t→0t\to 0. For a point xx in a tubular neighbourhood of MM, let π(x)\pi(x) be its unique nearest-point projection onto MM, and define the unit vector n=(π(x)−x)/∥π(x)−x∥n=(\pi(x)-x)/\|\pi(x)-x\|, pointing from xx to the manifold. Under the paper’s mild assumptions, the score ∇xlog⁡pt(x)\nabla_x\log p_t(x) aligns with nn as noise vanishes: its cosine similarity with nn tends to 11. Equivalently, the norm of its tangent-space projection divided by the norm of its normal-space projection tends to 00. Thus, at sufficiently low noise, the score points predominantly toward the manifold rather than along it.

  2. Knowl 2 — Intrinsic-dimension estimator from a trained diffusion model

    algorithm

    The estimator uses a trained score network sθ(x,t)s_\theta(x,t), which approximates ∇xlog⁡pt(x)\nabla_x\log p_t(x), to estimate the dimension near a data point x0x_0. Here dd is the ambient dimension, t=ϵt=\epsilon is a small positive diffusion time, σϵ2\sigma_\epsilon^2 is the variance of the forward Gaussian perturbation at that time, and KK is the number of score vectors. The paper uses K=4dK=4d. For each x0x_0, it estimates the intrinsic dimension by counting the near-zero singular values of the matrix of sampled scores; it operationalizes the spectral cutoff by the largest adjacent singular-value drop. For a dataset, it repeats the pointwise procedure at multiple randomly chosen data points and uses the maximum pointwise estimate, which the authors find works best on more complex data.

    Input: Trained score network s_theta, data points, small time epsilon
    Set d to the ambient dimension and K to 4d
    For each selected data point x0:
        Initialize S as a matrix with d rows and no columns
        For j from 1 to K:
            Sample xt_j from the Gaussian with mean x0 and covariance sigma_epsilon^2 I
            Evaluate v_j = s_theta(xt_j, epsilon)
            Append v_j as a column of S
        Compute the singular values s_1, ..., s_d of S in descending order
        Set i_star to the index i in {1, ..., d-1} maximizing s_i - s_(i+1)
        Set k_hat(x0) = d - i_star
    Return the maximum k_hat(x0) over the selected data points
  3. Knowl 3 — Score-vector spectra reveal the normal-space dimension

    theoretical result

    Suppose the data lie on a kk-dimensional manifold in Rd\mathbb{R}^d. Around a data point x0x_0, sample nearby points by applying small isotropic Gaussian perturbations, and evaluate the score at those points. If the perturbations are sufficiently small that the nearby manifold normal spaces are approximately the same as the normal space at x0x_0, low-noise score vectors are approximately contained in that (d−k)(d-k)-dimensional normal space. With enough sampled perturbations, they span it. Consequently, the score matrix has d−kd-k relatively large singular values and kk small or nearly vanishing singular values; the latter count estimates the manifold’s intrinsic dimension. At finite noise, imperfect score approximation, nonuniform density, and curvature can make the tangential singular values nonzero, so the method identifies a spectral gap rather than requiring exact zeros.

  4. Knowl 4 — Comparison across Euclidean, image, and MNIST datasets

    data/table

    The following comparison reports the paper’s intrinsic-dimension estimates for datasets with known dimension, plus MNIST, whose true dimension is unknown. “Ours” is the diffusion-score estimator; ID-NF extracts dimension from a trained normalizing flow; MLE is reported with nearest-neighbour counts m=5m=5 and m=20m=20; Local PCA and PPCA are statistical baselines. The diffusion estimator is exact or close to the known value throughout these tests. The results also illustrate failures on complex manifolds: for the 100-dimensional Gaussian-blob images, ID-NF and Local PCA substantially underestimate, while PPCA greatly overestimates. MNIST estimates vary by method, so its row is not a ground-truth accuracy comparison.

    Dataset Ground truth Ours ID-NF MLE (m=5m=5) MLE (m=20m=20) Local PCA PPCA
    10-sphere 10 11 11 9.61 9.46 11 11
    50-sphere 50 51 51 35.52 34.04 51 51
    Spaghetti line 1 1 1 1.01 1.00 32 98
    Squares, k=10k=10 10 11 9.7 8.48 8.17 10 10
    Squares, k=20k=20 20 22 19.5 14.96 14.36 20 20
    Squares, k=100k=100 100 100 94.2 37.69 34.42 78 99
    Gaussian blobs, k=10k=10 10 12 9.8 8.88 8.67 10 136
    Gaussian blobs, k=20k=20 20 21 17.8 16.34 15.75 20 264
    Gaussian blobs, k=100k=100 100 98 56.3 39.66 35.31 18 985
    MNIST N/A 152 182 14.12 13.27 38 706
  5. Knowl 5 — Pointwise estimates handle nonlinear and mixed-dimensional Euclidean data

    empirical result

    The paper tests local score-spectrum estimation on Euclidean manifolds embedded in an ambient dimension of d=100d=100. A “spaghetti line,” parameterized by τ↦(sin⁡τ,sin⁡(2τ),…,sin⁡(100τ))\tau\mapsto(\sin\tau,\sin(2\tau),\ldots,\sin(100\tau)), is a one-dimensional curve not contained in a low-dimensional linear subspace. The diffusion method estimates its dimension as 11, whereas PPCA estimates 9898, illustrating the advantage of a nonlinear, local estimator over a linear-subspace approach. The authors also form a union of two spheres, of dimensions 1010 and 3030 and radii 11 and 0.250.25, respectively. Applying the estimator at different data points gives estimates of 1010 and 3131, depending on the selected component and point. The separated spectral drops indicate that pointwise estimates can distinguish components of different dimensions in a union of manifolds.

  6. Knowl 6 — Synthetic image manifolds separate linear and nonlinear structure

    experimental setup

    The paper constructs two families of synthetic 32×3232\times32 grayscale images, each with ambient dimension 10241024 and controllable intrinsic dimension kk. In the squares family, square centres and side lengths (3 or 5 pixels) are fixed across images; each square receives a brightness sampled uniformly from [0,1][0,1], and overlapping contributions are summed. The kk independently varying square brightnesses give a kk-dimensional image manifold. In the Gaussian-blobs family, blob centres are fixed and each blob’s standard deviation is sampled uniformly from [1,5][1,5], giving dimension kk through those kk varying parameters. Unlike the squares manifold, the Gaussian-blobs manifold is not contained in a low-dimensional linear subspace. The experiments evaluate both families at k=10,20,100k=10,20,100; the comparison results show PPCA performs well on squares but can fail severely on Gaussian blobs, while the diffusion method remains close to the target dimension.

  7. Knowl 7 — MNIST estimates vary by digit and align with an autoencoder plateau

    empirical result

    For MNIST, whose intrinsic dimension is unknown, the paper reports diffusion-based estimates separately for each digit. Each reported estimate is the maximum pointwise estimate among spectra evaluated at 500 instances of that digit: digit 0: 113; 1: 66; 2: 131; 3: 120; 4: 107; 5: 129; 6: 126; 7: 100; 8: 148; 9: 152. The variation is interpreted as evidence that different digits have different geometric complexity; for example, the estimate for ‘9’ is higher than for ‘1’. As a separate check, the authors compare estimates with reconstruction error from autoencoders trained at different latent dimensions. The diffusion estimate of 152 and the ID-NF estimate of 182 are close to the latent-dimension region where reconstruction improvements begin to plateau; MLE and Local PCA estimates of 14.12, 13.27, and 38, respectively, lie where reconstruction error is still falling steeply. The paper treats these as indirect validation, not as ground-truth dimensions, and notes that its MNIST estimates may be slight overestimates or lower bounds on the true dimensions.

  8. Knowl 8 — Maximum pointwise estimates improve robustness to nonuniform sampling

    data/table

    The authors test a 10-sphere embedded in R100\mathbb{R}^{100} under increasingly concentrated sampling. They generate nonuniform points by sampling the sphere’s k−1k-1 angles from a Gaussian with covariance αI\alpha I; smaller α\alpha means greater concentration. They estimate dimension at 1,000 sampled points and use the maximum pointwise estimate for the diffusion method. This rule retains the correct estimate at α=1\alpha=1 and 0.750.75, while at the strongest tested concentration, α=0.5\alpha=0.5, the estimate falls to 7. The table gives all reported estimates; the competing MLE and Local PCA estimates decrease substantially as sampling becomes nonuniform, whereas PPCA remains at 11.

    Method Uniform α=1\alpha=1 α=0.75\alpha=0.75 α=0.5\alpha=0.5
    Ground truth 10 10 10 10
    Ours 11 10 10 7
    MLE (m=5m=5) 9.61 5.37 4.83 4.12
    MLE (m=20m=20) 9.46 4.99 4.49 3.82
    Local PCA 11 5 4 3
    PPCA 11 11 11 11
  9. Knowl 9 — Score error and data off the manifold limit practical reliability

    limitation

    The theoretical result assumes that the data distribution is supported on a manifold and that the score is accurately approximated at sufficiently low noise. In practice, two error sources can impair the estimate: score-approximation error, and geometric error when the sampling noise is too large. Larger noise can increase tangential score components and make the nearby normal spaces differ more because of curvature. In an empirical test on a 25-sphere, the authors add Gaussian noise to the score outputs and measure its strength as the expected noise-vector norm divided by the expected score-vector norm. The spectrum still has a visible drop at the correct dimension when this ratio is below 0.50.5; as the ratio increases, tangential singular values rise and the gap diminishes. In a separate test, they train on data blurred away from a 25-sphere by Gaussian noise: the method estimates the correct dimension for small blur, but the paper gives no numerical blur threshold. The authors therefore limit the off-manifold claim to distributions that remain tightly concentrated around a manifold.

  10. Knowl 10 — Diffusion-model training configurations used in the experiments

    experimental setup

    The score networks are trained with weighted denoising score matching and likelihood weighting, meaning the time-dependent loss weight is λ(t)=g(t)2\lambda(t)=g(t)^2, where g(t)g(t) is the diffusion coefficient of the forward stochastic differential equation. Euclidean experiments use a variance-exploding process with σmin⁡=0.01\sigma_{\min}=0.01 and σmax⁡=4\sigma_{\max}=4, and a fully connected score network with five hidden layers of 2,048 units each. The Euclidean models use Adam with learning rate 2×10−52\times10^{-5} and exponential moving average (EMA) decay 0.99990.9999. Image experiments use a DDPM U-Net architecture with a variance-exploding process. Their reported settings are:

    Hyperparameter MNIST Synthetic images
    Number of filters 128 128
    Channel multipliers (1,2,2,4)(1,2,2,4) (1,2,2,2)(1,2,2,2)
    Dropout 0.1 0.1
    EMA rate 0.999 0.999
    Normalization GroupNorm GroupNorm
    Nonlinearity Swish Swish
    Residual blocks 4 4
    Attention resolution 16 16
    Convolution size 3 3
    σmin⁡\sigma_{\min} 0.009 0.01
    σmax⁡\sigma_{\max} 50 50
    Learning rate Scheduler (10−4,10−5)(10^{-4},10^{-5}) 2×10−42\times10^{-4}

Coverage note — Detailed proof steps and individual score-spectrum plots are omitted because they support, but do not add standalone findings beyond, the geometric result and summarized experiments.

References

  1. 1.Brian D.O. Anderson. Reverse-time diffusion equation models. Stochastic Processes and their Applications, 12(3):313–326, 1982. ISSN 0304-4149. doi: https://doi.org/10.1016/0304-4149(82)90051-5. URL https://www.sciencedirect.com/science/article/pii/0304414982900515.
  2. 2.Jens Behrmann, Paul Vicol, Kuan-Chieh Wang, Roger Grosse, and Jörn-Henrik Jacobsen. Understanding and mitigating exploding inverses in invertible neural networks. In International Conference on Artificial Intelligence and Statistics, pages 1792–1800. PMLR, 2021.
  3. 3.C M Bishop and M E Tipping. Probabilistic principal component analysis. 2001. PPCA.
  4. 4.Johann Brehmer and Kyle Cranmer. Flows for simultaneous manifold learning and density estimation. Advances in Neural Information Processing Systems, 33:442–453, 2020.
  5. 5.F. Camastra and A. Vinciarelli. Estimating the intrinsic dimension of data with a fractal-based method. IEEE Transactions on Pattern Analysis and Machine Intelligence, 24(10):1404–1407, 2002. doi: 10.1109/TPAMI.2002.1039212.
  6. 6.Paola Campadelli, Elena Casiraghi, Claudio Ceruti, and Alessandro Rozza. Intrinsic dimension estimation: Relevant techniques and a benchmark framework. Mathematical Problems in Engineering, 2015:1–21, 2015.
  7. 7.Minshuo Chen, Kaixuan Huang, Tuo Zhao, and Mengdi Wang. Score approximation, estimation and distribution recovery of diffusion models on low-dimensional data, 2023.
  8. 8.Rob Cornish, Anthony Caterini, George Deligiannidis, and Arnaud Doucet. Relaxing bijectivity constraints with continuously indexed normalising flows. In International conference on machine learning, pages 2133–2143. PMLR, 2020.
  9. 9.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009. doi: 10.1109/CVPR.2009.5206848.
  10. 10.Mingyu Fan, Nannan Gu, Hong Qiao, and Bo Zhang. Intrinsic dimension estimation of data by principal component analysis, 2010. URL https://arxiv.org/abs/1002.2050.
  11. 11.Charles Fefferman, Sanjoy Mitter, and Hariharan Narayanan. Testing the manifold hypothesis, 2013. URL https://arxiv.org/abs/1310.0425.
  12. 12.K. Fukunaga and D.R. Olsen. An algorithm for finding intrinsic dimensionality of data. IEEE Transactions on Computers, C-20(2):176–183, 1971. doi: 10.1109/T-C.1971.223208.
  13. 13.Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks, 2014. URL https://arxiv.org/abs/1406.2661.
  14. 14.Anirudh Goyal and Yoshua Bengio. Inductive biases for deep learning of higher-level cognition. Proceedings of the Royal Society A, 478(2266):20210068, 2022.
  15. 15.Gloria Haro, Gregory Randall, and Guillermo Sapiro. Translated poisson mixture model for stratification learning. Int. J. Comput. Vis., 80(3):358–374, 2008.
  16. 16.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models, 2020a. URL https://arxiv.org/abs/2006.11239.
  17. 17.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020b.
  18. 18.Christian Horvat and Jean-Pascal Pfister. Intrinsic dimensionality estimation using normalizing flows. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 12225–12236. Curran Associates, Inc., 2022.
  19. 19.Aapo Hyvärinen. Estimation of non-normalized statistical models by score matching. Journal of Machine Learning Research, 6(24):695–709, 2005. URL http://jmlr.org/papers/v6/hyvarinen05a.html.
  20. 20.Priyank Jaini, Ivan Kobyzev, Yaoliang Yu, and Marcus Brubaker. Tails of lipschitz triangular flows. In International Conference on Machine Learning, pages 4673–4681. PMLR, 2020.
  21. 21.Kerstin Johnsson. intrinsicdimension: Intrinsic dimension estimation, 2016. URL https://cran.r-project.org/web/packages/intrinsicDimension/.
  22. 22.Balázs Kégl. Intrinsic dimension estimation using packing numbers. In S. Becker, S. Thrun, and K. Obermayer, editors, Advances in Neural Information Processing Systems, volume 15. MIT Press, 2002. URL https://proceedings.neurips.cc/paper/2002/file/1177967c7957072da3dc1db4ceb30e7a-Paper.pdf.
  23. 23.Jisu Kim, Jaehyeok Shin, Alessandro Rinaldo, and Larry Wasserman. Uniform convergence rate of the kernel density estimator adaptive to intrinsic volume dimension. In International Conference on Machine Learning, pages 3398–3407. PMLR, 2019.
  24. 24.Diederik P Kingma and Max Welling. Auto-encoding variational bayes, 2013. URL https://arxiv.org/abs/1312.6114.
  25. 25.Samory Kpotufe. k-nn regression adapts to local intrinsic dimension. Advances in neural information processing systems, 24, 2011.
  26. 26.Alex Krizhevsky. Learning multiple layers of features from tiny images. University of Toronto, 05 2012.
  27. 27.Mike Laszkiewicz, Johannes Lederer, and Asja Fischer. Copula-based normalizing flows. arXiv preprint arXiv:2107.07352, 2021.
  28. 28.Yann LeCun and Corinna Cortes. MNIST handwritten digit database. 2010. URL http://yann.lecun.com/exdb/mnist/.
  29. 29.J.M. Lee. Introduction to Riemannian Manifolds. Graduate Texts in Mathematics. Springer International Publishing, 2019. ISBN 9783319917542. URL https://books.google.co.uk/books?id=UIPltQEACAAJ.
  30. 30.Elizaveta Levina and Peter Bickel. Maximum likelihood estimation of intrinsic dimension. In L. Saul, Y. Weiss, and L. Bottou, editors, Advances in Neural Information Processing Systems, volume 17. MIT Press, 2004. URL https://proceedings.neurips.cc/paper/2004/file/74934548253bcab8490ebd74afed7031-Paper.pdf.
  31. 31.Thomas Minka. Automatic choice of dimensionality for pca. In T. Leen, T. Dietterich, and V. Tresp, editors, Advances in Neural Information Processing Systems, volume 13. MIT Press, 2000. URL https://proceedings.neurips.cc/paper/2000/file/7503cfacd12053d309b6bed5c89de212-Paper.pdf.
  32. 32.L. Nicolaescu. An Invitation to Morse Theory. Universitext. Springer New York, 2011. ISBN 9781461411055. URL https://books.google.co.uk/books?id=nCgvt2MY4QAC.
  33. 33.Kazusato Oko, Shunta Akiyama, and Taiji Suzuki. Diffusion models are minimax optimal distribution estimators, 2023.
  34. 34.Richard S. Palais and C. Terng. Critical Point Theory and Submanifold Geometry. Critical Point Theory and Submanifold Geometry. Springer, 1988. ISBN 9783540503996. URL https://books.google.co.uk/books?id=ViSHzQEACAAJ.
  35. 35.Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al. Scikit-learn: Machine learning in python. Journal of machine learning research, 12(Oct):2825–2830, 2011.
  36. 36.Karl W. Pettis, Thomas A. Bailey, Anil K. Jain, and Richard C. Dubes. An intrinsic dimensionality estimator from near-neighbor information. IEEE Transactions on Pattern Analysis and Machine Intelligence, PAMI-1(1):25–37, 1979. doi: 10.1109/TPAMI.1979.4766873.
  37. 37.Jakiw Pidstrigach. Score-based generative models detect manifolds, 2022.
  38. 38.Phillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum, and Tom Goldstein. The intrinsic dimension of images and its impact on learning. arXiv preprint arXiv:2104.08894, 2021.
  39. 39.Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics, 2015. URL https://arxiv.org/abs/1503.03585.
  40. 40.Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020.
  41. 41.Yang Song, Conor Durkan, Iain Murray, and Stefano Ermon. Maximum likelihood training of score-based diffusion models. Advances in Neural Information Processing Systems, 34:1415–1428, 2021.
  42. 42.Piotr Tempczyk, Rafał Michaluk, Łukasz Garncarek, Przemysław Spurek, Jacek Tabor, and Adam Golinski. Lidl: ´ Local intrinsic dimension estimation using approximate likelihood, 2022.
  43. 43.Pascal Vincent. A connection between score matching and denoising autoencoders. Neural Computation, 23(7):1661–1674, 2011. doi: 10.1162/NECO_a_00142.
  44. 44.Jonathan Weed and Francis Bach. Sharp asymptotic and finite-sample rates of convergence of empirical measures in wasserstein distance. 2019.

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/