KerJEPA: Kernel Discrepancies for Euclidean Self-Supervised Learning

Eric ZimmermannEric ZimmermannHarley WiltzerJustin SzetoDavid Alvarez-MelisLester Mackey

article2025arXiv4 citations

Proposes a flexible family of kernel-regularized Joint-Embedding Predictive Architectures that uses closed-form limits of sliced maximum mean discrepancies to improve training stability and design flexibility in self-supervised learning.

Listen

Self-supervised learning enables artificial intelligence models to learn useful visual features directly from unlabeled data, forming the backbone of modern computer vision systems. A persistent challenge in this approach is avoiding representation collapse, where a network generates identical outputs for all inputs. While recent methods prevent collapse by regularizing representations toward standard target distributions using random one-dimensional projections, this projection technique introduces significant estimation noise and restricts regularizers to limited geometric assumptions.

The article aims to introduce a unified framework called KerJEPA that expands the design space for self-supervised regularization using broad families of kernel discrepancies. It evaluates alternative statistical distance measures, non-Gaussian target distributions, and analytic high-dimensional limits to enhance training stability and flexibility.

The authors conducted a theoretical and empirical study comparing multiple regularization objectives across varied representation dimensions and training schedules. They evaluated the framework on the standard ImageNette benchmark using a Vision Transformer backbone trained for up to 800 epochs. The study analyzed Maximum Mean Discrepancy, which measures distributional distances in high-dimensional feature spaces, and Kernel Stein Discrepancy, which evaluates alignment using only the score function of the target distribution without requiring sampling from the target prior.

The investigation produced several key findings. First, the standard projection-based regularization method was proven mathematically equivalent to high-dimensional Maximum Mean Discrepancy using a heavy-tailed kernel, allowing the exact high-dimensional limit to be computed directly without random projections. Second, unsliced and analytically sliced objectives significantly improved training stability and accelerated early convergence, outperforming finite-slice approximations across output dimensions ranging from 16 to 1024. Third, Kernel Stein Discrepancy using an Inverse Multiquadric kernel achieved the highest overall top-1 accuracy at approximately 91.90%, outperforming the standard baseline of 91.13%. Fourth, finite-slice approximations exhibited rising gradient variance as embedding dimensions increased, causing noticeable training instability unless large numbers of projection slices were computed.

These findings indicate that teams developing self-supervised foundation models can avoid the optimization noise and instability of random projections by adopting closed-form, analytically sliced, or unsliced kernel regularizers. Because exact and analytically derived regularizers converge faster in earlier training stages, they can reduce computational risk and total training time. Furthermore, the decoupling between projector heads and inference backbones suggests that downstream model performance is primarily driven by smooth Euclidean gradient dynamics rather than the exact choice between Gaussian or non-Gaussian target priors.

Practitioners should select regularization strategies based on batch size and dimensional constraints. For typical batch sizes where the number of samples is smaller than the required number of random projections, teams should deploy analytically sliced or unsliced Kernel Stein Discrepancy estimators to ensure stable optimization. When scaling to massive batch sizes across distributed clusters, engineering teams should evaluate efficiency tools such as random Fourier features or coreset approximations to balance computation costs.

The primary limitation of the study is its empirical reliance on ImageNette, a relatively small visual classification benchmark. While the theoretical derivations provide high confidence in the statistical mechanics of the proposed discrepancies, broader validation on large-scale datasets, downstream object detection, and semantic segmentation tasks is recommended before large-scale production deployment.

arXiv: 2512.19605
  • Paper: Demystifying MMD GANs, Mikołaj Bińkowski et al. (2018). Its treatment of kernel-based maximum mean discrepancy gives useful grounding for KerJEPA’s sliced-MMD regularizers and their role in training objectives.

No sufficiently relevant recommendations were found.

Cover for KerJEPA: Kernel Discrepancies for Euclidean Self-Supervised Learning

Abstract

Recent breakthroughs in self-supervised Joint-Embedding Predictive Architectures (JEPAs) have established that regularizing Euclidean representations toward isotropic Gaussian priors yields provable gains in training stability and downstream generalization. We introduce a new, flexible family of KerJEPAs, self-supervised learning algorithms with kernel-based regularizers. One instance of this family corresponds to the recently-introduced LeJEPA Epps-Pulley regularizer which approximates a sliced maximum mean discrepancy (MMD) with a Gaussian prior and Gaussian kernel. By expanding the class of viable kernels and priors and computing the closed-form high-dimensional limit of sliced MMDs, we develop alternative KerJEPAs with a number of favorable properties including improved training stability and design flexibility.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Matching Distributions with Kernel Discrepancies
  • 3.1 Kernels, Reproducing Kernel Hilbert Spaces, and Distribution Comparison
  • 3.2 Maximum Mean Discrepancy
  • 3.2.1 Epps-Pulley via Maximum Mean Discrepancy
  • 3.3 Kernel Stein Discrepancy
  • 3.3.1 Spectral Representation of the Kernel Stein Discrepancy
  • 3.4 Sliced Integral Probability Metrics
  • 3.5 Tradeoffs with Sliced Metrics
  • 3.6 Sliced Kernel Stein Discrepancy
  • 4 Discrepancy Regularization in Self-Supervised Learning
  • 4.1 Sliced Discrepancy Regularization
  • 4.2 Generalized Discrepancy Regularization
  • 5 Experimental Results
  • 6 Discussion
  • 7 Limitations and Future Work
  • References
  • A Shift-Invariant Kernels
  • B Isotropic Prior Distributions
  • C Maximum Mean Discrepancy
  • D Kernel Stein Discrepancy
  • D.1 Spectral Representations of KSD
  • E Sliced Maximum Mean Discrepancies without Slicing with an Isotropic Gaussian Prior
  • F Sliced Kernel Stein Discrepancies without Slicing with an Isotropic Gaussian Prior
  • G Hypergeometric Integrals
  • H Learning with Hyperspherical Representations
  • H.1 Maximum Mean Discrepancy with Uniform Prior
  • H.2 Kernel Stein Discrepancy as a Uniformity Regularizer
  • I Extended Results

Knowls

  1. Knowl 1 — KerJEPA regularizes view-aligned Euclidean embeddings with a distribution discrepancy

    model/method

    KerJEPA combines a view-alignment objective with a regularizer that matches a batch of projected embeddings to a chosen isotropic prior. Given nn source images, let Pl\mathcal P_l be the set of positive view pairs for image ll, and let zli=fθ(vli)∈Rdz_l^i=f_\theta(v_l^i)\in\mathbb R^d be the embedding of its iith augmented view. The encoder fθf_\theta is the composition of a backbone and a projector. For the batch of embeddings ZZ, the objective is

    L(θ)=1n∣Pl∣∑l=1n∑(i,j)∈PlLalign(zli,zlj)+λ Ω(Z;Q),\mathcal L(\theta)=\frac{1}{n|\mathcal P_l|}\sum_{l=1}^n\sum_{(i,j)\in\mathcal P_l}\mathcal L_{\mathrm{align}}(z_l^i,z_l^j)+\lambda\,\Omega(Z;Q),

    where QQ is an isotropic target prior, λ\lambda weights the regularizer, and Ω\Omega can be an MMD-based or KSD-based discrepancy. The paper uses squared Euclidean distance, Lalign(z,z′)=∥z−z′∥22\mathcal L_{\mathrm{align}}(z,z')=\|z-z'\|_2^2, for alignment. The discrepancy regularizer supplies distributional structure to the projector outputs alongside view alignment, which by itself can collapse embeddings.

  2. Knowl 2 — Kernel Stein discrepancy enables prior matching through target scores

    model/method

    For a target distribution QQ with differentiable log-density and score sQ(x)=∇xlog⁡q(x)s_Q(x)=\nabla_x\log q(x), the squared Kernel Stein Discrepancy (KSD) can be computed using only samples from the embedding distribution PP and the score of QQ. For a positive-definite scalar kernel k:Rd×Rd→Rk:\mathbb R^d\times\mathbb R^d\to\mathbb R, define

    kQ(x,y)=sQ(x)⊤k(x,y)sQ(y)+sQ(x)⊤∇yk(x,y)+∇xk(x,y)⊤sQ(y)+tr⁡(∇x∇y⊤k(x,y)).k_{Q}(x,y)=s_Q(x)^\top k(x,y)s_Q(y)+s_Q(x)^\top\nabla_y k(x,y)+\nabla_x k(x,y)^\top s_Q(y)+\operatorname{tr}(\nabla_x\nabla_y^\top k(x,y)).

    Then KSD⁡k2(P,Q)=EX,X′∼P[kQ(X,X′)]\operatorname{KSD}_k^2(P,Q)=\mathbb E_{X,X'\sim P}[k_Q(X,X')], with independent dd-dimensional X,X′X,X'. This form does not require sampling from the prior or knowing its normalization constant. The paper considers Gaussian and inverse-multiquadric kernels with isotropic Gaussian, Laplace, and Student-tt priors, for which it reports tractable KSD expressions. It presents KSD as a way to choose target geometries through the score function rather than relying on the Gaussian-prior, Gaussian-kernel combination.

  3. Knowl 3 — SIGReg is an MMD with a dimension-dependent heavy-tailed kernel

    theoretical result

    Let PP be a distribution on Rd\mathbb R^d, let ξ\xi be uniform on the unit sphere Sd−1S^{d-1}, and let ξ⊤ ⁣#P\xi^\top\!\#P denote the distribution of the scalar projection ξ⊤X\xi^\top X for X∼PX\sim P. The Epps–Pulley discrepancy from a univariate Gaussian target N(0,σ2)N(0,\sigma^2), when weighted by the spectral density of the Gaussian kernel with bandwidth γ>0\gamma>0, is the squared MMD for that kernel. Averaging this univariate discrepancy over directions gives the SIGReg objective, which is exactly a multivariate MMD:

    SIGReg⁡(P)=MMD⁡κd2(P,N(0,σ2Id)),κd(x,y)=1F1 ⁣(12;d2;−γ∥x−y∥22).\operatorname{SIGReg}(P)=\operatorname{MMD}_{\kappa_d}^2\bigl(P,N(0,\sigma^2 I_d)\bigr),\qquad \kappa_d(x,y)={}_1F_1\!\left(\frac12;\frac d2;-\gamma\|x-y\|_2^2\right).

    Here IdI_d is the dd-dimensional identity matrix and 1F1{}_1F_1 is the confluent hypergeometric function. For d>1d>1, the induced kernel has polynomially decaying tails, unlike the exponentially decaying Gaussian kernel. Its geometry also depends on dd: the equivalent kernel changes with embedding dimension. Thus, sliced Gaussian-kernel matching does not impose the original Gaussian kernel directly in the ambient feature space.

  4. Knowl 4 — An unsliced U-statistic estimates SIGReg without random directions

    theoretical result

    The MMD representation of SIGReg yields an estimator that integrates over projection directions analytically. Let X1,…,XnX_1,\ldots,X_n be independent samples from P⊆RdP\subseteq\mathbb R^d, let κd\kappa_d be the kernel defined by 1F1(1/2;d/2;−γ∥x−y∥22){}_1F_1(1/2;d/2;-\gamma\|x-y\|_2^2), and let C=E[κd(Y,Y′)]C=\mathbb E[\kappa_d(Y,Y')] for independent Y,Y′∼N(0,σ2Id)Y,Y'\sim N(0,\sigma^2 I_d). The estimator

    S^n=1n(n−1)∑i≠jκd(Xi,Xj)−2n1+2γσ2∑i=1n1F1 ⁣(12;d2;−γ∥Xi∥221+2γσ2)+C\widehat{S}_n= \frac{1}{n(n-1)}\sum_{i\ne j}\kappa_d(X_i,X_j) -\frac{2}{n\sqrt{1+2\gamma\sigma^2}}\sum_{i=1}^n{}_1F_1\!\left(\frac12;\frac d2;-\frac{\gamma\|X_i\|_2^2}{1+2\gamma\sigma^2}\right)+C

    is an unbiased estimate of SIGReg⁡(P)\operatorname{SIGReg}(P). The constant CC depends on the target and kernel parameters, not on PP. The expected difference between this estimator and the empirical sliced SIGReg statistic is bounded by O(1/n)O(1/n). This removes Monte Carlo variance from sampling directions, but the pairwise sum costs O(n2d)O(n^2d) operations; the hypergeometric evaluations can also be expensive to approximate.

  5. Knowl 5 — The Gaussian-prior sliced KSD has a closed-form ambient-space expression

    theoretical result

    For the Gaussian base kernel k(x,y)=exp⁡(−γ∥x−y∥22)k(x,y)=\exp(-\gamma\|x-y\|_2^2) and isotropic Gaussian target N(0,σ2Id)N(0,\sigma^2 I_d), the sliced KSD can be evaluated without Monte Carlo sampling of projection directions. Let X,X′X,X' be independent samples from P⊆RdP\subseteq\mathbb R^d, with d>1d>1, γ>0\gamma>0, and σ>0\sigma>0. For X≠X′X\ne X', define η=(X−X′)/∥X−X′∥2\eta=(X-X')/\|X-X'\|_2, a=(η⊤X)(η⊤X′)a=(\eta^\top X)(\eta^\top X'), and b=X⊤X′−ab=X^\top X'-a. Then

    SKSD⁡2(P,N(0,σ2Id))=E[(2γ+bσ4(d−1))1F1 ⁣(12;d2;−γ∥X−X′∥22)]+E[1d(aσ4−bσ4(d−1))1F1 ⁣(32;d2+1;−γ∥X−X′∥22)]+E[1d(2γσ2+4γ2)∥X−X′∥221F1 ⁣(32;d2+1;−γ∥X−X′∥22)],\begin{aligned} \operatorname{SKSD}^2(P,N(0,\sigma^2I_d))={}&\mathbb E\left[\left(2\gamma+\frac{b}{\sigma^4(d-1)}\right){}_1F_1\!\left(\frac12;\frac d2;-\gamma\|X-X'\|_2^2\right)\right]\\ &+\mathbb E\left[\frac1d\left(\frac{a}{\sigma^4}-\frac{b}{\sigma^4(d-1)}\right){}_1F_1\!\left(\frac32;\frac d2+1;-\gamma\|X-X'\|_2^2\right)\right]\\ &+\mathbb E\left[\frac1d\left(\frac{2\gamma}{\sigma^2}+4\gamma^2\right)\|X-X'\|_2^2{}_1F_1\!\left(\frac32;\frac d2+1;-\gamma\|X-X'\|_2^2\right)\right], \end{aligned}

    where each expectation is over the same independent pair (X,X′)(X,X'). The expression gives the expectation over all one-dimensional projections in closed form. An empirical U-statistic uses distinct sample pairs, avoiding the undefined direction at X=X′X=X'. As with the analytic sliced MMD, evaluating the resulting pairwise terms entails quadratic cost in the number of samples.

  6. Knowl 6 — A spectral KSD formula supports characteristic-function estimation

    equation

    For a shift-invariant kernel with spectral density ρk\rho_k and a target distribution QQ with score sQ(x)=∇xlog⁡q(x)s_Q(x)=\nabla_x\log q(x), the squared KSD has the spectral representation

    KSD⁡k2(P,Q)=∫Rd∥EX∼P[(sQ(X)+iω)eiω⊤X]∥22ρk(ω) dω,\operatorname{KSD}_k^2(P,Q)=\int_{\mathbb R^d}\left\|\mathbb E_{X\sim P}\left[(s_Q(X)+i\omega)e^{i\omega^\top X}\right]\right\|_2^2\rho_k(\omega)\,d\omega,

    where ω∈Rd\omega\in\mathbb R^d is a frequency and ii is the imaginary unit. For Q=N(0,σ2Id)Q=N(0,\sigma^2I_d), with characteristic function φP(ω)=EX∼P[eiω⊤X]\varphi_P(\omega)=\mathbb E_{X\sim P}[e^{i\omega^\top X}], this becomes

    KSD⁡k2(P,N(0,σ2Id))=∫Rd∥ωφP(ω)+1σ2∇ωφP(ω)∥22ρk(ω) dω.\operatorname{KSD}_k^2(P,N(0,\sigma^2I_d))=\int_{\mathbb R^d}\left\|\omega\varphi_P(\omega)+\frac{1}{\sigma^2}\nabla_\omega\varphi_P(\omega)\right\|_2^2\rho_k(\omega)\,d\omega.

    These formulas express KSD using a characteristic-function integrand and the kernel spectrum. The paper identifies quadrature and random-feature approximations as computational options; substituting the empirical distribution for PP gives a V-estimator with bias O(1/n)O(1/n) for bounded kernels.

  7. Knowl 7 — Finite-sliced MMD and KSD regularizers use projected embeddings and quadrature

    algorithm

    The paper's finite-sliced regularizers estimate a distribution discrepancy by projecting a batch onto random one-dimensional directions and approximating the Gaussian-kernel spectral integral with Gauss–Hermite quadrature. Inputs are embeddings Z=(z1,…,zn)∈Rn×dZ=(z_1,\ldots,z_n)\in\mathbb R^{n\times d}, a Gaussian or Laplace isotropic prior with scale σ\sigma, a number rr of directions, and uu quadrature knots. Outputs are a scalar MMDReg or KSDReg loss.

    Input: Embeddings Z, prior Q in {Gaussian, Laplace}, prior scale sigma, number of directions r, Gauss-Hermite knots and weights (omega_j, a_j) for j = 1,...,u
    Output: Finite-sliced MMDReg loss and/or KSDReg loss
    Initialize MMD_total = 0 and KSD_total = 0
    For each direction l = 1,...,r:
        Sample g from the d-dimensional standard normal distribution and set theta = g / ||g||_2
        For each embedding i = 1,...,n, compute projected value t_i = theta^T z_i
        For each quadrature knot j = 1,...,u:
            Compute empirical characteristic-function real part c_j = mean_i cos(omega_j t_i)
            Compute target characteristic function phi_j = exp(-sigma^2 omega_j^2 / 2) for Gaussian Q
                or phi_j = 1 / (1 + sigma^2 omega_j^2) for Laplace Q
            Set delta_j = c_j - phi_j
            Add a_j delta_j^2 to the MMD loss for direction l
            Compute the scalar target score s_i at t_i: -t_i / sigma^2 for Gaussian Q,
                or -sign(t_i) / sigma for Laplace Q
            Set A_j = mean_i [s_i cos(omega_j t_i) - omega_j sin(omega_j t_i)]
            Set B_j = mean_i [s_i sin(omega_j t_i) + omega_j cos(omega_j t_i)]
            Add a_j (A_j^2 + B_j^2) to the KSD loss for direction l
        Add the direction-specific losses to MMD_total and KSD_total
    Return MMD_total / r and KSD_total / r

    The paper's ImageNette experiments use r=1024r=1024 directions and u=21u=21 quadrature points. It gives complexity Θ(nr(d+u))\Theta(nr(d+u)) for the projected, quadrature-based computation, compared with Θ(n2d)\Theta(n^2d) for ambient-space pairwise kernel evaluations. Finite direction sampling introduces variance; quadrature approximates the spectral integral.

  8. Knowl 8 — ImageNette comparisons show similar accuracy across many regularizers

    data/table

    The regularizers were compared on ImageNette using a ViT-s/8 trained for 800 epochs on one 80GB A100 GPU with AdamW in BF16, cosine scheduling, peak learning rate 0.00050.0005, weight decay 0.050.05, batch size 256, four views, and a 128-dimensional projector. The augmentation pipeline used random resize crops to 128 pixels, horizontal flips, color jitter, grayscale, blur, and solarization. The study grid-searched kernel bandwidth and regularization weight and evaluated an online instance-normalized linear probe. Finite-sliced rows use 1024 directions and 21 quadrature points; analytic-sliced rows use infinitely many directions in the closed-form limit; unsliced rows compute in Rd\mathbb R^d.

    Algorithm Type Base kernel Prior Slices Knots Accuracy (%) ±\pm std. err.
    LeJEPA Sliced (Finite) Gaussian Gaussian 1024 21 91.13±0.4591.13 \pm 0.45
    MMD Unsliced Gaussian Gaussian – – 91.29±0.4591.29 \pm 0.45
    MMD Sliced (Analytic) Gaussian Gaussian ∞\infty – 91.13±0.4591.13 \pm 0.45
    MMD Sliced (Finite) Gaussian Laplace 1024 21 90.25±0.4790.25 \pm 0.47
    KSD Unsliced Gaussian Gaussian – – 91.31±0.4591.31 \pm 0.45
    KSD Unsliced Gaussian Laplace – – 91.18±0.4591.18 \pm 0.45
    KSD Unsliced IMQ Gaussian – – 91.90±0.4491.90 \pm 0.44
    KSD Unsliced IMQ Laplace – – 91.12±0.4591.12 \pm 0.45
    KSD Sliced (Analytic) Gaussian Gaussian ∞\infty – 91.11±0.4591.11 \pm 0.45
    KSD Sliced (Finite) Gaussian Gaussian 1024 21 91.36±0.4591.36 \pm 0.45
    KSD Sliced (Finite) Gaussian Laplace 1024 21 90.70±0.4690.70 \pm 0.46

    The tested MMD and KSD variants perform similarly overall. The highest reported accuracy is 91.90±0.44%91.90 \pm 0.44\% for unsliced KSD with an IMQ kernel and Gaussian prior. The results also show that a heavy-tailed kernel is not necessary for competitive performance in this setup.

  9. Knowl 9 — Finite-direction slicing is less stable at larger embedding dimensions

    empirical result

    On ImageNette, the paper compared finite-sliced and analytic-sliced MMD and KSD across projector dimensions 16, 128, and 1024, training horizons of 100, 300, and 800 epochs, and finite direction counts of 16, 128, and 1024. The plotted learning curves show that the analytically sliced versions are the strongest performers across the tested settings. With finite slicing, too few directions slow convergence and produce greater training instability, especially at larger embedding dimensions and during the earlier part of training. The same qualitative pattern appears for both MMD and KSD. The analytic-sliced results show no observed performance effect from changing projector output dimension in these experiments.

  10. Knowl 10 — Evaluation scope leaves scaling and transfer behavior unresolved

    limitation

    The empirical conclusions are based on ImageNette, which the paper characterizes as a small dataset, and the reported evaluation focuses on classification through an online linear probe. The paper therefore leaves open whether the observed effects persist at larger model and dataset scales, how sensitive the variants are to kernel bandwidth, and how they perform in low-data fine-tuning, detection, or segmentation. It also notes that the regularizer shapes projector outputs while downstream inference uses backbone features; consequently, the experiments do not establish whether performance depends on the specified isotropic prior or on the learning dynamics of Euclidean gradients more generally.

Coverage note — The appendix's hyperspherical vMF/uniform-prior extension is omitted because it is peripheral to the paper's main Euclidean KerJEPA methods and experiments; proof-only derivations and auxiliary kernel-calculus tables are also omitted.

References

  1. 1.M. Abramowitz and I. A. Stegun. Handbook of mathematical functions with formulas, graphs, and mathematical tables, volume 55. US Government printing office, 1948.
  2. 2.N. Aronszajn. Theory of reproducing kernels. Transactions of the American Mathematical Society, 68(3):337–404, 1950. ISSN 1088-6850. doi: 10.1090/s0002-9947-1950-0051437-7.
  3. 3.R. Balestriero and Y. LeCun. LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics. arXiv preprint arXiv:2511.08544, 2025.
  4. 4.A. Bardes, J. Ponce, and Y. Lecun. VICReg: Variance-Invariance-Covariance Regularization For Self-Supervised Learning. In International Conference on Learning Representatins (ICLR), 2022.
  5. 5.L. Baringhaus and N. Henze. A consistent test for multivariate normality based on the empirical characteristic function. Metrika, 35(1):339–348, 1988. ISSN 1435-926X. doi: 10.1007/bf02613322.
  6. 6.A. Barp, C.-J. Simon-Gabriel, M. Girolami, and L. Mackey. Targeted separation and convergence with kernel discrepancies. Journal of Machine Learning Research, 25(378):1–50, 2024.
  7. 7.M. G. Bellemare, I. Danihelka, W. Dabney, S. Mohamed, B. Lakshminarayanan, S. Hoyer, and R. Munos. The cramer distance as a solution to biased wasserstein gradients. arXiv preprint arXiv:1705.10743, 2017.
  8. 8.C. Bonet, P. Berg, N. Courty, F. Septier, L. Drumetz, and M.-T. Pham. Spherical sliced-wasserstein. In International Conference on Learning Representatins (ICLR), 2023.
  9. 9.N. Bonneel, J. Rabin, G. Peyré, and H. Pfister. Sliced and Radon Wasserstein barycenters of measures. Journal of Mathematical Imaging and Vision, 51(1):22–45, 2015.
  10. 10.F. Bordes, R. Balestriero, Q. Garrido, A. Bardes, and P. Vincent. Guillotine regularization: Why removing layers is needed to improve generalization in self-supervised learning. Transactions on Machine Learning Research, 2023. ISSN 2835-8856.
  11. 11.M. Caron, I. Misra, J. Mairal, P. Goyal, P. Bojanowski, and A. Joulin. Unsupervised learning of visual features by contrasting cluster assignments. In Advances in Neural Information Processing Systems (NeurIPS), volume 33, pages 9912–9924, 2020.
  12. 12.T. Chen, S. Kornblith, M. Norouzi, and G. Hinton. A simple framework for contrastive learning of visual representations. In International Conference on Machine Learning (ICML), volume 119, pages 1597–1607. PMLR, 2020.
  13. 13.K. Chwialkowski, H. Strathmann, and A. Gretton. A kernel test of goodness of fit. In International Conference on Machine Learning (ICML), volume 48, pages 2606–2615. PMLR, 2016.
  14. 14.C. Domingo-Enrich, R. Dwivedi, and L. Mackey. Compress then test: Powerful kernel testing in near-linear time. In F. Ruiz, J. Dy, and J.-W. van de Meent, editors, International Conference on Artificial Intelligence and Statistics (AISTATS), volume 206, pages 1174–1218. PMLR, 2023.
  15. 15.T. W. Epps and L. B. Pulley. A test for normality based on the empirical characteristic function. Biometrika, 70 (3):723–726, 1983. ISSN 00063444.
  16. 16.fast.ai. imagenette, 2019. URL https://github.com/fastai/imagenette.
  17. 17.W. Gong, Y. Li, and J. M. Hernández-Lobato. Sliced kernelized Stein discrepancy. In International Conference on Learning Representatins (ICLR), 2021.
  18. 18.J. Gorham and L. Mackey. Measuring sample quality with Stein's method. In Advances in Neural Information Processing Systems (NeurIPS), volume 28, 2015.
  19. 19.J. Gorham and L. Mackey. Measuring sample quality with kernels. In International Conference on Machine Learning (ICML), volume 70, pages 1292–1301. PMLR, 2017.
  20. 20.A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Schölkopf, and A. Smola. A kernel two-sample test. Journal of Machine Learning Research, 13(25):723–773, 2012.
  21. 21.J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, B. Piot, k. kavukcuoglu, R. Munos, and M. Valko. Bootstrap your own latent - a new approach to self-supervised learning. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems (NeurIPS), volume 33, pages 21271–21284, 2020.
  22. 22.K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  23. 23.J. Huggins and L. Mackey. Random feature Stein discrepancies. In Advances in Neural Information Processing Systems (NeurIPS), volume 31, 2018.
  24. 24.S. Knop, P. Spurek, J. Tabor, I. Podolak, M. Mazur, and S. Jastrzebski. Cramer-Wold auto-encoder. Journal of Machine Learning Research, 21(164):1–28, 2020.
  25. 25.A. Korba, P.-C. Aubin-Frankowski, S. Majewski, and P. Ablin. Kernel Stein discrepancy descent. In International Conference on Machine Learning (ICML), volume 139, pages 5719–5730. PMLR, 2021.
  26. 26.L. Li, R. Dwivedi, and L. Mackey. Debiased distribution compression. In International Conference on Machine Learning (ICML), volume 235, pages 27675–27731. PMLR, 2024.
  27. 27.Y. Li, R. Pogodin, D. J. Sutherland, and A. Gretton. Self-supervised learning with kernel dependence maximization. In Advances in Neural Information Processing Systems (NeurIPS), volume 34, pages 15543–15556, 2021.
  28. 28.Q. Liu, J. Lee, and M. Jordan. A kernelized Stein discrepancy for goodness-of-fit tests. In International Conference on Machine Learning (ICML), volume 48, pages 276–284. PMLR, 2016.
  29. 29.T. Moutakanni, M. Oquab, M. Szafraniec, M. Vakalopoulou, and P. Bojanowski. You don’t need domain-specific data augmentations when scaling self-supervised learning. In Advances in Neural Information Processing Systems (NeurIPS), volume 37, pages 116106–116125, 2024.
  30. 30.K. Nadjahi, A. Durmus, L. Chizat, S. Kolouri, S. Shahrampour, and U. Simsekli. Statistical and topological properties of sliced probability divergences. In Advances in Neural Information Processing Systems (NeurIPS), volume 33, pages 20802–20812, 2020.
  31. 31.J. Rabin, G. Peyré, J. Delon, and M. Bernot. Wasserstein barycenter and its application to texture mixing. In A. M. Bruckstein, B. M. ter Haar Romeny, A. M. Bronstein, and M. M. Bronstein, editors, Scale Space and Variational Methods in Computer Vision, pages 435–446, Berlin, Heidelberg, 2012. Springer Berlin Heidelberg. ISBN 978-3-642-24785-9.
  32. 32.A. Rahimi and B. Recht. Random features for large-scale kernel machines. In Advances in Neural Information Processing Systems (NeurIPS), volume 20, 2007.
  33. 33.C. E. Rasmussen and C. K. I. Williams. Gaussian Processes for Machine Learning. The MIT Press, 2005. ISBN 9780262256834. doi: 10.7551/mitpress/3206.001.0001.
  34. 34.H. Sepanj and P. Fieguth. Aligning feature distributions in vicreg using maximum mean discrepancy for enhanced manifold awareness in self-supervised representation learning. Journal of Computational Vision and Imaging Systems, 10(1):13–18, Feb. 2025. doi: 10.15353/jcvis.v10i1.10002.
  35. 35.B. K. Sriperumbudur, A. Gretton, K. Fukumizu, B. Schölkopf, and G. R. Lanckriet. Hilbert space embeddings and metrics on probability measures. Journal of Machine Learning Research, 11(50):1517–1561, 2010.
  36. 36.C. Villani et al. Optimal transport: old and new, volume 338. Springer, 2008.
  37. 37.T. Wang and P. Isola. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In International Conference on Machine Learning (ICML), volume 119, pages 9929–9939. PMLR, 2020.
  38. 38.W. Zaremba, A. Gretton, and M. Blaschko. B-test: A non-parametric, low variance kernel two-sample test. In Advances in Neural Information Processing Systems (NeurIPS), volume 26, 2013.
  39. 39.J. Zbontar, L. Jing, I. Misra, Y. LeCun, and S. Deny. Barlow twins: Self-supervised learning via redundancy reduction. In International Conference on Machine Learning (ICML), volume 139, pages 12310–12320. PMLR, 2021.
  40. 40.J. Zhao and D. Meng. FastMMD: Ensemble of circular discrepancy for efficient two-sample test. Neural Computation, 27(6):1345–1372, 2015. doi: 10.1162/NECO_a_00732.

Citation

MLA
Zimmermann, E., et al. “KerJEPA: Kernel Discrepancies for Euclidean Self-Supervised Learning”. arXiv, 2025, http://arxiv.org/abs/2512.19605v1.
APA
Zimmermann, E., Wiltzer, H., Szeto, J., Alvarez-Melis, D., & Mackey, L. (2025). KerJEPA: Kernel Discrepancies for Euclidean Self-Supervised Learning. arXiv. http://arxiv.org/abs/2512.19605v1
Chicago
Zimmermann, E., H. Wiltzer, J. Szeto, D. Alvarez-Melis, and L. Mackey. 2025. “KerJEPA: Kernel Discrepancies for Euclidean Self-Supervised Learning”. arXiv. http://arxiv.org/abs/2512.19605v1.
Harvard
Zimmermann, E. et al. (2025) “KerJEPA: Kernel Discrepancies for Euclidean Self-Supervised Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2512.19605v1.
Vancouver
1. Zimmermann E, Wiltzer H, Szeto J, Alvarez-Melis D, Mackey L (2025) KerJEPA: Kernel Discrepancies for Euclidean Self-Supervised Learning. arXiv

BibTeX

@article{zimmermann2025kerjepa,
  title = {KerJEPA: Kernel Discrepancies for Euclidean Self-Supervised Learning},
  author = {Zimmermann, Eric and Wiltzer, Harley and Szeto, Justin and Alvarez-Melis, David and Mackey, Lester},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2512.19605v1},
  eprint = {2512.19605}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/