Built independently by an author, for readers. Read the story and support ChapterPal

keyword

kernel alignment

Kernel alignment is a statistical similarity metric in machine learning that quantifies the degree of agreement between two kernel functions or between a kernel matrix and target data labels over a given dataset. Mathematically, it computes a normalized inner product representing the cosine of the angle between two Gram matrices in the Frobenius inner product space. This metric is widely used to evaluate and optimize the selection of kernel functions prior to model training, as well as to assess how well a kernel captures the underlying class structure of a learning task. In extended forms, such as centered kernel alignment, it provides an invariance to isotropic scaling and orthogonal transformations, making it a foundational tool for comparing the geometry of internal representations across different neural network layers and architectures.

3 items

Text Classification using String Kernels

Text Classification using String Kernels

H. Lodhi, C. Saunders, J. Shawe-Taylor, N. Cristianini, Christopher J. C. H. Watkins

OrganizationsRoyal Holloway, University of London

Why you should read this

Proposes a string subsequence kernel for text categorization that captures non-contiguous character sequences via dynamic programming and outperforms traditional bag-of-words representations on benchmark corpora.

We propose a novel approach for categorizing text documents based on the use of a special kernel. The kernel is an inner product in the feature space generated by all subsequences of length k. A subsequence is any ordered sequence of k characters occurring in the text though not necessarily contiguously. The subsequences are weighted by an exponentially decaying factor of their full length in the text, hence emphasising those occurrences that are close to contiguous. A direct computation of this feature vector would involve a prohibitive amount of computation even for modest values of k, since the dimension of the feature space grows exponentially with k. The paper describes how despite this fact the inner product can be efficiently evaluated by a dynamic programming technique. Experimental comparisons of the performance of the kernel compared with a standard word feature space kernel (Joachims, 1998) show positive results on modestly sized datasets. The case of contiguous subsequences is also considered for comparison with the subsequences kernel with different decay factors. For larger documents and datasets the paper introduces an approximation technique that is shown to deliver good approximations efficiently for large datasets.

Added

2026-09-25

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, Saining Xie

OrganizationsKorea Advanced Institute of Science and TechnologyKorea UniversityNew York UniversityScaled Foundations

Why you should read this

Introduces REPA, a simple regularization framework that achieves over 17.5x faster training convergence and state-of-the-art image generation quality by aligning diffusion transformer representations with those from pretrained self-supervised visual encoders.

Recent studies have shown that the denoising process in (generative) diffusion models can induce meaningful (discriminative) representations inside the model, though the quality of these representations still lags behind those learned through recent self-supervised learning methods. We argue that one main bottleneck in training large-scale diffusion models for generation lies in effectively learning these representations. Moreover, training can be made easier by incorporating high-quality external visual representations, rather than relying solely on the diffusion models to learn them independently. We study this by introducing a straightforward regularization called REPresentation Alignment (REPA), which aligns the projections of noisy input hidden states in denoising networks with clean image representations obtained from external, pretrained visual encoders. The results are striking: our simple strategy yields significant improvements in both training efficiency and generation quality when applied to popular diffusion and flow-based transformers, such as DiTs and SiTs. For instance, our method can speed up SiT training by over 17.5×\times, matching the performance (without classifier-free guidance) of a SiT-XL model trained for 7M steps in less than 400K steps. In terms of final generation quality, our approach achieves state-of-the-art results of FID=1.42 using classifier-free guidance with the guidance interval.

Added

2026-05-18