Built independently by an author, for readers. Read the story and support ChapterPal

keyword

generative moment matching

Generative moment matching is a machine learning approach used to train generative models by aligning the statistical moments of generated data with those of the true target data distribution. Rather than calculating exact data likelihoods or relying on minimax adversarial training with a separate discriminator network, this framework typically minimizes a distance metric between probability distributions, most commonly maximum mean discrepancy computed with kernel methods. By embedding probability distributions into a reproducing kernel Hilbert space, generative moment matching compares and matches statistics across all orders between the empirical data samples and model-generated samples. The resulting discrepancy serves as a direct, differentiable loss function, enabling the generative network parameters to be trained straightforwardly through standard gradient descent and backpropagation to synthesize realistic, high-dimensional representations.

1 item

Primitive Generation and Semantic-Related Alignment for Universal Zero-Shot Segmentation

Primitive Generation and Semantic-Related Alignment for Universal Zero-Shot Segmentation

Shuting He, Henghui Ding, Wei Jiang

OrganizationsNanyang Technological UniversityZhejiang University

Why you should read this

Proposes PADing, a unified universal zero-shot segmentation framework that bridges the cross-modal domain gap by assembling learned fine-grained primitives to synthesize unseen visual features and aligning their semantic-related components with linguistic class relationships.

We study universal zero-shot segmentation in this work to achieve panoptic, instance, and semantic segmentation for novel categories without any training samples. Such zero-shot segmentation ability relies on inter-class relationships in semantic space to transfer the visual knowledge learned from seen categories to unseen ones. Thus, it is desired to well bridge semantic and visual spaces and apply the semantic relationships to visual feature learning. We introduce a generative model to synthesize features for unseen categories, which links semantic and visual spaces as well as addresses the issue of lack of unseen training data. Furthermore, to mitigate the domain gap between semantic and visual spaces, firstly, we enhance the vanilla generator with learned primitives, each of which contains fine-grained attributes related to categories, and synthesize unseen features by selectively assembling these primitives. Secondly, we propose to disentangle the visual feature into the semantic-related part and the semantic-unrelated part that contains useful visual classification clues but is less relevant to semantic representation. The inter-class relationships of semantic-related visual features are then required to be aligned with those in semantic space, thereby transferring semantic knowledge to visual feature learning. The proposed approach achieves impressively state-of-the-art performance on zero-shot panoptic segmentation, instance segmentation, and semantic segmentation.

Added

2026-09-26