Generalized Category Discovery with Decoupled Prototypical Network

Wenbin AnFeng TianQinghua ZhengWei DingQianYing WangPing Chen

article2023AAAI79 citations

Proposes Decoupled Prototypical Network, a framework that uses bipartite prototype matching and semantic-weighted soft assignment to separate known and novel classes, enabling explicit category-specific knowledge transfer in generalized category discovery.

Listen

Modern machine learning models deployed in real-world environments often encounter unlabeled data containing a mixture of previously learned classes and entirely new, unannotated categories. Generalized category discovery addresses this challenge by identifying both known and novel groups without manual annotation, reducing expensive labeling costs and keeping classification taxonomies up to date. Existing approaches treat known and novel categories identically during unsupervised training, which fails to transfer explicit, category-specific knowledge from labeled data and causes models to overfit to noisy automated cluster assignments.

The article develops and evaluates the Decoupled Prototypical Network, a machine learning framework designed to separate known and novel categories systematically and explicitly transfer category-specific guidance into unlabeled data.

The researchers formulated an optimization approach based on category prototypes, which are representative average feature vectors for each class. Using an efficient bipartite matching algorithm, the model aligns labeled prototypes with unlabeled clusters to identify which data belong to known classes and which represent novel categories. The model then applies semi-supervised learning to refine known categories and unsupervised learning to discover novel ones, using soft, semantic-similarity weighting rather than rigid, noisy cluster assignments. Category representations are maintained and stabilized over time using moving averages. The authors validated the method through comparative benchmarking against standard unsupervised and semi-supervised techniques across three standard text classification datasets spanning banking intent, technical discussions, and multi-domain queries.

The evaluation produced several key findings. First, the proposed framework surpassed all existing state-of-the-art methods, achieving an average overall accuracy improvement of 5.87 percentage points across the benchmark datasets, reaching up to 89.06% overall accuracy. Second, it improved classification accuracy on known categories by an average of 4.13 percentage points, demonstrating the benefit of explicit knowledge transfer over general pretraining. Third, the system boosted novel category discovery accuracy by an average of 5.21 percentage points, showing particular gains of nearly 7 percentage points on complex technical text datasets. Fourth, ablation studies demonstrated that semantic-aware soft assignment is critical to performance; removing semantic weighting caused overall accuracy to drop sharply from 84.23% to 35.70%. Finally, the framework demonstrated lower error rates (ranging between 8.7% and 13.0%) when automatically estimating the total number of unknown categories in unlabeled data.

These findings indicate that treating known and novel categories with distinct, decoupled learning objectives resolves major performance bottlenecks in autonomous data discovery. For organizations managing large-scale text systems, this approach reduces operational labeling overhead, enhances automated intent routing, and lowers the risk of misclassifying emergent user requests.

Organizations handling evolving text classification tasks should consider adopting decoupled prototypical architectures to expand existing taxonomies autonomously. Technical teams should implement semantic-weighted soft assignment rather than hard pseudo-label assignments when clustering unlabeled data to safeguard against boundary noise. Future development should focus on applying this discovery architecture to multimodal domains, such as image and audio classification, while validating category estimation stability in live operational pipelines.

The experimental findings rely on assumptions of benchmark text data where unlabeled datasets contain all known categories, and the initial experiments assumed prior knowledge of the total category count before testing automated estimation algorithms. Nonetheless, the consistent performance gains across varied category ratios and domains provide high confidence in the framework's effectiveness for automated category discovery.

Cover for Generalized Category Discovery with Decoupled Prototypical Network

Abstract

Generalized Category Discovery (GCD) aims to recognize both known and novel categories from a set of unlabeled data, based on another dataset labeled with only known categories. Without considering differences between known and novel categories, current methods learn about them in a coupled manner, which can hurt model’s generalization and discriminative ability. Furthermore, the coupled training approach prevents these models transferring category-specific knowledge explicitly from labeled data to unlabeled data, which can lose high-level semantic information and impair model performance. To mitigate above limitations, we present a novel model called Decoupled Prototypical Network (DPN). By formulating a bipartite matching problem for category prototypes, DPN can not only decouple known and novel categories to achieve different training targets effectively, but also align known categories in labeled and unlabeled data to transfer category-specific knowledge explicitly and capture high-level semantics. Furthermore, DPN can learn more discriminative features for both known and novel categories through our proposed Semantic-aware Prototypical Learning (SPL). Besides capturing meaningful semantic information, SPL can also alleviate the noise of hard pseudo labels through semantic-weighted soft assignment. Extensive experiments show that DPN outperforms state-of-the-art models by a large margin on all evaluation metrics across multiple benchmark datasets. Code and data are available at https://github.com/Lackel/DPN.

Table of Contents

  • Introduction
  • Related Work
  • Generalized Category Discovery
  • Prototypical Learning
  • Method
  • Problem Statement
  • Approach Overview
  • Representation Learning
  • Learning Category Prototypes
  • Alignment and Decoupling
  • Semantic-aware Prototypical Learning
  • Updating Category Prototypes
  • Experiments
  • Experimental Setup
  • Experimental Results
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Problem Formulation of Generalized Category Discovery

    definition

    Generalized Category Discovery (GCD) in text classification involves recognizing both known and novel categories from an unlabeled dataset using a partially labeled category taxonomy.

    Formally, the setting contains:

    • A labeled dataset Dl={(xi,yi)∣yi∈Yk}\mathcal{D}^l = \{(x_i, y_i) \mid y_i \in \mathcal{Y}_k\}, containing instances annotated strictly with labels from a set of known categories Yk\mathcal{Y}_k, where M=∣Yk∣M = |\mathcal{Y}_k| is the number of known classes.
    • An unlabeled dataset Du={xi∣yi∈{Yk∪Yn}}\mathcal{D}^u = \{x_i \mid y_i \in \{\mathcal{Y}_k \cup \mathcal{Y}_n\}\}, containing unannotated instances drawn from both known categories Yk\mathcal{Y}_k and novel categories Yn\mathcal{Y}_n, where ∣Yn∣|\mathcal{Y}_n| is the number of novel classes and K=∣Yk∣+∣Yn∣K = |\mathcal{Y}_k| + |\mathcal{Y}_n| is the total number of categories.

    The objective is to train a feature extractor FθF_\theta on Dl\mathcal{D}^l and Du\mathcal{D}^u without additional human annotation, such that it correctly partitions and assigns category labels to instances in a test set Dt={(xi,yi)∣yi∈{Yk∪Yn}}\mathcal{D}^t = \{(x_i, y_i) \mid y_i \in \{\mathcal{Y}_k \cup \mathcal{Y}_n\}\} spanning all known and novel classes. Model performance is evaluated by clustering accuracy on all instances (All), instances from known categories (Known), and instances from novel categories (Novel), using Hungarian bipartite matching between predicted cluster indices and ground-truth class labels.

  2. Knowl 2 — Prototype Alignment and Category Decoupling via Bipartite Matching

    model/method

    To decouple known and novel categories from unlabeled data and enable category-specific knowledge transfer without adding extra classifier parameters, Decoupled Prototypical Network (DPN) constructs two sets of prototypes and aligns them using minimum-cost bipartite matching:

    1. Labeled Prototypes: For each known category j∈{1,…,M}j \in \{1, \dots, M\}, the labeled prototype μjl\mu_j^l is the mean feature embedding over all labeled samples belonging to category jj: μjl=1∣Cj∣∑xi∈CjFθ(xi)\mu_j^l = \frac{1}{|C_j|} \sum_{x_i \in C_j} F_\theta(x_i) where Fθ:X→RdF_\theta: \mathcal{X} \to \mathbb{R}^d is a feature extractor, CjC_j is the set of labeled instances in category jj, and Pl={μjl}j=1M\mathcal{P}^l = \{\mu_j^l\}_{j=1}^M.

    2. Unlabeled Prototypes: KK-Means clustering is applied to all representations in Du\mathcal{D}^u to produce KK clusters {C1u,C2u,…,CKu}\{C_1^u, C_2^u, \dots, C_K^u\}, where K=∣Yk∣+∣Yn∣K = |\mathcal{Y}_k| + |\mathcal{Y}_n|. The centroid for cluster jj is: μju=1∣Cju∣∑xi∈CjuFθ(xi)\mu_j^u = \frac{1}{|C_j^u|} \sum_{x_i \in C_j^u} F_\theta(x_i) yielding the unlabeled prototype set Pu={μju}j=1K\mathcal{P}^u = \{\mu_j^u\}_{j=1}^K.

    3. Bipartite Matching: Under the assumption that Du\mathcal{D}^u contains all known categories and that the closest prototypes between labeled and unlabeled sets represent the same category, an optimal permutation P^\hat{P} over Pu\mathcal{P}^u is found by minimizing the total Euclidean matching cost: P^=arg⁡min⁡P∈Pall∑i=1M∥μil−μP(i)u∥2\hat{P} = \arg\min_{P \in \mathcal{P}_{all}} \sum_{i=1}^M \|\mu_i^l - \mu_{P(i)}^u\|_2 where Pall\mathcal{P}_{all} denotes the set of all possible permutations of length MM from {1,…,K}\{1, \dots, K\}. This assignment is solved in polynomial time using the Hungarian algorithm.

    4. Decoupling: The matched unlabeled prototypes Puk={μP^(i)u}i=1M\mathcal{P}^{uk} = \{\mu_{\hat{P}(i)}^u\}_{i=1}^M represent known categories, while the remaining unmatched prototypes Pun={μP^(i)u}i=M+1K\mathcal{P}^{un} = \{\mu_{\hat{P}(i)}^u\}_{i=M+1}^K represent novel categories. Correspondingly, unlabeled instances are partitioned into known data Duk\mathcal{D}^{uk} and novel data Dun\mathcal{D}^{un} according to their cluster assignments.

  3. Knowl 3 — Semantic-Aware Prototypical Learning Objective for Novel Categories

    model/method

    Traditional prototypical learning assigns each instance to a single prototype based on hard pseudo labels, which causes overfitting to clustering noise near decision boundaries. Semantic-aware Prototypical Learning (SPL) assigns each instance to all prototypes in a soft manner weighted by semantic similarity.

    For unlabeled instances decoupled into novel categories Dun\mathcal{D}^{un} with prototype set Pun={μku}k=1K′\mathcal{P}^{un} = \{\mu_k^u\}_{k=1}^{K'}, where K′=∣Yn∣K' = |\mathcal{Y}_n| and n=∣Dun∣n = |\mathcal{D}^{un}|, the SPL loss function is: Lnovel=Lspl(Dun,Pun)=1n∑i=1n∑k=1K′∥Fθ(xi)−μku∥2exp⁡(cos⁡(Fθ(xi),μku)/τ)∑j=1K′exp⁡(cos⁡(Fθ(xi),μju)/τ)\mathcal{L}_{novel} = \mathcal{L}_{spl}(\mathcal{D}^{un}, \mathcal{P}^{un}) = \frac{1}{n} \sum_{i=1}^n \sum_{k=1}^{K'} \|F_\theta(x_i) - \mu_k^u\|_2 \frac{\exp(\cos(F_\theta(x_i), \mu_k^u)/\tau)}{\sum_{j=1}^{K'} \exp(\cos(F_\theta(x_i), \mu_j^u)/\tau)} where:

    • Fθ(xi)∈RdF_\theta(x_i) \in \mathbb{R}^d is the feature embedding of instance xix_i,
    • ∥Fθ(xi)−μku∥2\|F_\theta(x_i) - \mu_k^u\|_2 is the Euclidean distance pulling instance embeddings toward category prototypes,
    • cos⁡(Fθ(xi),μku)=Fθ(xi)⊤μku∥Fθ(xi)∥2∥μku∥2\cos(F_\theta(x_i), \mu_k^u) = \frac{F_\theta(x_i)^\top \mu_k^u}{\|F_\theta(x_i)\|_2 \|\mu_k^u\|_2} is the cosine similarity weighting the soft assignment,
    • τ>0\tau > 0 is a temperature hyperparameter (set to τ=0.07\tau = 0.07).

    By weighting distances across all novel prototypes, SPL captures high-level semantic relationships while reducing the negative impact of noisy pseudo labels.

  4. Knowl 4 — Semi-Supervised Learning and Category-Specific Knowledge Regularization for Known Categories

    model/method

    For unlabeled instances decoupled into known categories Duk\mathcal{D}^{uk}, DPN transfers category-specific knowledge explicitly from labeled prototypes Pl={μkl}k=1M\mathcal{P}^l = \{\mu_k^l\}_{k=1}^M to guide the pseudo-label training process in a semi-supervised manner.

    The loss for known categories is formulated as: Lknown=Lspl(Duk,Puk)+Lce(Dl)+γ⋅Lreg(Duk,Pl)\mathcal{L}_{known} = \mathcal{L}_{spl}(\mathcal{D}^{uk}, \mathcal{P}^{uk}) + \mathcal{L}_{ce}(\mathcal{D}^l) + \gamma \cdot \mathcal{L}_{reg}(\mathcal{D}^{uk}, \mathcal{P}^l) where:

    1. Lspl(Duk,Puk)\mathcal{L}_{spl}(\mathcal{D}^{uk}, \mathcal{P}^{uk}) is the Semantic-aware Prototypical Learning loss applied to the MM aligned unlabeled known prototypes Puk\mathcal{P}^{uk}.
    2. Lce(Dl)\mathcal{L}_{ce}(\mathcal{D}^l) is standard cross-entropy classification loss on labeled data Dl\mathcal{D}^l to prevent catastrophic forgetting.
    3. Lreg(Duk,Pl)\mathcal{L}_{reg}(\mathcal{D}^{uk}, \mathcal{P}^l) is an explicit category knowledge regularization term that pulls unlabeled known instances toward labeled prototypes using semantic similarity weights: Lreg=1r∑i=1r∑k=1M(1−cos⁡(Fθ(xi),μkl))exp⁡(cos⁡(Fθ(xi),μkl)/τ)∑j=1Mexp⁡(cos⁡(Fθ(xi),μjl)/τ)\mathcal{L}_{reg} = \frac{1}{r} \sum_{i=1}^r \sum_{k=1}^M (1 - \cos(F_\theta(x_i), \mu_k^l)) \frac{\exp(\cos(F_\theta(x_i), \mu_k^l)/\tau)}{\sum_{j=1}^M \exp(\cos(F_\theta(x_i), \mu_j^l)/\tau)} where r=∣Duk∣r = |\mathcal{D}^{uk}| is the number of unlabeled instances in known categories, τ=0.07\tau = 0.07 is the temperature parameter, and γ>0\gamma > 0 is a regularization weight.

    The complete training objective of DPN is: Ldpn=Lnovel+Lknown\mathcal{L}_{dpn} = \mathcal{L}_{novel} + \mathcal{L}_{known}

  5. Knowl 5 — DPN Training Procedure and Exponential Moving Average Prototype Updates

    algorithm

    The Decoupled Prototypical Network training pipeline includes representation pre-training, prototype initialization via clustering and bipartite alignment, and decoupled optimization with Exponential Moving Average (EMA) prototype updates.

    Input: Labeled dataset Dl\mathcal{D}^l, unlabeled dataset Du\mathcal{D}^u, number of known categories MM, total categories KK, momentum factor α=0.9\alpha=0.9, regularization weight γ\gamma, temperature τ=0.07\tau=0.07
    Output: Trained feature extractor FθF_\theta and final cluster assignments for testing data Dt\mathcal{D}^t
    # Phase 1: Joint Pre-training
    Initialize BERT feature extractor FθF_\theta
    while pre-training not converged do
        Compute joint pre-training loss Lpre=Lce(Dl)+Lmlm(Dl,Du)\mathcal{L}_{pre} = \mathcal{L}_{ce}(\mathcal{D}^l) + \mathcal{L}_{mlm}(\mathcal{D}^l, \mathcal{D}^u)
        Update FθF_\theta parameters with AdamW
    end while
    # Phase 2: Category Prototype Initialization and Decoupled Training
    Compute initial labeled prototypes P0l={μjl}j=1M\mathcal{P}^l_0 = \{\mu_j^l\}_{j=1}^M using class averages on Dl\mathcal{D}^l
    Perform K-Means clustering on Fθ(Du)F_\theta(\mathcal{D}^u) to obtain KK clusters and centroids Pu={μju}j=1K\mathcal{P}^u = \{\mu_j^u\}_{j=1}^K
    Solve bipartite matching P^=arg⁡min⁡P∑i=1M∥μil−μP(i)u∥2\hat{P} = \arg\min_P \sum_{i=1}^M \|\mu_i^l - \mu_{P(i)}^u\|_2 using the Hungarian algorithm
    Partition Pu\mathcal{P}^u into known prototypes Puk\mathcal{P}^{uk} and novel prototypes Pun\mathcal{P}^{un}
    Partition Du\mathcal{D}^u into Duk\mathcal{D}^{uk} and Dun\mathcal{D}^{un} according to cluster assignments
    for epoch t=0,1,…,T−1t = 0, 1, \dots, T-1 do
        Compute novel loss Lnovel=Lspl(Dun,Pun)\mathcal{L}_{novel} = \mathcal{L}_{spl}(\mathcal{D}^{un}, \mathcal{P}^{un})
        Compute known loss Lknown=Lspl(Duk,Puk)+Lce(Dl)+γ⋅Lreg(Duk,Ptl)\mathcal{L}_{known} = \mathcal{L}_{spl}(\mathcal{D}^{uk}, \mathcal{P}^{uk}) + \mathcal{L}_{ce}(\mathcal{D}^l) + \gamma \cdot \mathcal{L}_{reg}(\mathcal{D}^{uk}, \mathcal{P}^l_t)
        Compute total loss Ldpn=Lnovel+Lknown\mathcal{L}_{dpn} = \mathcal{L}_{novel} + \mathcal{L}_{known}
        Update FθF_\theta (fine-tuning the last three Transformer layers) via AdamW
        
        # EMA Update for Labeled Prototypes
        Compute empirical mean prototype Pt+1l\mathcal{P}^l_{t+1} on Dl\mathcal{D}^l using updated FθF_\theta
        Update Pt+1l←α⋅Ptl+(1−α)⋅Pt+1l\mathcal{P}^l_{t+1} \leftarrow \alpha \cdot \mathcal{P}^l_t + (1 - \alpha) \cdot \mathcal{P}^l_{t+1}
        # Unlabeled prototypes Pu\mathcal{P}^u are kept fixed to avoid re-clustering overhead
    end for
    Extract representations Fθ(Dt)F_\theta(\mathcal{D}^t) and apply K-Means clustering to predict category labels
  6. Knowl 6 — Experimental Setup and Hyperparameter Configurations for GCD

    experimental setup

    DPN is evaluated on three text classification benchmarks under the Generalized Category Discovery setting:

    • BANKING: Intent classification in the banking domain with 77 total classes (∣Yk∣=58|\mathcal{Y}_k|=58 known, ∣Yn∣=19|\mathcal{Y}_n|=19 novel; ∣Dl∣=673|\mathcal{D}^l|=673, ∣Du∣=8,330|\mathcal{D}^u|=8,330, ∣Dt∣=3,080|\mathcal{D}^t|=3,080).
    • StackOverflow: Technical question classification with 20 total classes (∣Yk∣=15|\mathcal{Y}_k|=15 known, ∣Yn∣=5|\mathcal{Y}_n|=5 novel; ∣Dl∣=1,350|\mathcal{D}^l|=1,350, ∣Du∣=16,650|\mathcal{D}^u|=16,650, ∣Dt∣=1,000|\mathcal{D}^t|=1,000).
    • CLINC: Multi-domain intent classification with 150 total classes (∣Yk∣=113|\mathcal{Y}_k|=113 known, ∣Yn∣=37|\mathcal{Y}_n|=37 novel; ∣Dl∣=1,344|\mathcal{D}^l|=1,344, ∣Du∣=16,656|\mathcal{D}^u|=16,656, ∣Dt∣=2,250|\mathcal{D}^t|=2,250).

    Implementation and Hyperparameters:

    • Backbone: bert-base-uncased. During pre-training, all Transformer layers are updated. During DPN training, only the top three Transformer layers are fine-tuned, and the [CLS] token embedding represents each text instance.
    • Pre-training: Masked language model probability is 0.150.15; AdamW optimizer with learning rate 5×10−55 \times 10^{-5} and early stopping patience of 20 epochs based on validation accuracy on known classes.
    • DPN Training: AdamW optimizer with learning rate 1×10−51 \times 10^{-5}; EMA momentum factor α=0.9\alpha = 0.9; temperature τ=0.07\tau = 0.07. Regularization factor γ\gamma is set to 1010 for BANKING, 1010 for CLINC, and 9090 for StackOverflow. Training epochs are set to 6060 for BANKING, 1010 for StackOverflow, and 8080 for CLINC.
    • Evaluation: Accuracy on All, Known, and Novel categories calculated via Hungarian matching on Dt\mathcal{D}^t.
  7. Knowl 7 — Benchmark Performance of DPN on Generalized Category Discovery

    data/table

    DPN was evaluated against state-of-the-art unsupervised and semi-supervised category discovery baselines across the BANKING, StackOverflow, and CLINC datasets. Average accuracy (%) over 3 runs is reported on All instances, Known category instances, and Novel category instances.

    Method BANKING StackOverflow CLINC
    All Known Novel All Known Novel All Known Novel
    DeepCluster 13.95 13.94 13.99 17.37 18.22 14.80 26.92 27.34 25.67
    DCN 17.85 18.94 14.35 29.10 28.94 29.51 29.64 30.00 28.45
    DEC 19.30 20.36 15.84 19.30 20.36 15.84 19.99 20.18 19.40
    BERT 21.29 21.48 20.70 16.80 16.67 17.20 34.52 34.98 33.16
    KM-GloVe 29.18 29.11 29.39 28.40 28.60 28.05 51.64 51.74 51.50
    AG-GloVe 30.09 29.69 31.29 29.23 28.49 31.56 44.70 45.17 43.20
    SAE 38.05 38.29 37.27 60.33 57.36 69.02 46.59 47.35 44.24
    Semi-DC 50.73 53.37 42.63 64.90 66.13 61.20 74.52 75.60 71.34
    CDAC+ 53.09 55.42 46.01 76.67 77.51 74.13 69.75 70.08 68.77
    Self-Labeling 56.19 61.64 39.56 71.03 78.53 48.53 72.69 80.06 49.65
    DTC 56.56 59.98 46.10 70.50 80.93 51.87 76.42 82.34 58.95
    DAC 63.63 69.60 45.44 70.77 76.13 54.67 84.42 89.10 70.59
    Semi-KM 66.23 73.62 43.68 73.13 81.02 49.47 81.42 89.03 59.01
    LASKM 67.55 75.16 44.34 74.83 82.00 53.33 79.26 89.64 48.66
    DPN (Ours) 72.96 80.93 48.60 84.23 85.29 81.07 89.06 92.97 77.54
    Improvement +5.41 +5.77 +2.50 +7.56 +3.29 +6.94 +4.64 +3.33 +6.20

    DPN achieves state-of-the-art results across all datasets and evaluation metrics:

    • Overall accuracy (All) improves over the prior best method by +5.41%+5.41\% on BANKING, +7.56%+7.56\% on StackOverflow, and +4.64%+4.64\% on CLINC (an average improvement of 5.87%5.87\% across datasets).
    • Known category accuracy improves by an average of 4.13%4.13\% across benchmarks (80.93%80.93\%, 85.29%85.29\%, and 92.97%92.97\%) because prototype alignment explicitly transfers category-specific knowledge.
    • Novel category accuracy improves by an average of 5.21%5.21\% across benchmarks (48.60%48.60\%, 81.07%81.07\%, and 77.54%77.54\%) through soft semantic weighting in SPL.
  8. Knowl 8 — Ablation Analysis of Decoupled Prototypical Network Components

    empirical result

    An ablation study on the StackOverflow dataset isolates the impact of each core component in DPN:

    Model Variant All (%) Known (%) Novel (%)
    Full DPN 84.23 85.29 81.07
    w/o Cross Entropy 83.83 85.02 80.26
    w/o EMA 82.50 83.87 78.40
    w/o Decoupling 78.77 78.53 79.47
    w/o Soft Assignment 75.10 75.33 74.40
    w/o Semantic Weights 35.70 33.73 41.60

    Findings from component removal:

    1. Semantic Weights: Removing similarity weights s(Fθ(xi),μku)s(F_\theta(x_i), \mu_k^u) causes the largest performance drop (All accuracy plummets from 84.23%84.23\% to 35.70%35.70\%), indicating that semantic similarity weighting is essential for feature representation learning.
    2. Soft Assignment: Switching to hard cluster assignment reduces All accuracy by 9.13%9.13\% (to 75.10%75.10\%) because hard assignments overfit to clustering noise.
    3. Decoupling: Removing prototype alignment and decoupling drops All accuracy by 5.46%5.46\% (to 78.77%78.77\%), hurting both known (78.53%78.53\%) and novel (79.47%79.47\%) classes by preventing category-specific knowledge transfer.
    4. EMA Prototype Updates: Omitting EMA updates reduces All accuracy by 1.73%1.73\% (to 82.50%82.50\%) due to prototype representation instability from batch outliers.
    5. Cross-Entropy: Removing supervised cross-entropy loss Lce(Dl)\mathcal{L}_{ce}(\mathcal{D}^l) causes only a slight decrease (−0.40%-0.40\% All), confirming that labeled prototypes alone transfer strong supervisory guidance.
  9. Knowl 9 — Robustness to Known Category Ratio and Category Number Estimation

    empirical result

    DPN provides superior category count estimation and maintains robust performance across different known-to-novel class ratios:

    1. Category Number Estimation (KK): Using the estimation algorithm from Deep Aligned Clustering (DAC), DPN's learned representations predict the total number of categories K=∣Yk∣+∣Yn∣K = |\mathcal{Y}_k| + |\mathcal{Y}_n| more accurately than DAC across all benchmarks:
    • CLINC (Ground truth K=150K=150): DAC estimates 130130 (error 13.3%13.3\%), whereas DPN estimates 137137 (error 8.7%8.7\%).
    • BANKING (Ground truth K=77K=77): DAC estimates 6666 (error 14.3%14.3\%), whereas DPN estimates 6767 (error 13.0%13.0\%).
    • StackOverflow (Ground truth K=20K=20): DAC estimates 1515 (error 25.0%25.0\%), whereas DPN estimates 1818 (error 10.0%10.0\%).
    1. Varying Known Category Ratio: Evaluating DPN on BANKING with known category ratios in {0.25,0.50,0.75}\{0.25, 0.50, 0.75\} demonstrates consistent superiority:
    • DPN outperforms all comparison methods (DAC, CDAC+, Semi-KM, Semi-DC, Self-Labeling, LASKM, DTC) under all ratio settings on All, Known, and Novel metrics.
    • As the known category ratio increases from 0.250.25 to 0.750.75, DPN's Overall accuracy rises steadily from approximately 35%35\% to over 75%75\%, and Known accuracy rises from approximately 50%50\% to over 80%80\%, while Novel category accuracy consistently exceeds the nearest baseline by 3–6%3\text{--}6\%.

Coverage note — None was omitted; all key contributions including framework formulation, alignment mechanism, loss definitions, EMA updates, experiments, ablations, and parameter analyses are covered.

References

  1. 1.An, W.; Tian, F.; Chen, P.; Tang, S.; Zheng, Q.; and Wang, Q. 2022. Fine-grained Category Discovery under Coarse-grained supervision with Hierarchical Weighted Self-contrastive Learning. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 1314–1323. Association for Computational Linguistics.
  2. 2.Caron, M.; Bojanowski, P.; Joulin, A.; and Douze, M. 2018. Deep clustering for unsupervised learning of visual features. In Proceedings of the European Conference on Computer Vision (ECCV), 132–149.
  3. 3.Casanueva, I.; Temcinas, T.; Gerz, D.; Henderson, M.; and Vulić, I. 2020. Efficient intent detection with dual sentence encoders. arXiv preprint arXiv:2003.04807.
  4. 4.Chen, C.; Li, O.; Tao, D.; Barnett, A.; Rudin, C.; and Su, J. K. 2019. This looks like that: deep learning for interpretable image recognition. Advances in neural information processing systems, 32.
  5. 5.Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  6. 6.Gowda, K. C.; and Krishna, G. 1978. Agglomerative clustering using the concept of mutual nearest neighbourhood. Pattern recognition, 10(2): 105–112.
  7. 7.Han, K.; Vedaldi, A.; and Zisserman, A. 2019. Learning to discover novel visual categories via deep transfer clustering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 8401–8409.
  8. 8.Kuhn, H. W. 1955. The Hungarian method for the assignment problem. Naval research logistics quarterly, 2: 83–97.
  9. 9.Larson, S.; Mahendran, A.; Peper, J. J.; Clarke, C.; Lee, A.; Hill, P.; Kummerfeld, J. K.; Leach, K.; Laurenzano, M. A.; Tang, L.; et al. 2019. An evaluation dataset for intent classification and out-of-scope prediction. arXiv preprint arXiv:1909.02027.
  10. 10.Li, J.; Zhou, P.; Xiong, C.; and Hoi, S. C. 2020. Prototypical contrastive learning of unsupervised representations. arXiv preprint arXiv:2005.04966.
  11. 11.Lin, T.-E.; Xu, H.; and Zhang, H. 2020. Discovering new intents via constrained deep adaptive clustering with cluster refinement. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, 8360–8367.
  12. 12.MacQueen, J.; et al. 1967. Some methods for classification and analysis of multivariate observations. In Proceedings of the fifth Berkeley symposium on mathematical statistics and probability, volume 1, 281–297. Oakland, CA, USA.
  13. 13.Nauta, M.; van Bree, R.; and Seifert, C. 2021. Neural prototype trees for interpretable fine-grained image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 14933–14943.
  14. 14.Pan, Y.; Yao, T.; Li, Y.; Wang, Y.; Ngo, C.-W.; and Mei, T. 2019. Transferrable prototypical networks for unsupervised domain adaptation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2239–2247.
  15. 15.Pennington, J.; Socher, R.; and Manning, C. D. 2014. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), 1532–1543.
  16. 16.Snell, J.; Swersky, K.; and Zemel, R. 2017. Prototypical networks for few-shot learning. Advances in neural information processing systems, 30.
  17. 17.Vaze, S.; Han, K.; Vedaldi, A.; and Zisserman, A. 2022. Generalized category discovery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7492–7501.
  18. 18.Vedula, N.; Gupta, R.; Alok, A.; and Sridhar, M. 2020. Automatic discovery of novel intents & domains from text utterances. arXiv preprint arXiv:2006.01208.
  19. 19.Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; et al. 2019. HuggingFace’s Transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771.
  20. 20.Xie, J.; Girshick, R.; and Farhadi, A. 2016. Unsupervised deep embedding for clustering analysis. In International conference on machine learning, 478–487. PMLR.
  21. 21.Xu, J.; Wang, P.; Tian, G.; Xu, B.; Zhao, J.; Wang, F.; and Hao, H. 2015. Short text clustering via convolutional neural networks. In Proceedings of the 1st Workshop on Vector Space Modeling for Natural Language Processing, 62–69.
  22. 22.Yang, B.; Fu, X.; Sidiropoulos, N. D.; and Hong, M. 2017. Towards k-means-friendly spaces: Simultaneous deep learning and clustering. In international conference on machine learning, 3861–3870. PMLR.
  23. 23.Yu, Q.; Ikami, D.; Irie, G.; and Aizawa, K. 2022. Self-Labeling Framework for Novel Category Discovery over Domains. In Proceedings of the AAAI Conference on Artificial Intelligence.
  24. 24.Yue, X.; Zheng, Z.; Zhang, S.; Gao, Y.; Darrell, T.; Keutzer, K.; and Vincentelli, A. S. 2021. Prototypical cross-domain self-supervised learning for few-shot unsupervised domain adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13834–13844.
  25. 25.Zhang, H.; Li, X.; Xu, H.; Zhang, P.; Zhao, K.; and Gao, K. 2021a. TEXTOIR: An integrated and visualized platform for text open intent recognition. arXiv preprint arXiv:2110.15063.
  26. 26.Zhang, H.; Xu, H.; Lin, T.-E.; and Lyu, R. 2021b. Discovering New Intents with Deep Aligned Clustering. In Proceedings of the AAAI Conference on Artificial Intelligence.
  27. 27.Zhang, J.; Bui, T.; Yoon, S.; Chen, X.; Liu, Z.; Xia, C.; Tran, Q. H.; Chang, W.; and Yu, P. 2021c. Few-shot intent detection via contrastive pre-training and fine-tuning. arXiv preprint arXiv:2109.06349.
  28. 28.Zhang, Y.; Zhang, H.; Zhan, L.-M.; Wu, X.-M.; and Lam, A. 2022. New Intent Discovery with Pre-training and Contrastive Learning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics.

Citation

MLA
An, W., et al. “Generalized Category Discovery with Decoupled Prototypical Network”. arXiv, 2022, http://arxiv.org/abs/2211.15115v2.
APA
An, W., Tian, F., Zheng, Q., Ding, W., Wang, Q., & Chen, P. (2022). Generalized Category Discovery with Decoupled Prototypical Network. arXiv. http://arxiv.org/abs/2211.15115v2
Chicago
An, W., F. Tian, Q. Zheng, W. Ding, Q. Wang, and P. Chen. 2022. “Generalized Category Discovery with Decoupled Prototypical Network”. arXiv. http://arxiv.org/abs/2211.15115v2.
Harvard
An, W. et al. (2022) “Generalized Category Discovery with Decoupled Prototypical Network”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.15115v2.
Vancouver
1. An W, Tian F, Zheng Q, Ding W, Wang Q, Chen P (2022) Generalized Category Discovery with Decoupled Prototypical Network. arXiv

BibTeX

@article{an2022generalized,
  title = {Generalized Category Discovery with Decoupled Prototypical Network},
  author = {An, Wenbin and Tian, Feng and Zheng, Qinghua and Ding, Wei and Wang, QianYing and Chen, Ping},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.15115v2},
  eprint = {2211.15115}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF