CONTaiNER: Few-Shot Named Entity Recognition via Contrastive Learning

Sarkar Snigdha Sarathi DasArzoo KatiyarRebecca J. PassonneauRui Zhang

article2022ACL187 citations

Proposes a contrastive learning framework that models token representations as Gaussian distributions to optimize distributional divergence between entity types, achieving substantial gains over existing few-shot named entity recognition methods across multiple benchmarks.

Listen

Extracting structured information from unstructured text through named entity recognition is a core capability across modern automated workflows. However, standard systems rely heavily on massive, human-annotated datasets, making deployment slow and costly when expanding into specialized domains where labeled data is scarce. While few-shot learning methods attempt to train systems using only a handful of examples, current approaches frequently struggle because they overfit to specific source categories and misclassify non-entity words that later become valid target entities.

The article evaluates a new framework named CONTAINER, which demonstrates how contrastive learning over Gaussian distributions can improve few-shot entity recognition across unseen text domains. The main objective is to establish a generalized, class-agnostic representation of language tokens that avoids overfitting and easily adapts to new entity types with minimal target-domain data.

The researchers assessed this approach by pretraining language model representations to separate distinct word categories while pulling identical categories together using probability distributions rather than traditional point embeddings. They evaluated the framework across standard benchmark datasets spanning general text, news, biomedical records, social media, and mixed genres, as well as a large-scale specialized few-shot benchmark. These tests covered both tag-set expansion within the same domain and transfer across entirely different text domains, primarily testing performance under 1-shot and 5-shot data constraints.

The evaluation revealed several key findings. First, CONTAINER outperformed existing state-of-the-art methods by an average of 3% to 13% in absolute F1 score across multiple domains. Second, the system showed its largest advantages in difficult scenarios where previous models struggled, such as cross-domain transfers and datasets with entirely disjoint coarse entity types. Third, fine-tuning on multiple target examples (such as 5-shot setups) yielded substantial performance gains, proving the framework effectively utilizes small support datasets to model class distributions. Fourth, the model learned natural label dependencies directly from contrastive training, rendering additional sequence decoding steps largely unnecessary unless transferring across extreme domain shifts.

These findings indicate that adopting distribution-based contrastive learning significantly reduces the time and cost required to deploy language processing systems in new, data-poor areas. By learning how to differentiate words rather than memorizing domain-specific labels, organizations can reliably adapt existing language models without complex, brittle prompt-engineering or extensive hyperparameter tuning.

For practical implementation, technical teams should consider adopting Gaussian contrastive objectives when developing low-resource entity extraction pipelines. While the approach delivers benchmark-leading few-shot performance, the article cautions that absolute accuracy still trails fully supervised systems trained on extensive manual annotations. Consequently, stakeholders should avoid fully autonomous deployment in high-stakes environments—such as clinical medical record extraction—without human review, and should conduct targeted pilot studies before broad operational rollout.

arXiv: 2109.07589
Cover for CONTaiNER: Few-Shot Named Entity Recognition via Contrastive Learning

Abstract

Named Entity Recognition (NER) in Few-Shot setting is imperative for entity tagging in low resource domains. Existing approaches only learn class-specific semantic features and intermediate representations from source domains. This affects generalizability to unseen target domains, resulting in suboptimal performances. To this end, we present CONTaiNER, a novel contrastive learning technique that optimizes the inter-token distribution distance for Few-Shot NER. Instead of optimizing class-specific attributes, CONTaiNER optimizes a generalized objective of differentiating between token categories based on their Gaussian-distributed embeddings. This effectively alleviates overfitting issues originating from training domains. Our experiments in several traditional test domains (OntoNotes, CoNLL’03, WNUT ’17, GUM) and a new large scale Few-Shot NER dataset (Few-NERD) demonstrate that, on average, CONTaiNER outperforms previous methods by 3%-13% absolute F1 points while showing consistent performance trends, even in challenging scenarios where previous approaches could not achieve appreciable performance. The source code of CONTaiNER will be available at: https://github.com/psunlpgroup/CONTaiNER.

Table of Contents

  • 1 Introduction
  • 2 Task Formulation
  • 3 Method
  • 3.1 Model
  • 3.2 Training in Source Domain
  • 3.3 Finetuning to Target Domain using Support Set
  • 3.4 Instance Level Nearest Neighbor Inference
  • 4 Experiment Setups
  • 4.1 Tag-set Extension Setting
  • 4.2 Domain Transfer Setting
  • 4.3 Few-NERD Setting
  • 5 Results and Analysis
  • 5.1 Overall Results
  • 5.2 Training Objective
  • 5.3 Effect of Model Fine-tuning
  • 5.4 Modeling Label Dependencies
  • 6 Related Works
  • 7 Conclusion
  • Acknowledgement
  • Ethics Statement
  • References
  • A Implementation Details
  • B Fine-tuning Objective
  • C t-SNE Visualization: Point Embedding vs. Gaussian Embedding
  • D Comparison of Different Training Objectives
  • E Embedding Quality: Before vs. After Projection
  • F NER Prediction Examples

Knowls

  1. Knowl 1 — CONTAINER Framework for Few-Shot Named Entity Recognition

    model/method

    CONTAINER is a metric learning framework for few-shot Named Entity Recognition (NER). Conventional supervised few-shot NER models learn class-specific linear classifiers on source domain entity types, which causes representations of non-entity (Outside / O) tokens in the training set to collapse into a single cluster and drop target-relevant features. In contrast, CONTAINER optimizes a generalized contrastive learning objective that models token embeddings as Gaussian distributions rather than deterministic points.

    The framework consists of three stages:

    1. Source Domain Training: A Pretrained Language Model (PLM) encoder processes token sequences into contextual representations hi∈Rl′h_i \in \mathbb{R}^{l'}. Projection heads output Gaussian distribution parameters (mean μi∈Rl\mu_i \in \mathbb{R}^l and diagonal covariance Σi∈Rl×l\Sigma_i \in \mathbb{R}^{l \times l}). The model is optimized using an in-batch contrastive loss based on symmetrized Kullback-Leibler (KL) divergence between token Gaussian distributions.
    2. Target Domain Adaptation (Fine-Tuning): Using a KK-shot support set from the target domain, the PLM and projection heads are fine-tuned using the contrastive loss with early stopping based on the support contrastive loss (patience = 1).
    3. Instance-Level Nearest Neighbor Inference: At test time, the projection heads are discarded, and query token representations from the PLM are classified by assigning the label of the closest support token in the Euclidean representation space.
  2. Knowl 2 — Gaussian Token Embedding and Symmetrized KL-Divergence Metric

    equation

    Given a sequence of nn tokens [x1,x2,…,xn][x_1, x_2, \dots, x_n], a pretrained language model encoder generates contextual representations hi∈Rl′h_i \in \mathbb{R}^{l'} for each token xix_i:

    [h1,h2,…,hn]=PLM([x1,x2,…,xn])[h_1, h_2, \dots, h_n] = \text{PLM}([x_1, x_2, \dots, x_n])

    CONTAINER projects each hidden vector hih_i into the parameters of a multivariate Gaussian distribution N(μi,Σi)\mathcal{N}(\mu_i, \Sigma_i) using two single-layer neural networks fμf_\mu and fΣf_\Sigma preceded by ReLU activations:

    μi=fμ(hi),Σi=ELU(fΣ(hi))+(1+ϵ)I\mu_i = f_\mu(h_i), \quad \Sigma_i = \text{ELU}(f_\Sigma(h_i)) + (1 + \epsilon)I

    where μi∈Rl\mu_i \in \mathbb{R}^l is the mean vector, Σi∈Rl×l\Sigma_i \in \mathbb{R}^{l \times l} is a diagonal covariance matrix, ELU\text{ELU} denotes the Exponential Linear Unit, II is the identity matrix, and ϵ≈e−14\epsilon \approx e^{-14} prevents division by zero (with embedding dimension l=128l = 128).

    For two tokens xpx_p and xqx_q with Gaussian embeddings Np=N(μp,Σp)\mathcal{N}_p = \mathcal{N}(\mu_p, \Sigma_p) and Nq=N(μq,Σq)\mathcal{N}_q = \mathcal{N}(\mu_q, \Sigma_q), the directional KL-divergence DKL[Nq∥Np]D_{\text{KL}}[\mathcal{N}_q \parallel \mathcal{N}_p] is:

    DKL[Nq∥Np]=12(Tr(Σp−1Σq)+(μp−μq)TΣp−1(μp−μq)−l+log⁡∣Σp∣∣Σq∣)D_{\text{KL}}[\mathcal{N}_q \parallel \mathcal{N}_p] = \frac{1}{2} \left( \text{Tr}(\Sigma_p^{-1} \Sigma_q) + (\mu_p - \mu_q)^T \Sigma_p^{-1} (\mu_p - \mu_q) - l + \log \frac{|\Sigma_p|}{|\Sigma_q|} \right)

    The symmetric distributional distance d(p,q)d(p, q) between xpx_p and xqx_q is:

    d(p,q)=12(DKL[Nq∥Np]+DKL[Np∥Nq])d(p, q) = \frac{1}{2} \left( D_{\text{KL}}[\mathcal{N}_q \parallel \mathcal{N}_p] + D_{\text{KL}}[\mathcal{N}_p \parallel \mathcal{N}_q] \right)

  3. Knowl 3 — In-Batch Distributional Contrastive Loss for Token Representations

    equation

    In a sampled minibatch X={(xi,yi)}\mathcal{X} = \{(x_i, y_i)\} containing bb sequences, two tokens xpx_p and xqx_q form a positive pair if they have identical entity tag labels yp=yqy_p = y_q. The set of in-batch positive samples for token pp is defined as:

    Xp={(xq,yq)∈X∣yp=yq, p≠q}\mathcal{X}_p = \{(x_q, y_q) \in \mathcal{X} \mid y_p = y_q, \, p \neq q\}

    The contrastive loss ℓ(p)\ell(p) for token xpx_p pulls representations of identical entity types together while pushing differing categories apart in distribution space:

    ℓ(p)=−log⁡∑(xq,yq)∈Xpexp⁡(−d(p,q))/∣Xp∣∑(xq,yq)∈X, p≠qexp⁡(−d(p,q))\ell(p) = -\log \frac{\sum_{(x_q, y_q) \in \mathcal{X}_p} \exp(-d(p, q)) / |\mathcal{X}_p|}{\sum_{(x_q, y_q) \in \mathcal{X}, \, p \neq q} \exp(-d(p, q))}

    where d(p,q)d(p, q) is the symmetric KL-divergence distance between the Gaussian embeddings of xpx_p and xqx_q. The total batch training loss Ltr\mathcal{L}_{\text{tr}} is computed across all valid tokens in X\mathcal{X}:

    Ltr=1∣X∣∑i∈Xℓ(i)\mathcal{L}_{\text{tr}} = \frac{1}{|\mathcal{X}|} \sum_{i \in \mathcal{X}} \ell(i)

  4. Knowl 4 — CONTAINER Source Training and Target Support Adaptation

    algorithm

    The algorithm trains the model on source domain data Xtr\mathcal{X}_{\text{tr}} and subsequently adapts the representations to the target domain using a small support set Xsup\mathcal{X}_{\text{sup}}.

    Input: Training data Xtr\mathcal{X}_{\text{tr}}, Support data Xsup\mathcal{X}_{\text{sup}}, Projection heads fμ,fΣf_\mu, f_\Sigma, Encoder PLM\text{PLM}
    Output: Adapted encoder PLM\text{PLM}
    // Phase 1: Source Domain Contrastive Training
    for sampled minibatch X⊆Xtr\mathcal{X} \subseteq \mathcal{X}_{\text{tr}} do
        for each token (xi,yi)∈X(x_i, y_i) \in \mathcal{X} do
            hi=PLM(xi)h_i = \text{PLM}(x_i)
            μi=fμ(hi)\mu_i = f_\mu(h_i)
            Σi=ELU(fΣ(hi))+(1+ϵ)I\Sigma_i = \text{ELU}(f_\Sigma(h_i)) + (1 + \epsilon)I
        end for
        for each token (xi,yi)∈X(x_i, y_i) \in \mathcal{X} do
            Xi={(xq,yq)∈X∣yi=yq,q≠i}\mathcal{X}_i = \{(x_q, y_q) \in \mathcal{X} \mid y_i = y_q, q \neq i\}
            Compute ℓ(i)=−log⁡∑q∈Xiexp⁡(−d(i,q))/∣Xi∣∑q∈X,q≠iexp⁡(−d(i,q))\ell(i) = -\log \frac{\sum_{q \in \mathcal{X}_i} \exp(-d(i, q))/|\mathcal{X}_i|}{\sum_{q \in \mathcal{X}, q \neq i} \exp(-d(i, q))}
        end for
        Ltr=1∣X∣∑i∈Xℓ(i)\mathcal{L}_{\text{tr}} = \frac{1}{|\mathcal{X}|} \sum_{i \in \mathcal{X}} \ell(i)
        Update fμ,fΣ,PLMf_\mu, f_\Sigma, \text{PLM} via backpropagation on Ltr\mathcal{L}_{\text{tr}}
    end for
    // Phase 2: Target Domain Fine-Tuning with Early Stopping
    Lprev=∞L_{\text{prev}} = \infty
    Lft=Lprev−1L_{\text{ft}} = L_{\text{prev}} - 1
    while Lft<LprevL_{\text{ft}} < L_{\text{prev}} do
        Lprev=LftL_{\text{prev}} = L_{\text{ft}}
        for each token (xi,yi)∈Xsup(x_i, y_i) \in \mathcal{X}_{\text{sup}} do
            hi=PLM(xi)h_i = \text{PLM}(x_i)
            μi=fμ(hi)\mu_i = f_\mu(h_i)
            Σi=ELU(fΣ(hi))+(1+ϵ)I\Sigma_i = \text{ELU}(f_\Sigma(h_i)) + (1 + \epsilon)I
        end for
        for each token (xi,yi)∈Xsup(x_i, y_i) \in \mathcal{X}_{\text{sup}} do
            Xi={(xq,yq)∈Xsup∣yi=yq,q≠i}\mathcal{X}_i = \{(x_q, y_q) \in \mathcal{X}_{\text{sup}} \mid y_i = y_q, q \neq i\}
            Compute ℓ(i)\ell(i) using support distance function
        end for
        Lft=1∣Xsup∣∑i∈Xsupℓ(i)L_{\text{ft}} = \frac{1}{|\mathcal{X}_{\text{sup}}|} \sum_{i \in \mathcal{X}_{\text{sup}}} \ell(i)
        Update fμ,fΣ,PLMf_\mu, f_\Sigma, \text{PLM} via backpropagation on LftL_{\text{ft}}
    end while
    return PLM\text{PLM} (discard fμ,fΣf_\mu, f_\Sigma)

    During fine-tuning, when multiple shots are available (e.g., 5-shot), d(i,q)d(i, q) uses the symmetrized Gaussian KL-divergence. For 1-shot transfer without prior target knowledge, d′(i,q)=∥μi−μq∥22d'(i, q) = \|\mu_i - \mu_q\|_2^2 (squared Euclidean distance between means) is used instead.

  5. Knowl 5 — Instance-Level Nearest Neighbor Inference and Viterbi Decoding

    model/method

    Following target domain adaptation, projection networks fμf_\mu and fΣf_\Sigma are removed, and inference is conducted directly on the representations h=PLM(x)h = \text{PLM}(x) produced by the encoder before the projection layer, which retain broader semantic information.

    For each test token xitest∈Xtestx_i^{\text{test}} \in \mathcal{X}_{\text{test}} with encoder representation hitesth_i^{\text{test}}, CONTAINER identifies the closest support token xksup∈Xsupx_k^{\text{sup}} \in \mathcal{X}_{\text{sup}} in representation space and assigns its label:

    yitest=arg⁡min⁡yksup where (xksup,yksup)∈Xsup∥hitest−hksup∥22y_i^{\text{test}} = \arg\min_{y_k^{\text{sup}} \text{ where } (x_k^{\text{sup}}, y_k^{\text{sup}}) \in \mathcal{X}_{\text{sup}}} \|h_i^{\text{test}} - h_k^{\text{sup}}\|_2^2

    Viterbi Decoding: When substantial domain shift exists between training and evaluation text, an optional Viterbi decoding step is applied. An abstract transition distribution among abstract tags O\text{O}, I\text{I}, and I-other\text{I-other} is calculated by counting their frequencies in the source training data, and then distributed uniformly across the corresponding target entity tags. Nearest neighbor distances are converted into emission probabilities with temperature τ=0.1\tau = 0.1. When source and target share domain style, in-batch contrastive learning captures label dependencies directly, making Viterbi decoding unnecessary.

  6. Knowl 6 — Tag-Set Extension Performance on OntoNotes 5.0

    data/table

    Tag-set extension assesses few-shot adaptation when novel entity types appear within the same text domain. The 18 entity types in OntoNotes 5.0 are divided into three disjoint subsets of 6 classes each: Group A, Group B, and Group C. Models are trained on two groups (with target group entities labeled as O) and evaluated on the remaining target group using 1-shot or 5-shot support samples. All models use bert-base-cased as the backbone encoder. Results show mean micro-F1 scores (±\pm standard deviation across 5 support sample draws).

    Model 1-shot 5-shot
    Group A Group B Group C Avg. Group A Group B Group C Avg.
    Proto 19.3±3.919.3 \pm 3.9 22.7±8.922.7 \pm 8.9 18.9±7.918.9 \pm 7.9 20.3 30.5±3.530.5 \pm 3.5 38.7±5.638.7 \pm 5.6 41.1±3.341.1 \pm 3.3 36.7
    NNShot 28.5±9.228.5 \pm 9.2 27.3±12.327.3 \pm 12.3 21.4±9.721.4 \pm 9.7 25.7 44.0±2.144.0 \pm 2.1 51.6±5.951.6 \pm 5.9 47.6±2.847.6 \pm 2.8 47.7
    StructShot 30.5±12.330.5 \pm 12.3 28.8±11.228.8 \pm 11.2 20.8±9.920.8 \pm 9.9 26.7 47.5±4.047.5 \pm 4.0 53.0±7.953.0 \pm 7.9 48.7±2.748.7 \pm 2.7 49.8
    CONTAINER 32.2±5.3\mathbf{32.2 \pm 5.3} 30.9±11.6\mathbf{30.9 \pm 11.6} 32.9±12.7\mathbf{32.9 \pm 12.7} 32.0\mathbf{32.0} 51.2±5.9\mathbf{51.2 \pm 5.9} 55.9±6.2\mathbf{55.9 \pm 6.2} 61.5±2.7\mathbf{61.5 \pm 2.7} 56.2\mathbf{56.2}
    + Viterbi 32.4±5.132.4 \pm 5.1 30.9±11.630.9 \pm 11.6 33.0±12.833.0 \pm 12.8 32.1 51.2±6.051.2 \pm 6.0 56.0±6.256.0 \pm 6.2 61.5±2.761.5 \pm 2.7 56.2

    CONTAINER outperforms StructShot by +5.3 absolute F1 points on 1-shot (32.0% vs. 26.7%) and by +6.4 points on 5-shot (56.2% vs. 49.8%). Group C under 5-shot shows an improvement of 12.75 F1 points (61.5% vs. 48.7%).

  7. Knowl 7 — Cross-Domain Transfer Evaluation from OntoNotes to Diverse Datasets

    data/table

    In domain transfer, models are trained on OntoNotes (General domain) and evaluated on target datasets representing distinct text domains: I2B2 (Medical, 23 classes), CoNLL'03 (News, 4 classes), WNUT'17 (Social, 6 classes), and GUM (Mixed text types, 11 classes). Target classes in CoNLL'03 are a subset of OntoNotes types, whereas I2B2, WNUT'17, and GUM represent novel class distributions. All models use bert-base-cased.

    Model 1-shot 5-shot
    I2B2 CoNLL WNUT GUM Avg. I2B2 CoNLL WNUT GUM Avg.
    Proto 13.4 49.9 17.4 17.8 24.6 17.9 61.3 22.8 19.5 30.4
    NNShot 15.3 61.2 22.7 10.5 27.4 22.0 74.1 27.3 15.9 34.8
    StructShot 21.4 62.4 24.2 7.8 29.0 30.3 74.8 30.4 13.3 37.2
    CONTAINER 16.4 57.8 24.2 17.9 29.1 24.1 72.8 27.7 24.4 37.3
    + Viterbi 21.5\mathbf{21.5} 61.2\mathbf{61.2} 27.5\mathbf{27.5} 18.5\mathbf{18.5} 32.2\mathbf{32.2} 36.7\mathbf{36.7} 75.8\mathbf{75.8} 32.5\mathbf{32.5} 25.2\mathbf{25.2} 42.6\mathbf{42.6}

    On the challenging GUM corpus, CONTAINER + Viterbi achieves 25.2% F1 on 5-shot compared to StructShot's 13.3% (+11.9 points). Across all domains, CONTAINER + Viterbi outperforms StructShot by an average of +3.2 F1 points on 1-shot and +5.4 F1 points on 5-shot.

  8. Knowl 8 — Few-Shot NER Benchmark Results on Few-NERD (INTRA and INTER)

    data/table

    The Few-NERD dataset comprises 66 fine-grained entity types across 8 coarse-grained categories. In Few-NERD (INTRA), train, dev, and test splits have mutually exclusive coarse-grained types, preventing any category overlap. In Few-NERD (INTER), coarse-grained categories are shared but fine-grained types are disjoint. All models use bert-base-uncased and evaluate across 5-way and 10-way episodes in 12 shot and 510 shot configurations.

    Few-NERD (INTRA) Few-NERD (INTER)
    Model 5-way 10-way Avg. 5-way 10-way Avg.
    1∼\sim2 5∼\sim10 1∼\sim2 5∼\sim10 1∼\sim2 5∼\sim10 1∼\sim2 5∼\sim10
    ProtoBERT 23.45 41.93 19.76 34.61 29.94 44.44 58.80 39.09 53.97 49.08
    NNShot 31.01 35.74 21.88 27.67 29.08 54.29 50.56 46.98 50.00 50.46
    StructShot 35.92 38.83 25.38 26.39 31.63 57.33 57.16 49.46 49.39 53.34
    CONTAINER 40.43\mathbf{40.43} 53.70\mathbf{53.70} 33.84\mathbf{33.84} 47.49\mathbf{47.49} 43.87\mathbf{43.87} 55.95 61.83\mathbf{61.83} 48.35 57.12\mathbf{57.12} 55.81\mathbf{55.81}
    + Viterbi 40.40 53.71 33.82 47.51 43.86 56.10 61.90 48.36 57.13 55.87

    On the challenging Few-NERD (INTRA) split, CONTAINER outperforms StructShot by +12.24 average F1 points (43.87% vs. 31.63%), with gains exceeding 14 points on 5~10 shot tasks (53.70% vs. 38.83% for 5-way; 47.49% vs. 26.39% for 10-way).

  9. Knowl 9 — Ablation of Training Objectives: Gaussian Embedding vs. Point Embedding

    data/table

    To evaluate the impact of modeling entity distributions via Gaussian embeddings, CONTAINER was compared against point embedding models trained with Cosine similarity and Euclidean distance contrastive objectives on the OntoNotes tag-set extension task. To ensure fair comparison with Gaussian embeddings (ll-dimensional mean and ll-dimensional diagonal covariance with l=128l = 128), point embeddings used dimension 2l=2562l = 256.

    Model Objective 1-shot 5-shot
    Group A Group B Group C Avg. Group A Group B Group C Avg.
    Point Embedding + Cosine 7.73 11.27 15.57 11.52 17.33 30.08 22.51 23.31
    Point Embedding + Euclidean 14.96 13.67 11.12 13.25 25.35 41.56 43.11 36.67
    Gaussian Embedding + KL 32.2 30.9 32.9 32.0 51.2 55.9 61.5 56.2

    Gaussian embedding with symmetric KL-divergence outperforms point embedding with Euclidean distance by +18.75 F1 points in 1-shot (32.0% vs. 13.25%) and by +19.53 points in 5-shot (56.2% vs. 36.67%). Cosine similarity performs worst (11.52% and 23.31%). Point embeddings overfit support examples during fine-tuning by collapsing them into tight point clusters, whereas Gaussian embeddings maintain distribution variance and produce well-separated clusters.

  10. Knowl 10 — Analysis of Support Set Adaptation and Pre-Projection Feature Extraction

    empirical result

    Ablation experiments identify three key operational dynamics in CONTAINER:

    1. Effect of Support Set Fine-Tuning: On the OntoNotes tag-set extension task (evaluating target entities PERSON, DATE, MONEY, LOC, FAC, PRODUCT), fine-tuning with early stopping increases 1-shot F1 from 31.76% to 32.90% (+1.14 points) and 5-shot F1 from 56.99% to 61.48% (+4.49 points), demonstrating that contrastive fine-tuning effectively leverages few support instances without overfitting.
    2. Pre- vs. Post-Projection Representations: Evaluating inference on OntoNotes Group A using hidden states before the projection head (h∈R768h \in \mathbb{R}^{768}) versus parameters after projection (μ∈R128\mu \in \mathbb{R}^{128}) shows that pre-projection representations achieve higher F1 scores: 32.17% vs. 29.21% in 1-shot, and 51.19% vs. 49.78% in 5-shot. This occurs because the layer directly adjacent to the contrastive objective discards general semantic information.
    3. 1-Shot vs. 5-Shot Fine-Tuning Metric: In cross-domain transfer to WNUT'17 where target distributions are unknown, fine-tuning 1-shot models using Euclidean distance between mean embeddings outperforms symmetric KL-divergence (27.48% vs. 18.78% F1) because a single sample cannot reliably estimate class variance. In 5-shot transfer, KL-divergence adaptation outperforms Euclidean distance (32.50% vs. 31.12% F1) as multiple samples enable variance estimation.

Coverage note — Qualitative prediction examples from Appendix F (Table 10) were omitted as their conclusions are fully represented by the quantitative ablation and benchmark tables.

References

  1. 1.Ben Athiwaratkun and Andrew Gordon Wilson. 2018. Hierarchical density order embeddings. arXiv preprint arXiv:1804.09843.
  2. 2.Yujia Bao, Menghua Wu, Shiyu Chang, and Regina Barzilay. 2020. Few-shot text classification with distributional signatures. In ICLR.
  3. 3.Aleksandar Bojchevski and Stephan Günnemann. 2017. Deep gaussian embedding of graphs: Unsupervised inductive learning via ranking. arXiv preprint arXiv:1707.03815.
  4. 4.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In ICML.
  5. 5.Leyang Cui, Yu Wu, Jian Liu, Sen Yang, and Yue Zhang. 2021. Template-based named entity recognition using bart. arXiv preprint arXiv:2106.01760.
  6. 6.Leon Derczynski, Eric Nichols, Marieke van Erp, and Nut Limsopatham. 2017. Results of the wnut2017 shared task on novel and emerging entity recognition. In Proceedings of the 3rd Workshop on Noisy Usergenerated Text, pages 140–147.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  8. 8.Ning Ding, Guangwei Xu, Yulin Chen, Xiaobin Wang, Xu Han, Pengjun Xie, Haitao Zheng, and Zhiyuan Liu. 2021. Few-nerd: A few-shot named entity recognition dataset. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3198–3213.
  9. 9.Carl Doersch and Andrew Zisserman. 2017. Multi-task self-supervised visual learning. In Proceedings of the IEEE International Conference on Computer Vision, pages 2051–2060.
  10. 10.Alexey Dosovitskiy, Jost Tobias Springenberg, Martin Riedmiller, and Thomas Brox. 2014. Discriminative unsupervised feature learning with convolutional neural networks. Advances in neural information processing systems, 27:766–774.
  11. 11.Li Fei-Fei, Rob Fergus, and Pietro Perona. 2006. Oneshot learning of object categories. IEEE transactions on pattern analysis and machine intelligence, 28(4):594–611.
  12. 12.Alexander Fritzler, Varvara Logacheva, and Maksim Kretov. 2019. Few-shot classification in named entity recognition task. In Proceedings of the 34th ACM/SIGAPP Symposium on Applied Computing, pages 993–1000.
  13. 13.Ruiying Geng, Binhua Li, Yongbin Li, Xiaodan Zhu, Ping Jian, and Jian Sun. 2019. Induction networks for few-shot text classification. arXiv preprint arXiv:1902.10482.
  14. 14.Abbas Ghaddar and Philippe Langlais. 2017. Winer: A wikipedia annotated corpus for named entity recognition. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 413–422.
  15. 15.Raia Hadsell, Sumit Chopra, and Yann LeCun. 2006. Dimensionality reduction by learning an invariant mapping. In 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06), volume 2, pages 1735–1742. IEEE.
  16. 16.Xu Han, Hao Zhu, Pengfei Yu, Ziyun Wang, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2018. Fewrel: A large-scale supervised few-shot relation classification dataset with state-of-the-art evaluation. arXiv preprint arXiv:1810.10147.
  17. 17.Yutai Hou, Wanxiang Che, Yongkui Lai, Zhihan Zhou, Yijia Liu, Han Liu, and Ting Liu. 2020. Few-shot slot tagging with collapsed dependency transfer and labelenhanced task-adaptive projection network. arXiv preprint arXiv:2006.05702.
  18. 18.Jiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose, Shobana Balakrishnan, Weizhu Chen, Baolin Peng, Jianfeng Gao, and Jiawei Han. 2020. Few-shot named entity recognition: A comprehensive study. arXiv preprint arXiv:2012.14978.
  19. 19.Zhiheng Huang, Wei Xu, and Kai Yu. 2015. Bidirectional lstm-crf models for sequence tagging. arXiv preprint arXiv:1508.01991.
  20. 20.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. 2020. Supervised contrastive learning. arXiv preprint arXiv:2004.11362.
  21. 21.John Lafferty, Andrew McCallum, and Fernando CN Pereira. 2001. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In ICML.
  22. 22.Brenden Lake, Ruslan Salakhutdinov, Jason Gross, and Joshua Tenenbaum. 2011. One shot learning of simple visual concepts. In Proceedings of the annual meeting of the cognitive science society.
  23. 23.Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016. Neural architectures for named entity recognition. arXiv preprint arXiv:1603.01360.
  24. 24.Xuezhe Ma and Eduard Hovy. 2016. End-to-end sequence labeling via bi-directional lstm-cnns-crf. arXiv preprint arXiv:1603.01354.
  25. 25.Ishan Misra and Laurens van der Maaten. 2020. Selfsupervised learning of pretext-invariant representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6707–6717.
  26. 26.Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021. True few-shot learning with language models. arXiv preprint arXiv:2105.11447.
  27. 27.Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. arXiv preprint arXiv:1802.05365.
  28. 28.Chen Qian, Fuli Feng, Lijie Wen, and Tat-Seng Chua. 2021. Conceptualized and contextualized gaussian embedding. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 13683–13691.
  29. 29.Erik F Sang and Fien De Meulder. 2003. Introduction to the conll-2003 shared task: Language-independent named entity recognition. arXiv preprint cs/0306050.
  30. 30.Jake Snell, Kevin Swersky, and Richard S Zemel. 2017. Prototypical networks for few-shot learning. arXiv preprint arXiv:1703.05175.
  31. 31.Amber Stubbs and Özlem Uzuner. 2015. Annotating longitudinal clinical narratives for de-identification: The 2014 i2b2/uthealth corpus. Journal of biomedical informatics, 58:S20–S29.
  32. 32.Luke Vilnis and Andrew McCallum. 2014. Word representations via gaussian embedding. arXiv preprint arXiv:1412.6623.
  33. 33.Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et al. 2016. Matching networks for one shot learning. NeurIPS.
  34. 34.Yan Wang, Wei-Lun Chao, Kilian Q Weinberger, and Laurens van der Maaten. 2019. Simpleshot: Revisiting nearest-neighbor classification for few-shot learning. arXiv preprint arXiv:1911.04623.
  35. 35.Yaqing Wang, Subhabrata Mukherjee, Haoda Chu, Yuancheng Tu, Ming Wu, Jing Gao, and Ahmed Hassan Awadallah. 2021. Meta self-training for few-shot neural sequence labeling. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pages 1737–1747.
  36. 36.Ralph Weischedel, Martha Palmer, Mitchell Marcus, Eduard Hovy, Sameer Pradhan, Lance Ramshaw, Nianwen Xue, Ann Taylor, Jeff Kaufman, Michelle Franchini, et al. 2013. Ontonotes release 5.0 ldc2013t19. Linguistic Data Consortium, Philadelphia, PA.
  37. 37.Sam Wiseman and Karl Stratos. 2019. Label-agnostic sequence labeling by copying nearest neighbors. arXiv preprint arXiv:1906.04225.
  38. 38.Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin. 2018. Unsupervised feature learning via nonparametric instance discrimination. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3733–3742.
  39. 39.Yi Yang and Arzoo Katiyar. 2020. Simple and effective few-shot named entity recognition with structured nearest neighbor learning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6365–6375.
  40. 40.Mang Ye, Xu Zhang, Pong C Yuen, and Shih-Fu Chang. 2019. Unsupervised embedding learning via invariant and spreading instance feature. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6210–6219.
  41. 41.Amir Zeldes. 2017. The GUM corpus: Creating multilayer resources in the classroom. Language Resources and Evaluation.
  42. 42.Chengxu Zhuang, Alex Lin Zhai, and Daniel Yamins. 2019. Local aggregation for unsupervised learning of visual embeddings. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6002–6012.

Citation

MLA
Das, S. S. S., et al. “CONTaiNER: Few-Shot Named Entity Recognition via Contrastive Learning”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 6338–53, https://doi.org/10.18653/v1/2022.acl-long.439.
APA
Das, S. S. S., Katiyar, A., Passonneau, R. J., & Zhang, R. (2022). CONTaiNER: Few-Shot Named Entity Recognition via Contrastive Learning. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6338–6353. https://doi.org/10.18653/v1/2022.acl-long.439
Chicago
Das, S. S. S., A. Katiyar, R. J. Passonneau, and R. Zhang. 2022. “CONTaiNER: Few-Shot Named Entity Recognition via Contrastive Learning”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6338–53. https://doi.org/10.18653/v1/2022.acl-long.439.
Harvard
Das, S.S.S. et al. (2022) “CONTaiNER: Few-Shot Named Entity Recognition via Contrastive Learning”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6338–6353. Available at: https://doi.org/10.18653/v1/2022.acl-long.439.
Vancouver
1. Das SSS, Katiyar A, Passonneau RJ, Zhang R (2022) CONTaiNER: Few-Shot Named Entity Recognition via Contrastive Learning. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6338–6353

BibTeX

@inproceedings{das-etal-2022-container,
    title = "{CONT}ai{NER}: Few-Shot Named Entity Recognition via Contrastive Learning",
    author = "Das, Sarkar Snigdha Sarathi  and
      Katiyar, Arzoo  and
      Passonneau, Rebecca  and
      Zhang, Rui",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.439/",
    doi = "10.18653/v1/2022.acl-long.439",
    pages = "6338--6353"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/