MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NER

Ran ZhouXin LiRuidan HeLidong BingErik CambriaLuo SiChunyan Miao

article2022ACL120 citations

Proposes a data augmentation framework that injects entity labels directly into sentence contexts during masked language modeling, preventing token-label misalignment and generating diverse, high-quality synthetic training examples for low-resource named entity recognition.

Listen

Data-driven text analysis systems often depend on named entity recognition to identify key items such as organizations, locations, and individuals. While supervised systems excel with ample training data, manual annotation is prohibitively expensive across many low-resource domains and languages. Standard text augmentation approaches typically struggle in this setting because modifying sentences causes token-label misalignment, where newly generated words no longer match their original entity tags.

The article demonstrates a targeted data augmentation framework called Masked Entity Language Modeling to resolve entity-label misalignment and expand training data in low-resource environments. The primary objective is to evaluate how conditioning text generation directly on entity labels and original sentence contexts improves recognition performance across monolingual, cross-lingual, and multilingual settings.

The researchers developed a workflow that wraps entity mentions with explicit label tags before masking entity tokens for fine-tuning a multilingual language model. The model then predicts diverse replacement entities by drawing on both surrounding context and explicit class labels while keeping the outer sentence intact. In multilingual setups, the authors integrated this technique with a bilingual embedding search algorithm to substitute semantically aligned entities across languages. The approach was evaluated against existing augmentation baselines using standard benchmarks in four languages across varying sample constraints ranging from 100 to 800 examples.

The experimental findings show that the proposed framework consistently outperforms existing augmentation baselines. First, in monolingual low-resource setups, it achieves an absolute performance gain of up to 6.3 points over the best baselines, delivering its largest advantages when data is most scarce at 100 samples. Second, the system substantially outperforms baseline models in zero-shot cross-lingual transfer, reaching an average transfer score of 57.0 at the 100-sample level compared to 38.7 for training without augmentation. Third, ablating the label-linearization step leads to a noticeable performance decline, confirming that conditioning generation directly on labels is critical to preventing invalid entity predictions. Finally, combining label-conditioned masking with semantic entity substitution in multilingual scenarios yields the highest overall performance across all tested resource sizes.

These results indicate that augmenting entity diversity rather than sentence structure offers a reliable, low-risk way to train high-performing models when labeled data is scarce. Preserving natural sentence context avoids the ungrammatical text often generated by fully synthetic models, lowering error risks in downstream applications like search and text extraction without incurring large manual labeling costs.

Organizations operating in data-constrained environments should adopt label-conditioned entity masking to expand their training assets instead of relying on generic word substitutions or full-sentence generation. For multilingual operations, combining this approach with semantic similarity search provides a clear performance boost. Additional analysis across broader non-European languages and domain-specific vocabularies is recommended to confirm boundary conditions before full-scale deployment.

Cover for MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NER

Abstract

Data augmentation is an effective solution to data scarcity in low-resource scenarios. However, when applied to token-level tasks such as NER, data augmentation methods often suffer from token-label misalignment, which leads to unsatsifactory performance. In this work, we propose Masked Entity Language Modeling (MELM) as a novel data augmentation framework for low-resource NER. To alleviate the token-label misalignment issue, we explicitly inject NER labels into sentence context, and thus the fine-tuned MELM is able to predict masked entity tokens by explicitly conditioning on their labels. Thereby, MELM generates high-quality augmented data with novel entities, which provides rich entity regularity knowledge and boosts NER performance. When training data from multiple languages are available, we also integrate MELM with code-mixing for further improvement. We demonstrate the effectiveness of MELM on monolingual, cross-lingual and multilingual NER across various low-resource levels. Experimental results show that our MELM presents substantial improvement over the baseline methods.1

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.1 Labeled Sequence Linearization
  • 2.2 Fine-tuning MELM
  • 2.3 Data Generation
  • 2.4 Post-Processing
  • 2.5 Extending to Multilingual Scenarios
  • 3 Experiments
  • 3.1 Dataset
  • 3.2 Experimental Setting
  • 3.3 Baseline Methods
  • 3.4 Experimental Results
  • 3.4.1 Monolingual and Cross-lingual NER
  • 3.4.2 Multilingual NER
  • 4 Further Analysis
  • 4.1 Case Study
  • 4.2 Number of Unique Entities
  • 5 Related Work
  • 6 Conclusion
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Hyperparameter Tuning
  • A.2 Statistics for Reproducibility
  • A.3 Computing Infrastructure

Knowls

  1. Knowl 1 — Masked Entity Language Modeling (MELM) Data Augmentation Algorithm

    algorithm

    Masked Entity Language Modeling (MELM) is a data augmentation framework designed for low-resource Named Entity Recognition (NER). It addresses the token-label misalignment problem of standard masked language models by explicitly conditioning entity generation on label tokens interleaved directly into the context.

    The procedure takes a labeled training dataset Dtrain\mathcal{D}_{\text{train}}, a pretrained masked language model M\mathcal{M}, a fine-tuning entity masking probability η\eta, a generation Gaussian mean μ\mu, a candidate pool size kk, and an augmentation repetition count RR. It produces an enlarged training set Dtrain∪Daug\mathcal{D}_{\text{train}} \cup \mathcal{D}_{\text{aug}}.

    Input: Gold training set Dtrain\mathcal{D}_{\text{train}}, pretrained MLM M\mathcal{M}, fine-tuning masking rate η\eta, generation mean μ\mu, top-k parameter k=5k=5, augmentation rounds R=3R=3
    Output: Augmented dataset Dtrain∪Daug\mathcal{D}_{\text{train}} \cup \mathcal{D}_{\text{aug}}
    Dmasked←∅\mathcal{D}_{\text{masked}} \leftarrow \emptyset
    Daug←∅\mathcal{D}_{\text{aug}} \leftarrow \emptyset
    for each sentence-label pair (X,Y)∈Dtrain(X, Y) \in \mathcal{D}_{\text{train}} do
        X~←LINEARIZE(X,Y)\tilde{X} \leftarrow \text{LINEARIZE}(X, Y)
        X~←FINETUNEMASK(X~,η)\tilde{X} \leftarrow \text{FINETUNEMASK}(\tilde{X}, \eta)
        Dmasked←Dmasked∪{X~}\mathcal{D}_{\text{masked}} \leftarrow \mathcal{D}_{\text{masked}} \cup \{\tilde{X}\}
    end for
    Mfinetune←FINETUNE(M,Dmasked)\mathcal{M}_{\text{finetune}} \leftarrow \text{FINETUNE}(\mathcal{M}, \mathcal{D}_{\text{masked}})
    for each sentence-label pair (X,Y)∈Dtrain(X, Y) \in \mathcal{D}_{\text{train}} do
        repeat RR times:
            X~←LINEARIZE(X,Y)\tilde{X} \leftarrow \text{LINEARIZE}(X, Y)
            X~←GENMASK(X~,μ)\tilde{X} \leftarrow \text{GENMASK}(\tilde{X}, \mu)
            Xaug←RANDCHOICE(Mfinetune(X~),Top k=5)X_{\text{aug}} \leftarrow \text{RANDCHOICE}(\mathcal{M}_{\text{finetune}}(\tilde{X}), \text{Top } k=5)
            Daug←Daug∪{Xaug}\mathcal{D}_{\text{aug}} \leftarrow \mathcal{D}_{\text{aug}} \cup \{X_{\text{aug}}\}
        end repeat
    end for
    Daug←POSTPROCESS(Daug)\mathcal{D}_{\text{aug}} \leftarrow \text{POSTPROCESS}(\mathcal{D}_{\text{aug}})
    return Dtrain∪Daug\mathcal{D}_{\text{train}} \cup \mathcal{D}_{\text{aug}}

    In this algorithm, LINEARIZE inserts bounding label tokens around every entity mention. FINETUNEMASK masks entity tokens at rate η\eta. GENMASK dynamically samples an entity-specific masking rate ϵ∼N(μ,1/n2)\epsilon \sim \mathcal{N}(\mu, 1/n^2) for an entity of length nn. RANDCHOICE samples replacement tokens uniformly from the top-kk most probable vocabulary tokens predicted by Mfinetune\mathcal{M}_{\text{finetune}} at masked positions and removes inserted label tokens. POSTPROCESS discards generated sentences whose labels predicted by an initial NER model (trained on Dtrain\mathcal{D}_{\text{train}}) do not match the original target labels.

  2. Knowl 2 — Labeled Sequence Linearization for Entity Generation

    model/method

    Labeled Sequence Linearization is a structural formatting method that explicitly injects Named Entity Recognition (NER) label tokens into an input sentence sequence before masked entity prediction. Standard masked language models (MLMs) condition predictions solely on word context, which often generates semantically plausible words that mismatch the target entity category (token-label misalignment).

    Under labeled sequence linearization, a special label delimiter token is inserted immediately before and after each token in an entity mention. For example, given the sequence European Union agreed to the proposal with BIO tags B-ORG I-ORG O O O O, linearization converts it to:

    ⟨B-ORG⟩ European ⟨I-ORG⟩ Union ⟨I-ORG⟩ agreed to the proposal

    The inserted label tokens (such as ⟨B-ORG⟩ and ⟨I-ORG⟩) are treated as standard context tokens in the transformer input sequence. To facilitate language model fine-tuning and preserve natural sentence semantics, the word embeddings of the newly added label tokens are initialized using the embeddings of semantically related vocabulary words (such as using the embedding of the token "organization" for ⟨B-ORG⟩). During generation, the masked language model predicts masked entity positions while attending directly to both surrounding context words and the flanking label tokens. After generation, the inserted label tokens are removed to yield standard text sentences.

  3. Knowl 3 — MELM Training Objective with Entity-Restricted Masking

    equation

    Masked Entity Language Modeling (MELM) fine-tunes a pretrained masked language model M\mathcal{M} parameterized by θ\theta exclusively on entity tokens within linearized sequences, leaving non-entity context tokens unmasked during fine-tuning.

    Given a linearized token sequence X=(x1,x2,…,xn)X = (x_1, x_2, \dots, x_n) and its corrupted counterpart X~\tilde{X} where entity tokens are masked with probability η\eta, the training objective maximizes the log-likelihood of reconstructing the original entity tokens:

    max⁡θlog⁡pθ(X∣X~)≈∑i=1nmilog⁡pθ(xi∣X~)\max_{\theta} \log p_\theta(X \mid \tilde{X}) \approx \sum_{i=1}^{n} m_i \log p_\theta(x_i \mid \tilde{X})

    where:

    • θ\theta denotes the trainable parameters of the masked language model,
    • nn is the total length of the linearized token sequence X~\tilde{X},
    • xix_i is the original ii-th token in XX,
    • mi∈{0,1}m_i \in \{0, 1\} is a binary mask indicator where mi=1m_i = 1 if the ii-th token is a masked entity token, and mi=0m_i = 0 otherwise.
  4. Knowl 4 — Dynamic Gaussian Masking and Top-k Entity Sampling

    model/method

    To prevent generating identical copies of the training set during data augmentation, MELM employs dynamic Gaussian masking and top-kk candidate sampling during the generation phase.

    1. Dynamic Gaussian Masking: For an entity mention consisting of nn tokens, a masking rate ϵ\epsilon is independently drawn from a Gaussian distribution: ϵ∼N(μ,σ2),where σ2=1n2\epsilon \sim \mathcal{N}\left(\mu, \sigma^2\right), \quad \text{where } \sigma^2 = \frac{1}{n^2} Here, μ\mu is a fixed mean parameter (empirically set to 0.50.5). The variance scales inversely with the square of the entity length, ensuring that identical source sentences yield different masked spans across different augmentation passes r∈{1,…,R}r \in \{1, \dots, R\}.

    2. Top-kk Sampling: At each masked position ii, the fine-tuned MELM outputs a probability distribution P(xi∣X~)P(x_i \mid \tilde{X}) over the entire vocabulary VV. To avoid deterministically reproducing the original training entity (the mode of the distribution), candidate selection is restricted to the subset Vik⊆VV_i^k \subseteq V containing the kk most probable tokens (with k=5k=5). The replacement token x^i\hat{x}_i is then sampled at random from VikV_i^k.

  5. Knowl 5 — Bilingual Entity Similarity Search (ESS) Code-Mixing

    model/method

    In multilingual low-resource NER scenarios where training sets {Dtrainℓ∣ℓ∈L}\{\mathcal{D}_{\text{train}}^\ell \mid \ell \in \mathcal{L}\} are available across a set of languages L\mathcal{L}, MELM utilizes Entity Similarity Search (ESS) code-mixing to create semantically coherent cross-lingual training instances.

    Let Eℓ,y\mathcal{E}^{\ell, y} denote the set of entities of label type yy in language ℓ\ell. To substitute an entity mention ee of type yy in a source sentence XℓsrcX^{\ell_{\text{src}}}, a target language is sampled uniformly as ℓtgt∼U(L∖{ℓsrc})\ell_{\text{tgt}} \sim \mathcal{U}(\mathcal{L} \setminus \{\ell_{\text{src}}\}).

    Instead of random substitution, the substitute entity esub∈Eℓtgt,ye_{\text{sub}} \in \mathcal{E}^{\ell_{\text{tgt}}, y} is chosen as the entity with the highest cosine similarity in a shared cross-lingual space. The representation Emb(e)\text{Emb}(e) of entity ee is calculated by averaging the multilingual aligned word embeddings of its constituent tokens using MUSE bilingual embeddings:

    Emb(e)=1∣e∣∑i=1∣e∣MUSEℓsrc,ℓtgt(ei)\text{Emb}(e) = \frac{1}{|e|} \sum_{i=1}^{|e|} \text{MUSE}_{\ell_{\text{src}}, \ell_{\text{tgt}}}(e_i)

    where ∣e∣|e| is the token length of entity ee, MUSEℓsrc,ℓtgt\text{MUSE}_{\ell_{\text{src}}, \ell_{\text{tgt}}} denotes the aligned bilingual embedding space, and eie_i is the ii-th token of ee.

    The most semantically similar replacement entity esube_{\text{sub}} is selected via:

    esub=argmax⁡e~∈Eℓtgt,ycos⁡(Emb(e),Emb(e~))e_{\text{sub}} = \operatorname{argmax}_{\tilde{e} \in \mathcal{E}^{\ell_{\text{tgt}}, y}} \cos\left(\text{Emb}(e), \text{Emb}(\tilde{e})\right)

    When applying MELM fine-tuning on datasets containing ESS code-mixed samples, a language marker (such as ⟨Español⟩) is prepended to the linearized entity tokens to enable the model to distinguish entity language identities.

  6. Knowl 6 — Experimental Setup for Low-Resource NER Evaluation

    experimental setup

    The evaluation benchmark uses the CoNLL-2002 and CoNLL-2003 datasets across four languages: English (En), German (De), Spanish (Es), and Dutch (Nl). Low-resource regimes are simulated by subsampling N∈{100,200,400,800}N \in \{100, 200, 400, 800\} sentences from the full training set (denoted Dtrainℓ,N\mathcal{D}_{\text{train}}^{\ell, N}) and downscaling the development set to NN samples (Ddevℓ,N\mathcal{D}_{\text{dev}}^{\ell, N}), while evaluating on the full test sets (Dtestℓ\mathcal{D}_{\text{test}}^\ell).

    • MELM Generator: Initialized using XLM-RoBERTa-base with a language modeling head (270M parameters). Fine-tuned for 20 epochs using the Adam optimizer with batch size 30 and learning rate 1×10−51 \times 10^{-5}.
    • Target NER Model: XLM-RoBERTa-Large equipped with a Conditional Random Field (CRF) classification head. Trained for 10 epochs using the AdamW optimizer with batch size 16 and learning rate 2×10−52 \times 10^{-5}. The checkpoint with the highest Micro-F1 on Ddevℓ,N\mathcal{D}_{\text{dev}}^{\ell, N} is selected.
    • Hyperparameters: Fine-tuning mask rate η=0.7\eta = 0.7, generation Gaussian mask mean μ=0.5\mu = 0.5, augmentation rounds R=3R = 3, and top-k=5k = 5.
    • Evaluation Metric: Micro-averaged F1 score over 3 independent experimental runs.
  7. Knowl 7 — Monolingual and Cross-Lingual Low-Resource NER Results

    data/table

    The table below compares the Micro-F1 performance of MELM against data augmentation baselines across four low-resource sizes (N∈{100,200,400,800}N \in \{100, 200, 400, 800\}) on monolingual benchmarks (En, De, Es, Nl) and zero-shot cross-lingual transfer benchmarks (En→DeEn \to De, En→EsEn \to Es, En→NlEn \to Nl). Baselines include training only on gold data (Gold-Only), random label-wise entity substitution (Label-wise), entity masking with pretrained MLM without fine-tuning/linearization (MLM-Entity), autoregressive language model augmentation (DAGA), and MELM without labeled sequence linearization (MELM w/o linearize).

    Monolingual Cross-lingual
    #Gold Method En De Es Nl Avg En→\toDe En→\toEs En→\toNl Avg
    100 Gold-Only 50.57 39.47 42.93 21.63 38.65 39.54 37.40 39.27 38.74
    Label-wise 61.34 55.00 59.54 27.85 50.93 45.85 43.74 50.51 46.70
    MLM-Entity 61.22 50.96 61.29 46.59 55.02 47.96 45.42 49.34 47.57
    DAGA 68.06 59.15 69.33 45.64 60.54 52.95 46.72 54.63 51.43
    MELM w/o linearize 70.01 61.92 65.07 59.76 64.19 48.70 49.10 53.37 50.39
    MELM (Ours) 75.21 64.12 75.85 66.57 70.44 56.56 53.83 60.62 57.00
    200 Gold-Only 74.64 62.85 72.64 55.96 66.52 54.95 51.26 60.71 55.64
    Label-wise 76.82 67.31 78.34 66.52 72.25 55.01 53.14 63.30 57.15
    MLM-Entity 79.16 70.01 78.45 66.69 73.58 60.44 57.72 68.37 62.18
    DAGA 79.11 69.82 78.95 68.53 74.10 59.58 57.68 65.74 61.00
    MELM w/o linearize 81.77 71.41 80.43 72.92 76.63 62.57 63.49 70.18 65.41
    MELM (Ours) 82.91 72.71 80.46 77.02 78.27 65.01 63.71 70.37 66.36
    400 Gold-Only 81.85 70.77 80.02 74.60 76.81 65.76 61.57 71.04 66.12
    Label-wise 84.62 74.33 81.01 77.87 79.46 66.18 67.43 71.93 68.51
    MLM-Entity 83.82 74.66 81.08 77.90 79.37 67.41 70.28 74.31 70.67
    DAGA 84.36 72.95 82.83 78.99 79.78 66.77 67.13 72.40 68.77
    MELM w/o linearize 85.16 75.42 82.34 79.34 80.56 68.02 66.01 72.98 69.00
    MELM (Ours) 85.73 77.50 83.31 80.92 81.87 68.08 70.37 75.78 71.74
    800 Gold-Only 86.35 78.35 83.23 83.86 82.95 65.31 68.28 72.07 68.55
    Label-wise 86.72 78.21 84.42 84.26 83.40 65.60 72.22 74.77 70.86
    MLM-Entity 86.50 78.30 84.09 83.93 83.20 65.42 69.10 74.85 69.79
    DAGA 86.61 77.66 84.64 84.90 83.45 68.76 70.97 75.02 71.58
    MELM w/o linearize 87.35 78.58 84.59 84.94 83.99 67.37 71.53 75.20 71.37
    MELM (Ours) 87.59 79.32 85.40 85.17 84.37 67.95 75.72 75.25 72.97

    MELM achieves the highest average F1 across all resource tiers. In the most constrained setting (N=100N=100), MELM outperforms the strongest baseline (MELM w/o linearize) by +6.25+6.25 F1 points on monolingual benchmarks (70.4470.44 vs. 64.1964.19) and exceeds Gold-Only by +31.79+31.79 points (70.4470.44 vs. 38.6538.65). In zero-shot cross-lingual transfer, MELM improves over DAGA by +5.57+5.57 points on average at N=100N=100 (57.0057.00 vs. 51.4351.43).

  8. Knowl 8 — Multilingual Low-Resource NER Performance with Code-Mixing

    data/table

    The table below reports test Micro-F1 across En, De, Es, and Nl in multilingual training settings where N∈{100,200,400}N \in \{100, 200, 400\} gold examples per language are pooled. The evaluated methods are:

    • Gold-Only: Concatenated gold data from all four languages.
    • MulDA: Generation using an mBART model fine-tuned on linearized NER data.
    • MELM-gold: MELM applied directly to the multilingual gold training sets.
    • Code-Mix-random: Random entity replacement across languages within the same entity type.
    • Code-Mix-ess: Entity substitution guided by MUSE bilingual embedding cosine similarity.
    • MELM (Ours): MELM applied to the union of gold training data and Code-Mix-ess data.
    #Gold Method En De Es Nl Avg
    100 ×4\times 4 Gold-Only 75.62 69.35 75.85 74.33 73.79
    MulDA 73.67 70.47 75.53 72.40 73.02
    MELM-gold (Ours) 78.71 74.79 81.25 78.85 78.40
    Code-Mix-random 77.38 70.58 78.61 76.45 75.75
    Code-Mix-ess (Ours) 79.55 71.56 79.58 76.49 76.80
    MELM (Ours) 80.96 75.61 81.47 80.14 79.54
    200 ×4\times 4 Gold-Only 83.06 76.39 82.71 79.19 80.34
    MulDA 82.32 74.57 82.73 79.06 79.67
    MELM-gold (Ours) 82.90 78.05 85.93 81.00 81.97
    Code-Mix-random 82.86 75.70 83.13 79.08 80.19
    Code-Mix-ess (Ours) 83.34 76.64 82.02 82.27 81.07
    MELM (Ours) 83.56 78.24 84.98 82.79 82.39
    400 ×4\times 4 Gold-Only 83.92 77.40 83.22 84.04 82.14
    MulDA 84.37 78.41 84.54 83.09 82.60
    MELM-gold (Ours) 86.04 79.09 85.76 84.83 83.93
    Code-Mix-random 85.04 77.91 84.44 83.56 82.74
    Code-Mix-ess (Ours) 85.74 80.03 85.18 85.36 84.08
    MELM (Ours) 86.14 80.33 86.60 85.99 84.76

    Code-Mix-ess consistently outperforms Code-Mix-random across all low-resource tiers (76.8076.80 vs. 75.7575.75 at 100×4100\times 4; 81.0781.07 vs. 80.1980.19 at 200×4200\times 4; 84.0884.08 vs. 82.7482.74 at 400×4400\times 4). Applying MELM on top of Code-Mix-ess achieves the highest overall performance, improving average F1 over Gold-Only by +5.75+5.75, +2.05+2.05, and +2.62+2.62 points across the three low-resource levels.

  9. Knowl 9 — Label-Consistency and Unique Entity Count in Generated Data

    empirical result

    Analysis of generated entity tokens demonstrates how labeled sequence linearization controls token-label alignment and injects novel entity regularity knowledge:

    1. Prediction Label-Consistency: When predicting masked entity slots, unconstrained pretrained MLMs predominantly produce high-frequency non-entity tokens (e.g., "the", "he", "she"). Fine-tuning MELM without linearization produces named entities but fails to match the original class under ambiguous context (for example, generating "Pompeo" [PER] for an ORG slot). MELM with linearization strictly conditions generation on label tokens, successfully generating novel entities matching the target class (such as generating "Greenpeace" and "Amnesty" for ORG, or "French" and "British" for MISC).

    2. Oracle Filtering and Valid Unique Entity Counts: To quantify entity diversity, an 'oracle' NER tagger trained on the full CoNLL dataset was used to filter generated instances with misaligned labels. Generated instances from MLM-Entity, DAGA, and MELM without linearization experienced severe filtering due to label misalignment. In contrast, MELM generated significantly more unique entities that passed oracle verification across all sample sizes (N=100,200,400,800N=100, 200, 400, 800), reaching approximately 2,300 valid unique entities at N=800N=800 compared to under 1,600 for DAGA and MELM w/o linearize.

  10. Knowl 10 — Hyperparameter Sensitivity of Masking Rates and Augmentation Rounds

    data/table

    Grid searches on the English CoNLL development set were conducted to determine the optimal fine-tuning entity mask rate η∈{0.3,0.5,0.7}\eta \in \{0.3, 0.5, 0.7\}, generation dynamic mask mean μ∈{0.3,0.5,0.7}\mu \in \{0.3, 0.5, 0.7\}, and the number of augmentation rounds R∈{1,2,3,4,5,6}R \in \{1, 2, 3, 4, 5, 6\}.

    η\eta
    μ\mu 0.3 0.5 0.7
    0.3 76.90 75.64 78.08
    0.5 76.16 78.06 78.56
    0.7 75.94 78.09 78.37
    RR 1 2 3 4 5 6
    Dev F1 92.35 92.36 92.84 92.72 92.59 92.39

    Performance peaks at η=0.7\eta = 0.7 and μ=0.5\mu = 0.5, yielding a validation F1 of 78.5678.56. For augmentation rounds RR, performance improves from R=1R=1 (92.3592.35) up to R=3R=3 (92.8492.84). Beyond R=3R=3, performance drops progressively (92.7292.72 at R=4R=4, 92.3992.39 at R=6R=6) due to the amplification of synthetic noise.

Coverage note — None was omitted; all key algorithmic mechanisms, mathematical formulations, multilingual ESS code-mixing extensions, comprehensive empirical tables (monolingual, cross-lingual, multilingual), hyperparameter analyses, and label-consistency evaluations are fully captured.

References

  1. 1.Partha Sarathy Banerjee, Baisakhi Chakraborty, Deepak Tripathi, Hardik Gupta, and Sourabh S Kumar. 2019. A information retrieval based on question and answering and ner for unstructured information without using sql. Wireless Personal Communications, 108(3):1909–1931.
  2. 2.M Saiful Bari, Tasnim Mohiuddin, and Shafiq Joty. 2021. UXLA: A robust unsupervised data augmentation framework for zero-resource cross-lingual NLP. In Proceedings of the Annual Meeting of the Association for Computational Linguistics and the International Joint Conference on Natural Language Processing, pages 1978–1992.
  3. 3.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  4. 4.Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2017. Word translation without parallel data. arXiv preprint arXiv:1710.04087.
  5. 5.Ryan Cotterell and Kevin Duh. 2017. Low-resource named entity recognition with cross-lingual, character-level neural conditional random fields. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 91–96, Taipei, Taiwan. Asian Federation of Natural Language Processing.
  6. 6.Xiang Dai and Heike Adel. 2020. An analysis of simple data augmentation for named entity recognition. In Proceedings of the 28th International Conference on Computational Linguistics, pages 3861–3867, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  7. 7.Bosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai, Thien Hai Nguyen, Shafiq Joty, Luo Si, and Chunyan Miao. 2020. DAGA: Data augmentation with a generation approach for low-resource tagging tasks. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6045–6057, Online. Association for Computational Linguistics.
  8. 8.Li Dong, Jonathan Mallinson, Siva Reddy, and Mirella Lapata. 2017. Learning to paraphrase for question answering. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 875–886, Copenhagen, Denmark. Association for Computational Linguistics.
  9. 9.Alexander Fabbri, Patrick Ng, Zhiguo Wang, Ramesh Nallapati, and Bing Xiang. 2020. Template-based question generation from retrieved sentences for improved unsupervised question answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4508–4513, Online. Association for Computational Linguistics.
  10. 10.Marzieh Fadaee, Arianna Bisazza, and Christof Monz. 2017. Data augmentation for low-resource neural machine translation. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 567–573, Vancouver, Canada. Association for Computational Linguistics.
  11. 11.Xiaocheng Feng, Xiachong Feng, Bing Qin, Zhangyin Feng, and Ting Liu. 2018. Improving low resource named entity recognition using cross-lingual knowledge transfer. In Proceedings of the International Joint Conference on Artificial Intelligence, IJCAI-18, pages 4071–4077.
  12. 12.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  13. 13.Sosuke Kobayashi. 2018. Contextual augmentation: Data augmentation by words with paradigmatic relations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 452–457, New Orleans, Louisiana. Association for Computational Linguistics.
  14. 14.Varun Kumar, Ashutosh Choudhary, and Eunah Cho. 2020. Data augmentation using pre-trained transformer models. In Proceedings of the 2nd Workshop on Life-long Learning for Spoken Language Systems, pages 18–26.
  15. 15.Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016. Neural architectures for named entity recognition. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 260–270, San Diego, California. Association for Computational Linguistics.
  16. 16.Kun Li, Chengbo Chen, Xiaojun Quan, Qing Ling, and Yan Song. 2020a. Conditional augmentation for aspect term extraction via masked sequence-to-sequence generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7056–7066, Online. Association for Computational Linguistics.
  17. 17.Xin Li, Lidong Bing, Wenxuan Zhang, Zheng Li, and Wai Lam. 2020b. Unsupervised cross-lingual adaptation for sequence tagging and beyond. arXiv preprint arXiv:2010.12405.
  18. 18.Hongyu Lin, Yaojie Lu, Jialong Tang, Xianpei Han, Le Sun, Zhicheng Wei, and Nicholas Jing Yuan. 2020. A rigorous study on named entity recognition: Can fine-tuning pretrained model lead to the promised land? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7291–7300, Online. Association for Computational Linguistics.
  19. 19.Linlin Liu, Bosheng Ding, Lidong Bing, Shafiq Joty, Luo Si, and Chunyan Miao. 2021. MulDA: A multilingual data augmentation framework for low-resource cross-lingual NER. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5834–5846, Online. Association for Computational Linguistics.
  20. 20.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742.
  21. 21.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In International Conference on Learning Representations.
  22. 22.Ramesh Nallapati, Bowen Zhou, Caglar Gulcehre, Bing Xiang, et al. 2016. Abstractive text summarization using sequence-to-sequence rnns and beyond. arXiv preprint arXiv:1602.06023.
  23. 23.Libo Qin, Minheng Ni, Yue Zhang, and Wanxiang Che. 2020. Cosda-ml: Multi-lingual code-switching data augmentation for zero-shot cross-lingual nlp.
  24. 24.Shruti Rijhwani, Shuyan Zhou, Graham Neubig, and Jaime Carbonell. 2020. Soft gazetteers for low-resource named entity recognition. arXiv preprint arXiv:2005.01866.
  25. 25.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Improving neural machine translation models with monolingual data. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 86–96, Berlin, Germany. Association for Computational Linguistics.
  26. 26.Jasdeep Singh, Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. 2019. Xlda: Cross-lingual data augmentation for natural language inference and question answering. arXiv preprint arXiv:1905.11471.
  27. 27.Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2019. Mass: Masked sequence to sequence pre-training for language generation. In International Conference on Machine Learning, pages 5926–5936. PMLR.
  28. 28.Erik F. Tjong Kim Sang. 2002. Introduction to the CoNLL-2002 shared task: Language-independent named entity recognition. In COLING-02: The 6th Conference on Natural Language Learning 2002 (CoNLL-2002).
  29. 29.Erik F. Tjong Kim Sang and Fien De Meulder. 2003. Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition. In Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003, pages 142–147.
  30. 30.Chen-Tse Tsai, Stephen Mayhew, and Dan Roth. 2016. Cross-lingual named entity recognition via wikification. In Proceedings of The 20th SIGNLL Conference on Computational Natural Language Learning, pages 219–228.
  31. 31.Jason Wei and Kai Zou. 2019. EDA: Easy data augmentation techniques for boosting performance on text classification tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 6382–6388, Hong Kong, China. Association for Computational Linguistics.
  32. 32.Xing Wu, Shangwen Lv, Liangjun Zang, Jizhong Han, and Songlin Hu. 2019. Conditional bert contextual augmentation. In International Conference on Computational Science, pages 84–95. Springer.
  33. 33.Adams Wei Yu, David Dohan, Quoc Le, Thang Luong, Rui Zhao, and Kai Chen. 2018. Fast and accurate reading comprehension by combining self-attention and convolution. In International Conference on Learning Representations.
  34. 34.Wenxuan Zhang, Ruidan He, Haiyun Peng, Lidong Bing, and Wai Lam. 2021. Cross-lingual aspect-based sentiment analysis with aspect term code-switching. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9220–9230, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  35. 35.Xiaoshi Zhong and Erik Cambria. 2021. Time Expression and Named Entity Recognition. Springer.
  36. 36.Xiaoshi Zhong, Erik Cambria, and Amir Hussain. 2020. Extracting time expressions and named entities with constituent-based tagging schemes. Cognitive Computation, 12(4):844–862.
  37. 37.Joey Tianyi Zhou, Hao Zhang, Di Jin, Hongyuan Zhu, Meng Fang, Rick Siow Mong Goh, and Kenneth Kwok. 2019. Dual adversarial neural transfer for low-resource named entity recognition. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3461–3471, Florence, Italy. Association for Computational Linguistics.

Citation

MLA
Zhou, R., et al. “MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NER”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 2251–62, https://doi.org/10.18653/v1/2022.acl-long.160.
APA
Zhou, R., Li, X., He, R., Bing, L., Cambria, E., Si, L., & Miao, C. (2022). MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NER. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2251–2262. https://doi.org/10.18653/v1/2022.acl-long.160
Chicago
Zhou, R., X. Li, R. He, et al. 2022. “MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NER”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2251–62. https://doi.org/10.18653/v1/2022.acl-long.160.
Harvard
Zhou, R. et al. (2022) “MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NER”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2251–2262. Available at: https://doi.org/10.18653/v1/2022.acl-long.160.
Vancouver
1. Zhou R, Li X, He R, Bing L, Cambria E, Si L, Miao C (2022) MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NER. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 2251–2262

BibTeX

@inproceedings{zhou-etal-2022-melm,
    title = "{MELM}: Data Augmentation with Masked Entity Language Modeling for Low-Resource {NER}",
    author = "Zhou, Ran  and
      Li, Xin  and
      He, Ruidan  and
      Bing, Lidong  and
      Cambria, Erik  and
      Si, Luo  and
      Miao, Chunyan",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.160/",
    doi = "10.18653/v1/2022.acl-long.160",
    pages = "2251--2262"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/