MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NER
Ran ZhouXin LiRuidan HeLidong BingErik CambriaLuo SiChunyan Miao
Proposes a data augmentation framework that injects entity labels directly into sentence contexts during masked language modeling, preventing token-label misalignment and generating diverse, high-quality synthetic training examples for low-resource named entity recognition.
Data-driven text analysis systems often depend on named entity recognition to identify key items such as organizations, locations, and individuals. While supervised systems excel with ample training data, manual annotation is prohibitively expensive across many low-resource domains and languages. Standard text augmentation approaches typically struggle in this setting because modifying sentences causes token-label misalignment, where newly generated words no longer match their original entity tags.
The article demonstrates a targeted data augmentation framework called Masked Entity Language Modeling to resolve entity-label misalignment and expand training data in low-resource environments. The primary objective is to evaluate how conditioning text generation directly on entity labels and original sentence contexts improves recognition performance across monolingual, cross-lingual, and multilingual settings.
The researchers developed a workflow that wraps entity mentions with explicit label tags before masking entity tokens for fine-tuning a multilingual language model. The model then predicts diverse replacement entities by drawing on both surrounding context and explicit class labels while keeping the outer sentence intact. In multilingual setups, the authors integrated this technique with a bilingual embedding search algorithm to substitute semantically aligned entities across languages. The approach was evaluated against existing augmentation baselines using standard benchmarks in four languages across varying sample constraints ranging from 100 to 800 examples.
The experimental findings show that the proposed framework consistently outperforms existing augmentation baselines. First, in monolingual low-resource setups, it achieves an absolute performance gain of up to 6.3 points over the best baselines, delivering its largest advantages when data is most scarce at 100 samples. Second, the system substantially outperforms baseline models in zero-shot cross-lingual transfer, reaching an average transfer score of 57.0 at the 100-sample level compared to 38.7 for training without augmentation. Third, ablating the label-linearization step leads to a noticeable performance decline, confirming that conditioning generation directly on labels is critical to preventing invalid entity predictions. Finally, combining label-conditioned masking with semantic entity substitution in multilingual scenarios yields the highest overall performance across all tested resource sizes.
These results indicate that augmenting entity diversity rather than sentence structure offers a reliable, low-risk way to train high-performing models when labeled data is scarce. Preserving natural sentence context avoids the ungrammatical text often generated by fully synthetic models, lowering error risks in downstream applications like search and text extraction without incurring large manual labeling costs.
Organizations operating in data-constrained environments should adopt label-conditioned entity masking to expand their training assets instead of relying on generic word substitutions or full-sentence generation. For multilingual operations, combining this approach with semantic similarity search provides a clear performance boost. Additional analysis across broader non-European languages and domain-specific vocabularies is recommended to confirm boundary conditions before full-scale deployment.
- Paper: EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks, Jason Wei et al. (2019). Introduces foundational rule-based text data augmentation techniques whose limitations in entity-label alignment motivate MELM's label-conditioned masked modeling.
- Paper: SpanBERT: Improving Pre-training by Representing and Predicting Spans, Mandar Joshi et al. (2019). Establishes span-level masking and boundary-conditioned prediction techniques that underpin MELM's approach to predicting masked multi-token entities in context.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). Analyzes zero-shot cross-lingual sequence labeling capabilities in multilingual BERT, providing the essential cross-lingual foundation evaluated in MELM.
- Paper: Exploiting Cloze-Questions for Few-Shot Text Classification and Natural Language Inference, Timo Schick et al. (2020). Demonstrates how converting NLP inputs into label-conditioned cloze tasks unlocks low-resource transfer in pretrained language models, directly informing MELM's label-linearization strategy.
- Paper: Unsupervised Data Augmentation for Consistency Training, Qizhe Xie et al. (2020). Presents consistency-driven data augmentation principles for low-resource NLP that motivate generating semantically aligned entity variations.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). Introduces cross-lingual language model pretraining paradigms and translation language modeling that enable multilingual entity substitution and transfer.
- Paper: Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition, Erik F. Tjong Kim Sang et al. (2003). Defines the standard multilingual named entity recognition benchmarks, tag formats, and evaluation metrics used to measure MELM's performance.
- Paper: Good Examples Make A Faster Learner: Simple Demonstration-based Learning for Low-resource NER, Dong-Ho Lee et al. (2022). Explores an alternative in-context demonstration paradigm to overcome data scarcity in low-resource NER without synthetic data generation.
- Paper: GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer, Urchade Zaratiana et al. (2024). Extends low-resource named entity recognition to open-vocabulary zero-shot entity extraction using span-prompt matching representations.
- Paper: Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions, John Joon Young Chung et al. (2023). Builds on generative text augmentation for low-resource tasks by studying how to balance synthetic text diversity and label accuracy using LLMs.
- Paper: Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages, Ayyoob Imani et al. (2023). Scales low-resource multilingual representation and sequence labeling to hundreds of tail languages beyond the standard multilingual benchmark settings.
