Consistent Prototype Learning for Few-Shot Continual Relation Extraction
Xiudi ChenHui WuXiaodong Shi
Proposes a consistent prototype learning framework with memory refinement and prompt-based representations to prevent catastrophic forgetting and class confusion in few-shot continual relation extraction.
Modern information extraction systems must continually learn new relationship types from streaming text without forgetting previously learned knowledge. While existing continual learning techniques rely on storing past examples in memory, they struggle when training data is extremely scarce. Prior benchmarks masked this weakness by providing large datasets for initial tasks. In realistic deployment conditions where only a handful of examples per relation are available across all stages, current models suffer from severe prototype distortion—shifts in class representations—and confusion among semantically similar relations, leading to catastrophic forgetting of past knowledge.
The article establishes a rigorous task framework where every training stage contains only a small number of labeled instances per relation. To address this challenge, the authors introduce and evaluate Consistent Prototype Learning, a method designed to preserve old relational knowledge, maintain distribution stability, and distinguish closely related classes.
The evaluated framework integrates prompt learning with a standard language encoder to extract rich semantic representations, combined with a multi-part memory system that stores both representative text samples and fixed class prototype vectors. A three-stage training procedure uses classification and distribution consistency losses to align current predictions with stored representations, alongside a specialized loss function that directs model focus toward distinguishing easily confused relation classes. The approach was tested against established continual learning baselines across two standard relation extraction benchmarks, covering sequences of eight consecutive tasks under varying few-shot conditions.
The experimental results demonstrate that Consistent Prototype Learning significantly outperforms existing approaches. In an eight-task sequence with five examples per class, the proposed method achieved a final overall accuracy of 85.77% on the first benchmark and 76.38% on the second, outperforming the best non-prompt baselines by 26.48% and 41.19%, respectively, and maintaining superiority over prompt-enhanced baselines. The average forgetting rate dropped to 3.31%, closely approaching the theoretical upper-bound performance where all historical data is retained. Ablation testing revealed that the specialized loss for separating confusing relations provided the largest single performance gain of 10.66%, while maintaining prototype memory vectors contributed an additional 3.56% boost and reduced overall variance.
These findings show that tracking and preserving prototypical class vectors directly, rather than relying solely on raw text replays, resolves the primary driver of catastrophic forgetting in data-scarce environments. By mitigating confusion between similar relation categories, systems can scale incrementally without requiring costly historical retraining or massive initial labeling efforts. Organizations deploying ongoing information extraction pipelines can significantly reduce computational retraining overhead and annotation costs by adopting prototype-consistent architectures.
Decision-makers should consider adopting prototype-preserving architectures for production pipelines that require continuous updates with limited annotated data. When deploying such models, practitioners must carefully calibrate loss weights to balance memory preservation against new task learning. Before broad deployment across enterprise text streams, teams should conduct pilot evaluations to assess performance on industry-specific domain shifts, as the current study tested general benchmark corpora. Further investigation is also recommended to optimize the marginal memory storage overhead introduced by saving class prototype vectors.
- Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). iCaRL introduces exemplar-based class prototypes and distillation for incremental learning, providing a direct foundation for understanding how prototype memory can preserve old classes.
- Paper: Learning to Prompt for Continual Learning, Zifeng Wang et al. (2021). Learning to Prompt establishes prompt pools as a way to adapt pretrained transformers across sequential tasks, clarifying the prompt-learning component used by the source.
- Paper: Meta-Learning for Semi-Supervised Few-Shot Classification, Mengye Ren et al. (2018). This work extends Prototypical Networks to few-shot settings, making its class-mean prototype approach useful groundwork for the source’s few-shot relational representations.
No sufficiently relevant recommendations were found.
