Consistent Prototype Learning for Few-Shot Continual Relation Extraction

Xiudi ChenHui WuXiaodong Shi

article2023ACL45 citations

Proposes a consistent prototype learning framework with memory refinement and prompt-based representations to prevent catastrophic forgetting and class confusion in few-shot continual relation extraction.

Listen

Modern information extraction systems must continually learn new relationship types from streaming text without forgetting previously learned knowledge. While existing continual learning techniques rely on storing past examples in memory, they struggle when training data is extremely scarce. Prior benchmarks masked this weakness by providing large datasets for initial tasks. In realistic deployment conditions where only a handful of examples per relation are available across all stages, current models suffer from severe prototype distortion—shifts in class representations—and confusion among semantically similar relations, leading to catastrophic forgetting of past knowledge.

The article establishes a rigorous task framework where every training stage contains only a small number of labeled instances per relation. To address this challenge, the authors introduce and evaluate Consistent Prototype Learning, a method designed to preserve old relational knowledge, maintain distribution stability, and distinguish closely related classes.

The evaluated framework integrates prompt learning with a standard language encoder to extract rich semantic representations, combined with a multi-part memory system that stores both representative text samples and fixed class prototype vectors. A three-stage training procedure uses classification and distribution consistency losses to align current predictions with stored representations, alongside a specialized loss function that directs model focus toward distinguishing easily confused relation classes. The approach was tested against established continual learning baselines across two standard relation extraction benchmarks, covering sequences of eight consecutive tasks under varying few-shot conditions.

The experimental results demonstrate that Consistent Prototype Learning significantly outperforms existing approaches. In an eight-task sequence with five examples per class, the proposed method achieved a final overall accuracy of 85.77% on the first benchmark and 76.38% on the second, outperforming the best non-prompt baselines by 26.48% and 41.19%, respectively, and maintaining superiority over prompt-enhanced baselines. The average forgetting rate dropped to 3.31%, closely approaching the theoretical upper-bound performance where all historical data is retained. Ablation testing revealed that the specialized loss for separating confusing relations provided the largest single performance gain of 10.66%, while maintaining prototype memory vectors contributed an additional 3.56% boost and reduced overall variance.

These findings show that tracking and preserving prototypical class vectors directly, rather than relying solely on raw text replays, resolves the primary driver of catastrophic forgetting in data-scarce environments. By mitigating confusion between similar relation categories, systems can scale incrementally without requiring costly historical retraining or massive initial labeling efforts. Organizations deploying ongoing information extraction pipelines can significantly reduce computational retraining overhead and annotation costs by adopting prototype-consistent architectures.

Decision-makers should consider adopting prototype-preserving architectures for production pipelines that require continuous updates with limited annotated data. When deploying such models, practitioners must carefully calibrate loss weights to balance memory preservation against new task learning. Before broad deployment across enterprise text streams, teams should conduct pilot evaluations to assess performance on industry-specific domain shifts, as the current study tested general benchmark corpora. Further investigation is also recommended to optimize the marginal memory storage overhead introduced by saving class prototype vectors.

  • Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). iCaRL introduces exemplar-based class prototypes and distillation for incremental learning, providing a direct foundation for understanding how prototype memory can preserve old classes.
  • Paper: Learning to Prompt for Continual Learning, Zifeng Wang et al. (2021). Learning to Prompt establishes prompt pools as a way to adapt pretrained transformers across sequential tasks, clarifying the prompt-learning component used by the source.
  • Paper: Meta-Learning for Semi-Supervised Few-Shot Classification, Mengye Ren et al. (2018). This work extends Prototypical Networks to few-shot settings, making its class-mean prototype approach useful groundwork for the source’s few-shot relational representations.

No sufficiently relevant recommendations were found.

Cover for Consistent Prototype Learning for Few-Shot Continual Relation Extraction

Abstract

Few-shot continual relation extraction aims to continually train a model on incrementally few-shot data to learn new relations while avoiding forgetting old ones. However, current memory-based methods are prone to overfitting memory samples, resulting in insufficient activation of old relations and limited ability to handle the confusion of similar classes. In this paper, we design a new N-way-K-shot Continual Relation Extraction (NK-CRE) task and propose a novel few-shot continual relation extraction method with Consistent Prototype Learning (ConPL) to address the aforementioned issues. Our proposed ConPL is mainly composed of three modules: 1) a prototype-based classification module that provides primary relation predictions under few-shot continual learning; 2) a memory-enhanced module designed to select vital samples and refined prototypical representations as a novel multi-information episodic memory; 3) a consistent learning module to reduce catastrophic forgetting by enforcing distribution consistency. To effectively mitigate catastrophic forgetting, ConPL ensures that the samples and prototypes in the episodic memory remain consistent in terms of classification and distribution. Additionally, ConPL uses prompt learning to extract better representations and adopts a focal loss to alleviate the confusion of similar classes. Experimental results on two commonly-used datasets show that our model consistently outperforms other competitive baselines¹.

Knowls

  1. Knowl 1 — N-way-K-shot continual relation extraction

    definition

    NK-CRE is a continual relation extraction setting in which a model learns a sequence of tasks, with each task providing only KK labeled training examples for each of its NN relation classes. Each example is a sentence containing a head–tail entity pair and a relation label. After learning task kk, the model is evaluated on the combined test sets of tasks 11 through kk, so evaluation measures both learning the new relations and retaining the earlier ones. Unlike the CFRL setup discussed in the paper, NK-CRE keeps every task few-shot rather than allowing the first task to have substantially more labeled examples.

  2. Knowl 2 — Prompted prototype classifier

    model/method

    ConPL represents a relation example with a BERT encoder applied to a prompted input containing the head entity, a mask token, the tail entity, and the original sentence: [CLS],eh,[MASK],et,[SEP],x,[SEP][\mathrm{CLS}], e_h, [\mathrm{MASK}], e_t, [\mathrm{SEP}], x, [\mathrm{SEP}]. The contextual vector at [MASK][\mathrm{MASK}] is the example representation h=fθ(x)h=f_\theta(x), where xx denotes the prompted sequence and θ\theta denotes the encoder parameters. For each new relation, ConPL initializes a prototype by averaging the representations of that relation’s training examples. It classifies an example among all previously learned and current relations by applying a softmax to cosine similarities between its representation and the corresponding prototypes; cross-entropy on these probabilities supplies the basic classification loss. Current-task prototypes are temporary during initial training, while previously learned prototypes come from memory.

  3. Knowl 3 — Multi-information episodic memory

    model/method

    ConPL stores two kinds of information for each learned relation: a labeled sample and a prototype representation. For a new relation, it first computes the mean representation of that relation’s training examples, then selects the example closest to that mean as its single key sample. After selecting the key sample, ConPL uses its encoder representation to initialize the relation’s stored prototype; the prototype is then updated during training. The sample memory supports replay, while the prototype memory preserves a class representation for classification without recomputing it from a small set of replay examples. The added prototype memory requires one additional vector per relation.

  4. Knowl 4 — Classification and distribution consistency losses

    equation

    ConPL uses two consistency constraints to keep stored examples and relation prototypes aligned. The classification consistency loss pulls each replayed example toward the stored prototype for its labeled relation. The distribution consistency loss compares the example’s current similarity distribution over stored prototypes with the similarity distribution of its labeled prototype over that same prototype bank:

    Lcc=∑(xi,yi)∈S^k−1∥fθ(xi)−pyi∥,Ldc=∑(xi,yi)∈S^k∥d(fθ(xi),P^k)−d(pyi,P^k)∥.L_{cc}=\sum_{(x_i,y_i)\in\hat S^{k-1}}\|f_\theta(x_i)-p_{y_i}\|, \qquad L_{dc}=\sum_{(x_i,y_i)\in\hat S^k}\left\|d(f_\theta(x_i),\hat P^k)-d(p_{y_i},\hat P^k)\right\|.

    Here, S^k−1\hat S^{k-1} is the sample memory from tasks before task kk, S^k\hat S^k is the memory including task kk, pyip_{y_i} is the stored prototype for example xix_i’s label, and P^k\hat P^k is the prototype memory through task kk. The function d(v,P^k)d(v,\hat P^k) denotes the vector of cosine similarities from vector vv to the prototypes in that bank. The first loss constrains example-to-prototype classification alignment; the second constrains the broader pattern of relations to remain aligned. The distribution consistency loss is applied in a memory-only training stage.

  5. Knowl 5 — Focal loss for confusing relation classes

    model/method

    To address confusion between similar relations, ConPL computes an additional classification distribution over a restricted set of prototypes for each example. This set contains the target relation’s prototype, the highest-similarity negative prototype, and other negative prototypes whose cosine similarity is within threshold α\alpha of the target prototype’s similarity. If CiC_i is this set for example xix_i, the probability assigned to its true relation is qi=exp⁡(cos⁡(hi,pyi))/∑p∈Ciexp⁡(cos⁡(hi,p))q_i=\exp(\cos(h_i,p_{y_i}))/\sum_{p\in C_i}\exp(\cos(h_i,p)), where hih_i is the encoded example and pyip_{y_i} is its target prototype. ConPL minimizes Lfc=−∑ilog⁡qiL_{fc}=-\sum_i\log q_i over training examples. Restricting this loss to confusing classes emphasizes distinctions among similar relations; the paper uses α=0.1\alpha=0.1.

  6. Knowl 6 — Three-stage ConPL training procedure

    algorithm

    For each task, ConPL first builds temporary prototypes for the new relations, trains using new-task examples plus the previous sample memory, selects one key sample per new relation, and then trains with the enlarged memory. A final memory-only stage applies distribution consistency. The classification-stage objective combines cross-entropy, classification consistency, and the similar-class loss; the final-stage objective adds distribution consistency. The paper uses one epoch for each of the first two stages and three for the final stage.

    Input: New-task training set D_k and relation set R_k; previous sample memory S_hat^(k-1); previous prototype memory P_hat^(k-1).
    Output: Updated encoder and sample and prototype memories for task k.
    Initialize a temporary prototype for each relation in R_k by averaging its encoded training examples.
    Combine previous prototypes with the temporary prototypes, and combine D_k with S_hat^(k-1).
    For 1 epoch:
        Update the encoder and current-task prototypes using the classification-stage objective on the combined training data and previous memory.
    For each relation in R_k:
        Select the training example closest to that relation's mean representation and add it to the sample memory.
        Initialize that relation's stored prototype from the selected example's representation.
    Combine D_k with the enlarged sample memory, and combine previous prototypes with the new stored prototypes.
    For 1 epoch:
        Update the encoder and current-task prototypes using the classification-stage objective on the combined data and enlarged memory.
    For 3 epochs:
        Train using the enlarged sample memory alone, adding distribution consistency to the classification-stage objective.
        Update the encoder and current-task prototypes.
    Save the enlarged sample and prototype memories.
  7. Knowl 7 — Datasets, task sequences, and training configuration

    experimental setup

    Experiments use FewRel and TACRED in NK-CRE. FewRel’s 80 publicly accessible training and validation relations are divided into eight 10-relation tasks; the settings are 10-way-2-shot, 10-way-5-shot, and 10-way-10-shot. TACRED is filtered to remove the n/a relation, leaving 41 relations divided into eight tasks, one with six relations and seven with five; the settings are 5-way-5-shot and 5-way-10-shot. The encoder is BERTBASE, trained with Adam at learning rate 2×10−52\times10^{-5} and gradient clipping value 10. All four loss weights are set to 1.0. Accuracy results average six randomly generated task sequences, with the same sequences used for compared methods. The main table below reports reproduced prompt-based baselines and ConPL in the 10-way-5-shot FewRel and 5-way-5-shot TACRED settings.

  8. Knowl 8 — Whole accuracy across continual tasks

    data/table

    The table reports whole accuracy (%) after each task: evaluation is on the union of test sets for all tasks learned so far. It compares ConPL with the three reproduced prompt-based baselines under the same few-shot continual settings. ConPL is strongest from task 2 onward on FewRel and on every TACRED task; its final-task accuracies are 85.77% on FewRel and 76.38% on TACRED. On the first FewRel task, ERDA(PT) is higher than ConPL, 96.55% versus 95.72%; the paper notes that ERDA(PT) uses external Wikipedia sentences for augmentation.

    Benchmark and methodT1T2T3T4T5T6T7T8
    FewRel EMAR(PT)95.2892.7590.5788.1587.0484.9083.0781.34
    FewRel RP-CRE(PT)93.5289.4887.6785.2884.4883.0881.9280.87
    FewRel ERDA(PT)96.5592.5688.5684.4784.1479.9478.4577.02
    FewRel ConPL95.7293.5391.3189.9588.9388.3987.4385.77
    TACRED EMAR(PT)94.8886.3683.2580.6077.7675.4373.8568.67
    TACRED RP-CRE(PT)90.9884.7178.9675.6473.0771.1466.9965.31
    TACRED ERDA(PT)95.7786.5578.5974.5869.3166.5361.9255.97
    TACRED ConPL97.3489.8586.3382.5381.2179.5678.3876.38
  9. Knowl 9 — Ablation effects on FewRel

    data/table

    Ablations on FewRel 10-way-5-shot measure the effect of removing one ConPL component at a time. The values below are whole accuracy (%) after the first and eighth tasks. Removing prototype memory reduces task-8 accuracy by 3.56 percentage points, removing the consistent-learning module by 1.52 points, and removing the similar-class loss by 10.66 points. The classification and distribution consistency losses have smaller individual effects in this configuration. Without consistent learning, task-1 accuracy is 96.43%, compared with 95.72% for full ConPL; the authors attribute this to the absence of forgetting on the first task and potential overfitting to memory.

    ModelT1T8
    ConPL95.7285.77
    Without prototype memory95.7282.21
    Without consistent-learning module96.4384.25
    Without classification consistency loss95.7285.69
    Without distribution consistency loss95.7285.40
    Without similar-class loss94.8775.11
  10. Knowl 10 — Forgetting and prototype distortion analysis

    empirical result

    After all tasks are learned in FewRel 10-way-5-shot, ConPL’s mean forgetting is 3.31%, close to the JointTrain reference at 3.29% and far below the SeqRun reference at 82.08%. JointTrain retains all past training data, while SeqRun fine-tunes on new-task data without past data. The other reported means are EMAR 13.34%, RP-CRE 10.70%, ERDA 20.32%, EMAR(PT) 7.24%, RP-CRE(PT) 5.18%, and ERDA(PT) 12.78%. The paper also compares prototype embedding distortion with forgetting across 50 randomly generated task sequences: its analysis reports that greater prototype distortion is highly related to greater forgetting, and the scatter plot places ConPL prototypes at comparatively low distortion and forgetting. The paper does not provide a numerical correlation coefficient and acknowledges that it does not explain the exceptional prototypes that do not follow the general pattern.

Coverage note — No substantial contributed material was omitted; the paper’s storage-overhead caveat is included with the memory method, and its unanalysed prototype exceptions are noted with the distortion result.

References

  1. 1.Sagie Benaim and Lior Wolf. 2018. One-shot unsupervised cross domain translation. In Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901.
  3. 3.Hyuntak Cha, Jaeho Lee, and Jinwoo Shin. 2021. Co2l: Contrastive continual learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 9516–9525.
  4. 4.Arslan Chaudhry, Marc’Aurelio Ranzato, Marcus Rohrbach, and Mohamed Elhoseiny. 2018. Efficient lifelong learning with a-gem.
  5. 5.Kuilin Chen and Chi-Guhn Lee. 2021. Incremental few-shot learning via vector quantization in deep embedded space. In International Conference on Learning Representations.
  6. 6.Xiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng, Yunzhi Yao, Chuanqi Tan, Fei Huang, Luo Si, and Huajun Chen. 2022. Knowprompt: Knowledge-aware prompt-tuning with synergistic optimization for relation extraction. In Proceedings of the ACM Web Conference 2022, page 2778–2788.
  7. 7.Li Cui, Deqing Yang, Jiaxin Yu, Chengwei Hu, Jiayang Cheng, Jingjie Yi, and Yanghua Xiao. 2021. Refining sample embeddings with relation prototypes to enhance continual relation extraction. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 232–243.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding.
  9. 9.Chrisantha Fernando, Dylan Banarse, Charles Blundell, Yori Zwols, David Ha, Andrei A. Rusu, Alexander Pritzel, and Daan Wierstra. 2017. Pathnet: Evolution channels gradient descent in super neural networks.
  10. 10.Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017. Model-agnostic meta-learning for fast adaptation of deep networks. In Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 1126–1135. PMLR.
  11. 11.Jiale Han, Bo Cheng, and Wei Lu. 2021a. Exploring task difficulty for few-shot relation extraction. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2605–2616.
  12. 12.Xu Han, Yi Dai, Tianyu Gao, Yankai Lin, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. 2020. Continual relation learning via episodic memory activation and reconsolidation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6429–6440.
  13. 13.Xu Han, Weilin Zhao, Ning Ding, Zhiyuan Liu, and Maosong Sun. 2021b. Ptr: Prompt tuning with rules for text classification.
  14. 14.Xu Han, Hao Zhu, Pengfei Yu, Ziyun Wang, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2018. FewRel: A large-scale supervised few-shot relation classification dataset with state-of-the-art evaluation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4803–4809.
  15. 15.Chengwei Hu, Deqing Yang, Haoliang Jin, Zhen Chen, and Yanghua Xiao. 2022. Improving continual relation extraction through prototypical contrastive learning. In Proceedings of the 29th International Conference on Computational Linguistics, pages 1885–1895.
  16. 16.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. 2017. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526.
  17. 17.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3045–3059.
  18. 18.Abiola Obamuyide and Andreas Vlachos. 2019. Meta-learning improves lifelong relation extraction. In Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019), pages 224–229.
  19. 19.Chengwei Qin and Shafiq Joty. 2022. Continual few-shot relation learning via embedding space regularization and data augmentation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2776–2789.
  20. 20.Meng Qu, Tianyu Gao, Louis-Pascal Xhonneux, and Jian Tang. 2020. Few-shot relation extraction via Bayesian meta-learning on relation graphs. In Proceedings of the 37th International Conference on Machine Learning, volume 119, pages 7867–7876.
  21. 21.David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy Lillicrap, and Gregory Wayne. 2019. Experience replay for continual learning. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  22. 22.Timo Schick, Helmut Schmid, and Hinrich Schütze. 2020. Automatically identifying words that can serve as labels for few-shot text classification. In Proceedings of the 28th International Conference on Computational Linguistics, pages 5569–5578.
  23. 23.Jake Snell, Kevin Swersky, and Richard Zemel. 2017. Prototypical networks for few-shot learning. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  24. 24.Yao-Hung Hubert Tsai and Ruslan Salakhutdinov. 2017. Improving one-shot learning through fusing side information.
  25. 25.Tu Vu, Brian Lester, Noah Constant, Rami Al-Rfou’, and Daniel Cer. 2022. SPoT: Better frozen model adaptation through soft prompt transfer. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5039–5059.
  26. 26.Hong Wang, Wenhan Xiong, Mo Yu, Xiaoxiao Guo, Shiyu Chang, and William Yang Wang. 2019. Sentence embedding alignment for lifelong relation extraction. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 796–806.
  27. 27.Peiyi Wang, Yifan Song, Tianyu Liu, Binghuai Lin, Yunbo Cao, Sujian Li, and Zhifang Sui. 2022. Learning robust representations for continual relation extraction via adversarial class augmentation.
  28. 28.Mitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi, Mohammad Rastegari, Jason Yosinski, and Ali Farhadi. 2020. Supermasks in superposition. In Advances in Neural Information Processing Systems, volume 33, pages 15173–15184. Curran Associates, Inc.
  29. 29.Kaijia Yang, Nantao Zheng, Xinyu Dai, Liang He, Shujian Huang, and Jiajun Chen. 2020. Enhance prototypical network with text descriptions for few-shot relation classification. In Proceedings of the 29th ACM International Conference on Information Knowledge Management, page 2273–2276.
  30. 30.Friedemann Zenke, Ben Poole, and Surya Ganguli. 2017. Continual learning through synaptic intelligence. In Proceedings of the 34th International Conference on Machine Learning - Volume 70, ICML’17, page 3987–3995. JMLR.org.
  31. 31.Yuhao Zhang, Victor Zhong, Danqi Chen, Gabor Angeli, and Christopher D. Manning. 2017. Position-aware attention and supervised data improve slot filling. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 35–45.
  32. 32.Kang Zhao, Hua Xu, Jiangong Yang, and Kai Gao. 2022. Consistent representation learning for continual relation extraction. In Findings of the Association for Computational Linguistics: ACL 2022, pages 3402–3411.

Citation

MLA
Chen, X., et al. “Consistent Prototype Learning for Few-Shot Continual Relation Extraction”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 7409–22, https://doi.org/10.18653/V1/2023.ACL-LONG.409.
APA
Chen, X., Wu, H., & Shi, X. (2023). Consistent Prototype Learning for Few-Shot Continual Relation Extraction. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7409–7422. https://doi.org/10.18653/V1/2023.ACL-LONG.409
Chicago
Chen, X., H. Wu, and X. Shi. 2023. “Consistent Prototype Learning for Few-Shot Continual Relation Extraction”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7409–22. https://doi.org/10.18653/V1/2023.ACL-LONG.409.
Harvard
Chen, X., Wu, H. and Shi, X. (2023) “Consistent Prototype Learning for Few-Shot Continual Relation Extraction”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7409–7422. Available at: https://doi.org/10.18653/V1/2023.ACL-LONG.409.
Vancouver
1. Chen X, Wu H, Shi X (2023) Consistent Prototype Learning for Few-Shot Continual Relation Extraction. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7409–7422

BibTeX

@inproceedings{Chen_2023, title={Consistent Prototype Learning for Few-Shot Continual Relation Extraction}, url={http://dx.doi.org/10.18653/V1/2023.ACL-LONG.409}, DOI={10.18653/v1/2023.acl-long.409}, booktitle={Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)}, publisher={Association for Computational Linguistics}, author={Chen, Xiudi and Wu, Hui and Shi, Xiaodong}, year={2023}, pages={7409–7422} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/