A Survey of Active Learning for Natural Language Processing

Zhisong ZhangEmma StrubellEduard H. Hovy

article2022EMNLP90 citations

Presents a structured overview of active learning query strategies and practical challenges specific to natural language processing, offering essential guidance on handling deep neural architectures, structured prediction, and annotation costs.

Listen

Modern natural language processing systems rely heavily on data-driven machine learning, but gathering sufficient high-quality training labels remains expensive and labor-intensive. Active learning, an approach where the model strategically selects which data points to annotate to reach target accuracy with fewer labeled examples, addresses this bottleneck. As deep learning has transformed the field over the last decade, active learning methods have evolved significantly. The article provides a comprehensive literature survey categorizing active learning techniques specifically tailored to language technologies, spanning query design, complex data structures, real-world costs, advanced model training, and lifecycle management.

The article establishes that effective sample selection hinges on balancing two core criteria: informativeness, which measures the utility or uncertainty of an individual instance, and representativeness, which ensures selected instances reflect broader data distributions and avoid repetitive outliers. Selecting instances solely by individual difficulty often introduces sampling bias or captures unusable noise, making hybrid and dynamic multi-step strategies essential. When dealing with structured language tasks like parsing or sequence tagging, querying partial sub-structures instead of full instances significantly enhances labeling efficiency, provided models can process incomplete annotations. Furthermore, the analysis emphasizes that assuming equal labeling effort per text instance distorts real-world utility; true economic gains require cost-sensitive selection policies that account for actual annotator time.

These findings have direct operational implications for language technology development budgets, deployment timelines, and data management. Deploying active learning requires accounting for real human workflow constraints, such as annotator wait times during model retraining and the risks of model mismatch. Specifically, querying data with one model architecture and subsequently training a different, downstream architecture can entirely eliminate active learning efficiency gains. Consequently, organizations cannot treat active learning merely as an isolated algorithmic selection step, but must integrate it with semi-supervised learning, pre-trained language representations, and interactive computer-assisted labeling interfaces to maximize return on investment.

To implement active learning effectively, practitioners should adopt cost-sensitive return-on-investment selection strategies rather than relying solely on single-instance uncertainty metrics. Teams should incorporate warm-start methods, such as clustering centroids or language model representations, to address initial data cold-starts, and implement principled stopping criteria based on stabilization metrics across separate evaluation sets to avoid wasted labeling spend. Where possible, workflows should combine active sampling with pre-annotation interfaces to directly lower manual labor times. Because the article is a qualitative literature synthesis rather than a standardized benchmark study, decisions should be approached with caution regarding domain-specific performance, and pilot testing should precede wide-scale production deployment.

arXiv: 2210.10109
Cover for A Survey of Active Learning for Natural Language Processing

Abstract

In this work, we provide a literature review of active learning (AL) for its applications in natural language processing (NLP). In addition to a fine-grained categorization of query strategies, we also investigate several other important aspects of applying AL to NLP problems. These include AL for structured prediction tasks, annotation cost, model learning (especially with deep neural models), and starting and stopping AL. Finally, we conclude with a discussion of related topics and future directions.

Table of Contents

  • 1 Introduction
  • 1.1 Overview
  • 2 Query Strategies
  • 2.1 Informativeness
  • 2.1.1 Output Uncertainty
  • 2.1.2 Disagreement
  • 2.1.3 Gradient
  • 2.1.4 Performance Prediction
  • 2.2 Representativeness
  • 2.2.1 Density
  • 2.2.2 Discriminative
  • 2.2.3 Batch Diversity
  • 2.3 Hybrid
  • 3 Query and Annotation
  • 3.1 AL for Structured Prediction
  • 3.1.1 Full-structure AL
  • 3.1.2 Partial-structure AL
  • 3.2 Annotation Cost
  • 3.2.1 Cost Measurement
  • 3.2.2 Cost-sensitive Querying
  • 3.2.3 Directly Reducing Cost
  • 3.2.4 Wait Time
  • 4 Model and Learning
  • 4.1 Model Mismatch
  • 4.2 Learning
  • 5 Starting and Stopping AL
  • 5.1 Starting AL
  • 5.2 Stopping AL
  • 6 Related Topics and Future Directions
  • 6.1 Related Topics
  • 6.2 Future Directions
  • Limitations
  • References
  • A Tasks
  • B Other Aspects
  • C Surveying Process

Knowls

  1. Knowl 1 — Pool-based active-learning workflow

    algorithm

    The survey adopts pool-based active learning, in which an unlabeled pool is available and a model repeatedly chooses instances for annotation. Let UU be the current unlabeled pool, LL the labeled dataset, and MM the current task model. A typical procedure first creates a seed labeled set, trains a model, alternates querying and annotation, and retrains after each update.

    Input: An unlabeled data pool U
    Output: The final labeled dataset L and trained model M
    L, U ← seed(U)
    M ← train(L, U)
    while the stopping criterion is not satisfied do
        I ← query(M, U)
        I′ ← annotate(I)
        U ← U − I
        L ← L ∪ I′
        M ← train(L, U)
    return L, M

    The query, annotation, model-learning, initialization, and stopping choices are coupled: each query is made using the current model, each annotation changes the labeled and unlabeled pools, and the updated model determines subsequent queries.

  2. Knowl 2 — Informativeness-based query strategies

    model/method

    In active learning, informativeness-based strategies assign a utility score to each unlabeled instance independently and select instances with the highest scores. The survey groups the main signals into four families.

    • Output uncertainty: For probabilistic models, entropy, least confidence, and prediction margin favor instances whose predicted label distribution is uncertain. For non-probabilistic models, instances close to a decision boundary can serve the same purpose. Uncertainty can also be estimated from prediction changes in a local neighborhood, using nearest neighbors, adversarial perturbations, or data augmentation.
    • Disagreement: Query-by-committee methods use several models, or several approximate samples from a Bayesian parameter distribution, and select instances on which the models disagree. Vote entropy, KL divergence, and variation ratio are representative disagreement measures. Dropout-based Bayesian approximations make this approach applicable to neural NLP models.
    • Gradient information: Expected gradient length selects examples expected to produce large parameter updates. Because the gold label of an unlabeled instance is unknown, the gradient norm is evaluated in expectation over possible labels. Variants may restrict the gradient to selected parameters, such as word embeddings.
    • Performance prediction: Expected error reduction selects an instance according to the expected future error after labeling and retraining on it, but evaluating every candidate may require expensive retraining. Learned selection policies, loss-prediction models, and task-specific quality estimators provide cheaper approximations; these approaches may require labeled data for the policy and can have unstable learning signals on complex tasks.
  3. Knowl 3 — Representativeness and diversity in query selection

    model/method

    Representativeness-based active learning complements instance-level informativeness by modeling relationships among examples. It is intended to reduce sampling bias, avoid selecting outliers, and cover the unlabeled distribution.

    • Density: A candidate is favored when it is similar to many other unlabeled instances. Simple proxies include n-gram or word frequency, while a general density score is the candidate’s average similarity to the unlabeled pool. Computing all pairwise similarities can be replaced by considering only kk nearest neighbors.
    • Discriminativeness relative to labeled data: A candidate is favored when it differs from the current labeled set, for example by containing unseen n-grams or out-of-vocabulary words, having low similarity to labeled examples, or being judged by a classifier that separates labeled from unlabeled data. Domain separators provide a related strategy for domain adaptation.
    • Batch diversity: When several instances are selected together, the selected examples should also be dissimilar to one another. Greedy iterative selection adds examples while penalizing redundancy; clustering selects representatives from different clusters. Coreset methods and determinantal point processes provide more advanced diversity criteria. Similarity can be computed from inputs, neural representations, model-based features, gradients, or masked-language-model surprisal embeddings.
  4. Knowl 4 — Hybrid and dynamically combined query strategies

    model/method

    Active-learning query strategies can combine informativeness and representativeness because uncertainty alone may repeatedly select similar examples or outliers, whereas representativeness alone may miss decision-boundary cases. A simple hybrid combines utility scores using a weighted sum or multiplication. Other methods integrate the criteria structurally, such as uncertainty-weighted clustering, diversity selection in gradient space, or determinantal point processes that separate quality from diversity.

    Hybrid querying can also be multi-step: one stage may filter highly uncertain instances and a later stage may select a diverse subset, or the most uncertain example may be selected separately within each cluster. The survey further identifies dynamic combinations whose weights change during active learning. Representativeness is often more reliable at the beginning, when few labeled examples make uncertainty estimates unstable; uncertainty can become more useful later, when the model has enough data to refine decision boundaries. Dynamic strategies therefore switch or gradually transition from density-based selection to uncertainty-based selection.

  5. Knowl 5 — Active learning for structured prediction

    model/method

    NLP structured-prediction tasks produce interdependent outputs such as label sequences, parse trees, dependency edges, or coreference links, so active learning must choose both an annotation unit and a way to measure uncertainty over structured outputs.

    For full-structure annotation, the entire output for an input is queried and labeled. Because the output space may be exponentially large, full-structure entropy is computed with dynamic-programming procedures analogous to decoding or inference, or approximated using the top kk predicted structures. Disagreement can use partial-match metrics such as sequence-labeling F1 rather than requiring exact agreement. Length normalization is often used to prevent longer inputs from receiving larger uncertainty scores merely because they contain more components, although longer sequences may also contain more useful information. Structured utility can alternatively be aggregated from token- or factor-level utilities by summing, averaging, frequency weighting, or selecting the single least-probable component.

    For partial-structure annotation, the query unit is a substructure rather than a complete output. Examples include tokens, word boundaries, dependency edges, coreference links, subsequences, phrases, or groups of nearby instances. Larger substructures can preserve context, reduce annotation overhead from switching between examples, and lessen the risk that rare classes are missed. Local models can learn directly from independently labeled substructures, whereas global models require methods for learning from incomplete structures. Global constraints and partial labels can reduce the feasible output space for unannotated components, thereby improving later queries; partial feedback such as yes/no questions or ruling out labels can similarly simplify classification annotation.

  6. Knowl 6 — Annotation cost, cost-sensitive querying, and interaction latency

    model/method

    The survey emphasizes that active learning should be evaluated by real annotation effort rather than only by the number of labeled instances. Unit-cost assumptions can be misleading because longer, ambiguous, or difficult examples may take more time. Token-based costs are sometimes used as a refinement, but actual annotation time is the most direct measure, especially when comparing full and partial annotation, annotation with rule writing, or different annotation interfaces. Cost can also be predicted before annotation using regression models based on input features.

    Cost-sensitive querying favors examples with high expected utility and low annotation cost. If q(x)q(x) is the utility of querying instance xx and c(x)>0c(x)>0 is its annotation cost in time or another cost unit, a return-on-investment strategy ranks examples using q(x)/c(x)q(x)/c(x). Real settings may also involve multiple annotators with different expertise, refusals, or errors; proactive learning jointly selects an annotator and an instance.

    Annotation effort can be reduced through pre-annotation, in which model predictions or top-kk predictions are presented for correction, post-editing in machine translation, and interactive systems in which corrections immediately update subsequent predictions. Waiting for training and scoring between iterations is another cost. Subsampling, cached features, nearest-neighbor approximations, efficient query models, and parallel or stale-information workflows reduce latency. Continuing neural training from the previous iteration can be computationally attractive, but the survey notes evidence that warm starts may produce suboptimal active-learning performance, motivating retraining from scratch in many recent systems.

  7. Knowl 7 — Model mismatch and complementary learning methods

    model/method

    The model used to select queries need not be the model ultimately used for deployment. A smaller or faster query model may be preferred to reduce active-learning latency, and labeled data collected by active learning may later be reused to train a different successor model. The survey reports that this query–successor model mismatch can substantially reduce active-learning gains and can sometimes make them negligible or negative.

    Distillation offers one compromise: a smaller model can be distilled from a larger pretrained model for querying while retaining much of the query model’s useful behavior. Proxy models can also be synchronized with the main model through distillation, and pseudo-labeling or pool subsampling can further reduce computation.

    Active learning is compatible with several forms of data-efficient learning. Semi-supervised learning uses predictions on the unlabeled pool through self-training or pseudo-labeling; transfer learning supplies pretrained, cross-domain, or cross-language representations; weak supervision contributes rules, dictionaries, execution results, or other noisy signals; and data augmentation expands the queried training data or generates alternative views for uncertainty estimation. Query-synthesis methods go further by generating new instances for annotation instead of selecting only from the existing pool, although text synthesis remains difficult beyond relatively simple classification settings.

  8. Knowl 8 — Initialization and cold-start strategies

    model/method

    Active learning faces a cold-start problem when no reliable model exists to score the initial unlabeled pool. The initial seed set can strongly affect performance during the early iterations. Random sampling is the most common initialization because it approximately preserves the original data distribution. Representativeness-based initialization instead clusters the unlabeled data and selects points near cluster centroids to obtain coverage and diversity.

    Other initialization methods use information available without task labels. Transfer learning and unsupervised methods can provide useful representations or selection signals. A language model can select low-probability words in context, or use surprisal-based contextual embeddings and cluster centers to construct a seed set. These strategies are alternatives for creating an informative initial labeled dataset before ordinary model-based querying becomes reliable.

  9. Knowl 9 — Stopping criteria for active learning

    model/method

    A practical active-learning system needs a stopping rule indicating when additional annotation is unlikely to justify its cost. The survey organizes stopping design around three choices: the metric, the dataset on which it is measured, and the stopping condition.

    Candidate metrics include development-set performance, uncertainty or confidence, committee disagreement, estimated performance or expected error, confidence variation, performance on selected instances, and prediction changes between consecutive active-learning iterations. Development sets can be unstable when small, while cross-validation on actively labeled data can be biased because the queried sample is not representative. A separate unlabeled stop set is often preferable because it avoids requiring gold labels and can be made large enough for stable estimates, although processing it adds latency.

    Stopping based on a fixed threshold at one iteration is simple but task- and model-dependent. More stable alternatives monitor several iterations and stop when confidence consistently decreases, metric changes flatten, or predictions stabilize. Some task-specific rules stop when no unlabeled examples are closer to a support-vector boundary than existing support vectors, or when the unlabeled pool contains no new n-grams.

  10. Knowl 10 — Scope and practical limitations identified by the survey

    limitation

    The survey’s evidence base is primarily NLP research, and its descriptions are intentionally brief so that many query strategies, structured-prediction settings, annotation-cost issues, learning methods, and lifecycle decisions can be covered. It does not experimentally compare active-learning strategies: the work is a literature survey rather than a new empirical study.

    The survey identifies a practical gap between simulated and deployed active learning. Most studies simulate annotation by revealing labels from an existing labeled corpus, which can hide real annotation time, annotator noise, model-training wait time, data reuse, cold-start choices, stopping decisions, and the difficulty of selecting query hyperparameters when the process cannot be repeated. It also highlights underexplored settings such as complex generation and question-answering tasks, annotation of rationales or explanations rather than target labels, and other human-in-the-loop objectives. These limitations motivate empirical comparisons and more realistic deployment studies.

Coverage note — Deliberately omitted the appendix-style task bibliography and shorter discussions of crowdsourcing, multiple targets, data imbalance, curriculum learning, and related human-in-the-loop extensions because they are secondary to the survey’s core query, annotation, learning, and lifecycle taxonomy.

References

  1. 1.Charu C Aggarwal, Xiangnan Kong, Quanquan Gu, Jiawei Han, and S Yu Philip. 2014. Active learning: A survey. In Data Classification, pages 599–634. Chapman and Hall/CRC.
  2. 2.Vamshi Ambati, Sanjika Hewavitharana, Stephan Vogel, and Jaime Carbonell. 2011a. Active learning with multiple annotations for comparable data classification task. In Proceedings of the 4th Workshop on Building and Using Comparable Corpora: Comparable Corpora and the Web, pages 69–77, Portland, Oregon. Association for Computational Linguistics.
  3. 3.Vamshi Ambati, Stephan Vogel, and Jaime Carbonell. 2010a. Active learning and crowd-sourcing for machine translation. In Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC’10), Valletta, Malta. European Language Resources Association (ELRA).
  4. 4.Vamshi Ambati, Stephan Vogel, and Jaime Carbonell. 2010b. Active learning-based elicitation for semi-supervised word alignment. In Proceedings of the ACL 2010 Conference Short Papers, pages 365–370, Uppsala, Sweden. Association for Computational Linguistics.
  5. 5.Vamshi Ambati, Stephan Vogel, and Jaime Carbonell. 2010c. Active semi-supervised learning for improving word alignment. In Proceedings of the NAACL HLT 2010 Workshop on Active Learning for Natural Language Processing, pages 10–17, Los Angeles, California. Association for Computational Linguistics.
  6. 6.Vamshi Ambati, Stephan Vogel, and Jaime Carbonell. 2011b. Multi-strategy approaches to active learning for statistical machine translation. In Proceedings of Machine Translation Summit XIII: Papers, Xiamen, China.
  7. 7.Sankaranarayanan Ananthakrishnan, Rohit Prasad, David Stallard, and Prem Natarajan. 2010a. Discriminative sample selection for statistical machine translation. In Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing, pages 626–635, Cambridge, MA. Association for Computational Linguistics.
  8. 8.Sankaranarayanan Ananthakrishnan, Rohit Prasad, David Stallard, and Prem Natarajan. 2010b. A semi-supervised batch-mode active learning strategy for improved statistical machine translation. In Proceedings of the Fourteenth Conference on Computational Natural Language Learning, pages 126–134, Uppsala, Sweden. Association for Computational Linguistics.
  9. 9.Shilpa Arora, Eric Nyberg, and Carolyn P. Rosé. 2009. Estimating annotation cost for active learning in a multi-annotator environment. In Proceedings of the NAACL HLT 2009 Workshop on Active Learning for Natural Language Processing, pages 18–26, Boulder, Colorado. Association for Computational Linguistics.
  10. 10.Jordan Ash and Ryan P Adams. 2020. On warm-starting neural network training. Advances in Neural Information Processing Systems, 33:3884–3894.
  11. 11.Jordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford, and Alekh Agarwal. 2020. Deep batch active learning by diverse, uncertain gradient lower bounds. In International Conference on Learning Representations.
  12. 12.Seyed Arad Ashrafi Asli, Behnam Sabeti, Zahra Majdabadi, Preni Golazizian, Reza Fahmi, and Omid Momenzadeh. 2020. Optimizing annotation effort using active learning strategies: A sentiment analysis case study in Persian. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 2855–2861, Marseille, France. European Language Resources Association.
  13. 13.Jordi Atserias, Giuseppe Attardi, Maria Simi, and Hugo Zaragoza. 2010. Active learning for building a corpus of questions for parsing. In Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC’10), Valletta, Malta. European Language Resources Association (ELRA).
  14. 14.Josh Attenberg and ¸Seyda Ertekin. 2013. Class imbalance and active learning. Imbalanced Learning: Foundations, Algorithms, and Applications, pages 101–149.
  15. 15.Josh Attenberg and Foster Provost. 2011. Inactive learning? difficulties employing active learning in practice. ACM SIGKDD Explorations Newsletter, 12(2):36–41.
  16. 16.Philip Bachman, Alessandro Sordoni, and Adam Trischler. 2017. Learning algorithms for active learning. In International Conference on Machine Learning, pages 301–310. PMLR.
  17. 17.Guirong Bai, Shizhu He, Kang Liu, Jun Zhao, and Zaiqing Nie. 2020. Pre-trained language model based active learning for sentence matching. In Proceedings of the 28th International Conference on Computational Linguistics, pages 1495–1504, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  18. 18.Jason Baldridge and Miles Osborne. 2003. Active learning for HPSG parse selection. In Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003, pages 17–24.
  19. 19.Jason Baldridge and Miles Osborne. 2004. Active learning and the total cost of annotation. In Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing, pages 9–16, Barcelona, Spain. Association for Computational Linguistics.
  20. 20.Jason Baldridge and Alexis Palmer. 2009. How well does active learning actually work? Time-based evaluation of cost-reduction strategies for language documentation. In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing, pages 296–305, Singapore. Association for Computational Linguistics.
  21. 21.Garrett Beatty, Ethan Kochis, and Michael Bloodgood. 2019. The use of unlabeled data versus labeled data for stopping active learning for text classification. In 2019 IEEE 13th International Conference on Semantic Computing (ICSC), pages 287–294. IEEE.
  22. 22.Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009. Curriculum learning. In Proceedings of the 26th annual international conference on machine learning, pages 41–48.
  23. 23.Michael Bloodgood and Chris Callison-Burch. 2010. Bucking the trend: Large-scale cost-focused active learning for statistical machine translation. In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, pages 854–864, Uppsala, Sweden. Association for Computational Linguistics.
  24. 24.Michael Bloodgood and John Grothendieck. 2013. Analysis of stopping active learning based on stabilizing predictions. In Proceedings of the Seventeenth Conference on Computational Natural Language Learning, pages 10–19, Sofia, Bulgaria. Association for Computational Linguistics.
  25. 25.Michael Bloodgood and K. Vijay-Shanker. 2009a. A method for stopping active learning based on stabilizing predictions and the need for user-adjustable stopping. In Proceedings of the Thirteenth Conference on Computational Natural Language Learning (CoNLL-2009), pages 39–47, Boulder, Colorado. Association for Computational Linguistics.
  26. 26.Michael Bloodgood and K. Vijay-Shanker. 2009b. Taking into account the differences between actively and passively acquired data: The case of active learning with support vector machines for imbalanced datasets. In Proceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics, Companion Volume: Short Papers, pages 137–140, Boulder, Colorado. Association for Computational Linguistics.
  27. 27.Kianté Brantley, Amr Sharaf, and Hal Daumé III. 2020. Active imitation learning with noisy guidance. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2093–2105, Online. Association for Computational Linguistics.
  28. 28.Klaus Brinker. 2003. Incorporating diversity in active learning with support vector machines. In Proceedings of the 20th International Conference on Machine Learning (ICML-03), pages 59–66.
  29. 29.Tingting Cai, Zhiyuan Ma, Hong Zheng, and Yangming Zhou. 2021. Ne–lp: normalized entropy-and loss prediction-based sampling for active learning in chinese word segmentation on ehrs. Neural Computing and Applications, 33(19):12535–12549.
  30. 30.Tingting Cai, Yangming Zhou, and Hong Zheng. 2020. Cost-quality adaptive active learning for chinese clinical named entity recognition. In 2020 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pages 528–533. IEEE.
  31. 31.Hian Cañizares-Díaz, Alejandro Piad-Morffis, Suilan Estevez-Velarde, Yoan Gutiérrez, Yudivián Almeida Cruz, Andres Montoyo, and Rafael Muñoz-Guillena. 2021. Active learning for assisted corpus construction: A case study in knowledge discovery from biomedical text. In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021), pages 216–225, Held Online. INCOMA Ltd.
  32. 32.Kai Cao, Xiang Li, Miao Fan, and Ralph Grishman. 2015. Improving event detection with active learning. In Proceedings of the International Conference Recent Advances in Natural Language Processing, pages 72–77, Hissar, Bulgaria. INCOMA Ltd. Shoumen, BULGARIA.
  33. 33.Yee Seng Chan and Hwee Tou Ng. 2007. Domain adaptation with active learning for word sense disambiguation. In Proceedings of the 45th Annual Meeting of the Association of Computational Linguistics, pages 49–56, Prague, Czech Republic. Association for Computational Linguistics.
  34. 34.Aditi Chaudhary, Antonios Anastasopoulos, Zaid Sheikh, and Graham Neubig. 2021. Reducing confusion in active learning for part-of-speech tagging. Transactions of the Association for Computational Linguistics, 9:1–16.
  35. 35.Aditi Chaudhary, Jiateng Xie, Zaid Sheikh, Graham Neubig, and Jaime Carbonell. 2019. A little annotation does a lot of good: A study in bootstrapping low-resource named entity recognizers. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5164–5174, Hong Kong, China. Association for Computational Linguistics.
  36. 36.Chenhua Chen, Alexis Palmer, and Caroline Sporleder. 2011. Enhancing active learning for semantic role labeling via compressed dependency trees. In Proceedings of 5th International Joint Conference on Natural Language Processing, pages 183–191, Chiang Mai, Thailand. Asian Federation of Natural Language Processing.
  37. 37.Jinying Chen, Andrew Schein, Lyle Ungar, and Martha Palmer. 2006. An empirical study of the behavior of active learning for word sense disambiguation. In Proceedings of the Human Language Technology Conference of the NAACL, Main Conference, pages 120–127, New York City, USA. Association for Computational Linguistics.
  38. 38.Yukun Chen, Thomas A Lasko, Qiaozhu Mei, Joshua C Denny, and Hua Xu. 2015. A study of active learning methods for named entity recognition in clinical text. Journal of biomedical informatics, 58:11–18.
  39. 39.Gui Citovsky, Giulia DeSalvo, Claudio Gentile, Lazaros Karydas, Anand Rajagopalan, Afshin Rostamizadeh, and Sanjiv Kumar. 2021. Batch active learning at scale. Advances in Neural Information Processing Systems, 34.
  40. 40.David Cohn, Les Atlas, and Richard Ladner. 1994. Improving generalization with active learning. Machine learning, 15(2):201–221.
  41. 41.David A Cohn, Zoubin Ghahramani, and Michael I Jordan. 1996. Active learning with statistical models. Journal of artificial intelligence research, 4:129–145.
  42. 42.Aron Culotta and Andrew McCallum. 2005. Reducing labeling effort for structured prediction tasks. In AAAI, volume 5, pages 746–751.
  43. 43.Aswarth Abhilash Dara, Josef van Genabith, Qun Liu, John Judge, and Antonio Toral. 2014. Active learning for post-editing based incrementally retrained MT. In Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, volume 2: Short Papers, pages 185–189, Gothenburg, Sweden. Association for Computational Linguistics.
  44. 44.Sajib Dasgupta and Vincent Ng. 2009. Mine the easy, classify the hard: A semi-supervised approach to automatic sentiment classification. In Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP, pages 701–709, Suntec, Singapore. Association for Computational Linguistics.
  45. 45.Sanjoy Dasgupta. 2011. Two faces of active learning. Theoretical computer science, 412(19):1767–1781.
  46. 46.Yue Deng, KaWai Chen, Yilin Shen, and Hongxia Jin. 2018. Adversarial active learning for sequences labeling and generation. In IJCAI, pages 4012–4018.
  47. 47.Dmitriy Dligach and Martha Palmer. 2011. Good seed makes a good crop: Accelerating active learning using language modeling. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 6–10, Portland, Oregon, USA. Association for Computational Linguistics.
  48. 48.Pinar Donmez and Jaime G Carbonell. 2008. Proactive learning: cost-sensitive active learning with multiple imperfect oracles. In Proceedings of the 17th ACM conference on Information and knowledge management, pages 619–628.
  49. 49.Pinar Donmez, Jaime G Carbonell, and Paul N Bennett. 2007. Dual strategy active learning. In European Conference on Machine Learning, pages 116–127. Springer.
  50. 50.Gregory Druck, Burr Settles, and Andrew McCallum. 2009. Active learning by labeling features. In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing, pages 81–90, Singapore. Association for Computational Linguistics.
  51. 51.Long Duong, Hadi Afshar, Dominique Estival, Glen Pink, Philip Cohen, and Mark Johnson. 2018. Active learning for deep semantic parsing. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 43–48, Melbourne, Australia. Association for Computational Linguistics.
  52. 52.Matthias Eck, Stephan Vogel, and Alex Waibel. 2005. Low cost portability for statistical machine translation based on n-gram frequency and TF-IDF. In Proceedings of the Second International Workshop on Spoken Language Translation, Pittsburgh, Pennsylvania, USA.
  53. 53.Liat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch, Lena Dankin, Leshem Choshen, Marina Danilevsky, Ranit Aharonov, Yoav Katz, and Noam Slonim. 2020. Active Learning for BERT: An Empirical Study. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7949–7962, Online. Association for Computational Linguistics.
  54. 54.Sean P. Engelson and Ido Dagan. 1996. Minimizing manual annotation cost in supervised training from corpora. In 34th Annual Meeting of the Association for Computational Linguistics, pages 319–326, Santa Cruz, California, USA. Association for Computational Linguistics.
  55. 55.Alexander Erdmann, David Joseph Wrisley, Benjamin Allen, Christopher Brown, Sophie Cohen-Bodénès, Micha Elsner, Yukun Feng, Brian Joseph, Béatrice Joyeux-Prunel, and Marie-Catherine de Marneffe. 2019. Practical, efficient, and customizable active learning for named entity recognition in the digital humanities. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2223–2234, Minneapolis, Minnesota. Association for Computational Linguistics.
  56. 56.Seyda Ertekin, Jian Huang, Leon Bottou, and Lee Giles. 2007. Learning on the border: active learning in imbalanced data classification. In Proceedings of the sixteenth ACM conference on Conference on information and knowledge management, pages 127–136.
  57. 57.Nuno Escudeiro and Alípio Jorge. 2010. D-confidence: An active learning strategy which efficiently identifies small classes. In Proceedings of the NAACL HLT 2010 Workshop on Active Learning for Natural Language Processing, pages 18–26, Los Angeles, California. Association for Computational Linguistics.
  58. 58.Vebjørn Espeland, Beatrice Alex, and Benjamin Bach. 2020. Enhanced labelling in active learning for coreference resolution. In Proceedings of the Third Workshop on Computational Models of Reference, Anaphora and Coreference, pages 111–121, Barcelona, Spain (online). Association for Computational Linguistics.
  59. 59.Meng Fang and Trevor Cohn. 2017. Model transfer for tagging low-resource languages using a bilingual dictionary. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 587–593, Vancouver, Canada. Association for Computational Linguistics.
  60. 60.Meng Fang, Yuan Li, and Trevor Cohn. 2017. Learning how to active learn: A deep reinforcement learning approach. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 595–605, Copenhagen, Denmark. Association for Computational Linguistics.
  61. 61.Meng Fang, Jie Yin, and Dacheng Tao. 2014. Active learning for crowdsourcing using knowledge transfer. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 28.
  62. 62.Daniel Flannery and Shinsuke Mori. 2015. Combining active learning and partial annotation for domain adaptation of a Japanese dependency parser. In Proceedings of the 14th International Conference on Parsing Technologies, pages 11–19, Bilbao, Spain. Association for Computational Linguistics.
  63. 63.Linton C Freeman. 1965. Elementary applied statistics: for students in behavioral science. New York: Wiley.
  64. 64.Lisheng Fu and Ralph Grishman. 2013. An efficient active learning framework for new relation types. In Proceedings of the Sixth International Joint Conference on Natural Language Processing, pages 692–698, Nagoya, Japan. Asian Federation of Natural Language Processing.
  65. 65.Yifan Fu, Xingquan Zhu, and Bin Li. 2013. A survey on instance selection for active learning. Knowledge and information systems, 35(2):249–283.
  66. 66.Atsushi Fujii, Kentaro Inui, Takenobu Tokunaga, and Hozumi Tanaka. 1998. Selective sampling for example-based word sense disambiguation. Computational Linguistics, 24(4):573–597.
  67. 67.Yarin Gal and Zoubin Ghahramani. 2016. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In International Conference on Machine Learning, pages 1050–1059. PMLR.
  68. 68.Yarin Gal, Riashat Islam, and Zoubin Ghahramani. 2017. Deep bayesian active learning with image data. In International Conference on Machine Learning, pages 1183–1192. PMLR.
  69. 69.Caroline Gasperin. 2009. Active learning for anaphora resolution. In Proceedings of the NAACL HLT 2009 Workshop on Active Learning for Natural Language Processing, pages 1–8, Boulder, Colorado. Association for Computational Linguistics.
  70. 70.Jakob Gawlikowski, Cedrique Rovile Njieutcheu Tassi, Mohsin Ali, Jongseok Lee, Matthias Humt, Jianxiang Feng, Anna Kruspe, Rudolph Triebel, Peter Jung, Ribana Roscher, et al. 2021. A survey of uncertainty in deep neural networks. arXiv preprint arXiv:2107.03342.
  71. 71.Yonatan Geifman and Ran El-Yaniv. 2017. Deep active learning over the long tail. arXiv preprint arXiv:1711.00941.
  72. 72.Masood Ghayoomi. 2010. Using variance as a stopping criterion for active learning of frame assignment. In Proceedings of the NAACL HLT 2010 Workshop on Active Learning for Natural Language Processing, pages 1–9, Los Angeles, California. Association for Computational Linguistics.
  73. 73.Daniel Gissin and Shai Shalev-Shwartz. 2019. Discriminative active learning. arXiv preprint arXiv:1907.06347.
  74. 74.Jesús González-Rubio, Daniel Ortiz-Martínez, and Francisco Casacuberta. 2012. Active learning for interactive machine translation. In Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics, pages 245–254, Avignon, France. Association for Computational Linguistics.
  75. 75.Daniel Grießhaber, Johannes Maucher, and Ngoc Thang Vu. 2020. Fine-tuning BERT for low-resource natural language understanding via active learning. In Proceedings of the 28th International Conference on Computational Linguistics, pages 1158–1171, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  76. 76.Kamal Gupta, Dhanvanth Boppana, Rejwanul Haque, Asif Ekbal, and Pushpak Bhattacharyya. 2021. Investigating active learning in interactive neural machine translation. In Proceedings of Machine Translation Summit XVIII: Research Track, pages 10–22, Virtual. Association for Machine Translation in the Americas.
  77. 77.Suchin Gururangan, Ana Marasovic, Swabha ´ Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360, Online. Association for Computational Linguistics.
  78. 78.Ben Hachey, Beatrice Alex, and Markus Becker. 2005. Investigating the effects of selective sampling on the annotation task. In Proceedings of the Ninth Conference on Computational Natural Language Learning (CoNLL-2005), pages 144–151, Ann Arbor, Michigan. Association for Computational Linguistics.
  79. 79.Hossein Hadian and Hossein Sameti. 2014. Active learning in noisy conditions for spoken language understanding. In Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers, pages 1081–1090, Dublin, Ireland. Dublin City University and Association for Computational Linguistics.
  80. 80.Robbie Haertel, Paul Felt, Eric K. Ringger, and Kevin Seppi. 2010. Parallel active learning: Eliminating wait time with minimal staleness. In Proceedings of the NAACL HLT 2010 Workshop on Active Learning for Natural Language Processing, pages 33–41, Los Angeles, California. Association for Computational Linguistics.
  81. 81.Robbie Haertel, Eric Ringger, Kevin Seppi, James Carroll, and Peter McClanahan. 2008a. Assessing the costs of sampling methods in active learning for annotation. In Proceedings of ACL-08: HLT, Short Papers, pages 65–68, Columbus, Ohio. Association for Computational Linguistics.
  82. 82.Robbie Haertel, Eric Ringger, Kevin Seppi, and Paul Felt. 2015. An analytic and empirical evaluation of return-on-investment-based active learning. In Proceedings of The 9th Linguistic Annotation Workshop, pages 11–20, Denver, Colorado, USA. Association for Computational Linguistics.
  83. 83.Robbie A Haertel, Kevin D Seppi, Eric K Ringger, and James L Carroll. 2008b. Return on investment for active learning. In Proceedings of the NIPS workshop on cost-sensitive learning, volume 72.
  84. 84.Gholamreza Haffari, Maxim Roy, and Anoop Sarkar. 2009. Active learning for statistical phrase-based machine translation. In Proceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics, pages 415–423, Boulder, Colorado. Association for Computational Linguistics.
  85. 85.Gholamreza Haffari and Anoop Sarkar. 2009. Active learning for multilingual statistical machine translation. In Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP, pages 181–189, Suntec, Singapore. Association for Computational Linguistics.
  86. 86.Rishi Hazra, Parag Dutta, Shubham Gupta, Mohammed Abdul Qaathir, and Ambedkar Dukkipati. 2021. Active2 learning: Actively reducing redundancies in active learning methods for sequence tagging and machine translation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1982–1995, Online. Association for Computational Linguistics.
  87. 87.Rui He, Shan He, and Ke Tang. 2021. Multi-domain active learning: A comparative study. arXiv preprint arXiv:2106.13516.
  88. 88.Hideitsu Hino. 2020. Active learning: Problem settings and recent developments. arXiv preprint arXiv:2012.04225.
  89. 89.Victoria Hodge and Jim Austin. 2004. A survey of outlier detection methodologies. Artificial intelligence review, 22(2):85–126.
  90. 90.Andrea Horbach and Alexis Palmer. 2016. Investigating active learning for short-answer scoring. In Proceedings of the 11th Workshop on Innovative Use of NLP for Building Educational Applications, pages 301–311, San Diego, CA. Association for Computational Linguistics.
  91. 91.Neil Houlsby, Ferenc Huszár, Zoubin Ghahramani, and Máté Lengyel. 2011. Bayesian active learning for classification and preference learning. arXiv preprint arXiv:1112.5745.
  92. 92.Junjie Hu and Graham Neubig. 2021. Phrase-level active learning for neural machine translation. In Proceedings of the Sixth Conference on Machine Translation, pages 1087–1099, Online. Association for Computational Linguistics.
  93. 93.Peiyun Hu, Zack Lipton, Anima Anandkumar, and Deva Ramanan. 2019. Active learning with partial feedback. In International Conference on Learning Representations.
  94. 94.Rong Hu, Brian Mac Namee, and Sarah Jane Delany. 2010. Off to a good start: Using clustering to select the initial training set in active learning. In Twenty-Third International FLAIRS Conference.

Citation

MLA
Zhang, Z., et al. “A Survey of Active Learning for Natural Language Processing”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 6166–90, https://doi.org/10.18653/v1/2022.emnlp-main.414.
APA
Zhang, Z., Strubell, E., & Hovy, E. (2022). A Survey of Active Learning for Natural Language Processing. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 6166–6190. https://doi.org/10.18653/v1/2022.emnlp-main.414
Chicago
Zhang, Z., E. Strubell, and E. Hovy. 2022. “A Survey of Active Learning for Natural Language Processing”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 6166–90. https://doi.org/10.18653/v1/2022.emnlp-main.414.
Harvard
Zhang, Z., Strubell, E. and Hovy, E. (2022) “A Survey of Active Learning for Natural Language Processing”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 6166–6190. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.414.
Vancouver
1. Zhang Z, Strubell E, Hovy E (2022) A Survey of Active Learning for Natural Language Processing. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 6166–6190

BibTeX

@inproceedings{zhang-etal-2022-survey,
    title = "A Survey of Active Learning for Natural Language Processing",
    author = "Zhang, Zhisong  and
      Strubell, Emma  and
      Hovy, Eduard",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.414/",
    doi = "10.18653/v1/2022.emnlp-main.414",
    pages = "6166--6190"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/