Get another label? improving data quality and data mining using multiple, noisy labelers
Victor S. ShengFoster ProvostPanagiotis G. Ipeirotis
Demonstrates when and how repeatedly acquiring noisy labels from crowdsourced annotators improves supervised model performance more cost-effectively than collecting single labels for new instances.
Modern data science workflows increasingly rely on online outsourcing platforms to collect labels for machine learning tasks. While these services make label acquisition inexpensive, the labels provided by non-expert workers are often noisy and imperfect. At the same time, gathering and preparing the underlying raw examples often remains costly. Organizations face a critical resource-allocation problem: when labeling is error-prone, should a fixed budget be spent on collecting brand-new data points or on obtaining multiple, repeated labels for existing data to improve data quality?
The article evaluates whether, when, and how repeatedly acquiring noisy labels improves training data quality and the performance of supervised learning models. It specifically investigates the economic trade-offs between acquiring new examples versus refining existing ones and develops targeted strategies to guide which data points should receive additional labels.
The researchers conducted an analytical assessment of majority voting across varying worker quality levels and ran extensive empirical simulations across 12 real-world benchmark datasets. They simulated label acquisition under varying degrees of label noise, compared standard single-label collection against round-robin and selective re-labeling strategies, and tested two methods for integrating multi-label data: simple majority voting and an uncertainty-preserving weighting approach.
The findings demonstrate clear, quantifiable advantages for repeated labeling. First, when data preparation carries any noticeable cost relative to label acquisition, round-robin repeated labeling consistently delivers higher classification accuracy per dollar spent than traditional single labeling. Second, selectively choosing which examples to re-label provides substantial performance gains over round-robin approaches; a hybrid method combining label uncertainty and predictive model uncertainty achieved the highest average accuracy across all 12 datasets (averaging roughly 79.5% accuracy in high-noise environments compared to 73.7% for round-robin). Third, when worker accuracy is low (e.g., 60% accuracy on binary tasks), preserving label uncertainty via instance weighting outperforms majority voting, whereas majority voting suffices when worker quality is high. Finally, repeated labeling yields diminishing returns when worker accuracy is very high or when the training set is already small and positioned on a steep learning curve.
These results show that organizations can significantly improve model accuracy without increasing budgets by strategically allocating multiple low-cost workers to individual examples rather than constantly gathering new data. In contrast to traditional active learning, which assumes labels are costly and unlabeled data is free, low-cost crowdsourcing makes data preparation the primary bottleneck. Selective re-labeling mitigates this bottleneck and lowers project risk by prioritizing items with ambiguous label distributions and borderline model confidence.
Organizations deploying crowdsourced annotation should adopt repeated labeling and prioritize examples using combined label-and-model uncertainty algorithms rather than uniform round-robin allocation. When managing high worker noise, teams should utilize uncertainty-preserving data weighting rather than forcing hard majority votes. For complex operational settings, teams should pilot dynamic allocation schemes that balance the cost of new data against repeated annotations. These findings are established under benchmark conditions assuming equal worker quality and constant labeling difficulty per instance; decision-makers should exercise caution when dealing with correlated worker errors or varying task difficulties until further validation is conducted.
- Paper: A sequential algorithm for training text classifiers, David D. Lewis et al. (1994). It introduces the foundational concept of uncertainty sampling for active data acquisition, providing the theoretical basis for selectively acquiring labels based on model uncertainty.
- Paper: Active Learning with Statistical Models, David Cohn et al. (1996). It establishes statistical frameworks for variance reduction and optimal data selection in active learning that underpin selective labeling strategies.
- Paper: Query by committee, H. Seung et al. (1992). It formalizes ensemble disagreement as an informative criterion for query-based active learning, motivating multi-notion uncertainty measures.
- Paper: Improving Generalization with Active Learning, David Cohn et al. (1994). It provides foundational principles for selective sampling by querying regions where models exhibit highest uncertainty and disagreement.
- Paper: Neural Network Ensembles, Cross Validation, and Active Learning, Anders Krogh et al. (1994). It establishes the link between ensemble ambiguity, prediction error, and active learning query selection.
- Paper: Support Vector Machine Active Learning with Applications to Text Classification, Simon Tong et al. (2001). It demonstrates practical pool-based active learning and margin-based uncertainty sampling for text classification.
- Paper: MetaCost: a general method for making classifiers cost-sensitive, Pedro M. Domingos (1999). It demonstrates how relabeling training instances based on ensemble probability estimates can optimize downstream model quality under asymmetric costs.
- Paper: Learning From Crowds, V. Raykar et al. (2010). It develops a unified probabilistic framework to jointly infer true labels, evaluate annotator quality, and train classifiers directly from multiple noisy crowdsourced annotations.
- Paper: Whose Vote Should Count More: Optimal Integration of Labels from Labelers of Unknown Expertise, Jacob Whitehill et al. (2009). It advances repeated noisy labeling by modeling individual annotator expertise and task difficulty simultaneously when aggregating votes.
- Paper: Cheap and Fast – But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks, Rion Snow et al. (2008). It provides comprehensive empirical validation showing that aggregating repeated non-expert annotations matches expert-level data quality across multiple language tasks.
- Paper: Snorkel: Rapid Training Data Creation with Weak Supervision, Alexander J. Ratner et al. (2017). It generalizes multi-labeler aggregation by modeling noisy, overlapping programmatic labeling functions without requiring ground truth.
- Paper: Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach, Giorgio Patrini et al. (2016). It formalizes loss correction techniques for deep neural networks trained under class-conditional label noise resulting from imperfect or outsourced annotations.
- Paper: Learning From Noisy Labels With Deep Neural Networks: A Survey, Hwanjun Song et al. (2020). It surveys modern methodologies for training neural networks on noisy labels generated by crowdsourcing and web scraping.
- Paper: DivideMix: Learning with Noisy Labels as Semi-supervised Learning, Junnan Li et al. (2020). It extends learning from noisy data by dynamically separating reliable labels from noisy samples and leveraging semi-supervised learning.
- Paper: Learning to Reweight Examples for Robust Deep Learning, Mengye Ren et al. (2018). It introduces meta-learning mechanisms to dynamically reweight training examples to make models robust against imperfect and noisy labeling.
