Get another label? improving data quality and data mining using multiple, noisy labelers

Victor S. ShengFoster ProvostPanagiotis G. Ipeirotis

article2008KDD1,325 citationsKDD 2008 Best Research Paper Award (Runner-Up)

Demonstrates when and how repeatedly acquiring noisy labels from crowdsourced annotators improves supervised model performance more cost-effectively than collecting single labels for new instances.

Listen

Modern data science workflows increasingly rely on online outsourcing platforms to collect labels for machine learning tasks. While these services make label acquisition inexpensive, the labels provided by non-expert workers are often noisy and imperfect. At the same time, gathering and preparing the underlying raw examples often remains costly. Organizations face a critical resource-allocation problem: when labeling is error-prone, should a fixed budget be spent on collecting brand-new data points or on obtaining multiple, repeated labels for existing data to improve data quality?

The article evaluates whether, when, and how repeatedly acquiring noisy labels improves training data quality and the performance of supervised learning models. It specifically investigates the economic trade-offs between acquiring new examples versus refining existing ones and develops targeted strategies to guide which data points should receive additional labels.

The researchers conducted an analytical assessment of majority voting across varying worker quality levels and ran extensive empirical simulations across 12 real-world benchmark datasets. They simulated label acquisition under varying degrees of label noise, compared standard single-label collection against round-robin and selective re-labeling strategies, and tested two methods for integrating multi-label data: simple majority voting and an uncertainty-preserving weighting approach.

The findings demonstrate clear, quantifiable advantages for repeated labeling. First, when data preparation carries any noticeable cost relative to label acquisition, round-robin repeated labeling consistently delivers higher classification accuracy per dollar spent than traditional single labeling. Second, selectively choosing which examples to re-label provides substantial performance gains over round-robin approaches; a hybrid method combining label uncertainty and predictive model uncertainty achieved the highest average accuracy across all 12 datasets (averaging roughly 79.5% accuracy in high-noise environments compared to 73.7% for round-robin). Third, when worker accuracy is low (e.g., 60% accuracy on binary tasks), preserving label uncertainty via instance weighting outperforms majority voting, whereas majority voting suffices when worker quality is high. Finally, repeated labeling yields diminishing returns when worker accuracy is very high or when the training set is already small and positioned on a steep learning curve.

These results show that organizations can significantly improve model accuracy without increasing budgets by strategically allocating multiple low-cost workers to individual examples rather than constantly gathering new data. In contrast to traditional active learning, which assumes labels are costly and unlabeled data is free, low-cost crowdsourcing makes data preparation the primary bottleneck. Selective re-labeling mitigates this bottleneck and lowers project risk by prioritizing items with ambiguous label distributions and borderline model confidence.

Organizations deploying crowdsourced annotation should adopt repeated labeling and prioritize examples using combined label-and-model uncertainty algorithms rather than uniform round-robin allocation. When managing high worker noise, teams should utilize uncertainty-preserving data weighting rather than forcing hard majority votes. For complex operational settings, teams should pilot dynamic allocation schemes that balance the cost of new data against repeated annotations. These findings are established under benchmark conditions assuming equal worker quality and constant labeling difficulty per instance; decision-makers should exercise caution when dealing with correlated worker errors or varying task difficulties until further validation is conducted.

Cover for Get another label? improving data quality and data mining using multiple, noisy labelers

Abstract

This paper addresses the repeated acquisition of labels for data items when the labeling is imperfect. We examine the improvement (or lack thereof) in data quality via repeated labeling, and focus especially on the improvement of training labels for supervised induction. With the outsourcing of small tasks becoming easier, for example via Rent-A-Coder or Amazon's Mechanical Turk, it often is possible to obtain less-than-expert labeling at low cost. With low-cost labeling, preparing the unlabeled part of the data can become considerably more expensive than labeling. We present repeated-labeling strategies of increasing complexity, and show several main results. (i) Repeated-labeling can improve label quality and model quality, but not always. (ii) When labels are noisy, repeated labeling can be preferable to single labeling even in the traditional setting where labels are not particularly cheap. (iii) As soon as the cost of processing the unlabeled data is not free, even the simple strategy of labeling everything multiple times can give considerable advantage. (iv) Repeatedly labeling a carefully chosen set of points is generally preferable, and we present a robust technique that combines different notions of uncertainty to select data points for which quality should be improved. The bottom line: the results show clearly that when labeling is not perfect, selective acquisition of multiple labels is a strategy that data miners should have in their repertoire; for certain label-quality/cost regimes, the benefit is substantial.

Table of Contents

  • 1. INTRODUCTION
  • Categories and Subject Descriptors
  • General Terms
  • Keywords
  • 2. RELATED WORK
  • 3. REPEATED LABELING: THE BASICS
  • 3.1 Notation and Assumptions
  • 3.2 Majority Voting and Label Quality
  • 3.2.1 Uniform Labeler Quality
  • 3.2.2 Different Labeler Quality
  • 3.3 Uncertainty-preserving Labeling
  • 4. REPEATED-LABELING AND MODELING
  • 4.1 Experimental Setup
  • 4.2 Generalized Round-robin Strategies
  • 4.2.1 Round-robin Strategies, C_U ≪ C_L
  • 4.2.2 Round-robin Strategies, General Costs
  • 4.3 Selective Repeated-Labeling
  • 4.3.1 What Not To Do
  • 4.3.2 Estimating Label Uncertainty
  • 4.3.3 Using Model Uncertainty
  • 4.3.4 Model Performance with Selective ML
  • 5. CONCLUSIONS, LIMITATIONS, AND FUTURE WORK
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Combined Label and Model Uncertainty for Selective Repeated Labeling

    model/method

    The Label and Model Uncertainty (LMU) strategy selects instances for which to acquire additional noisy training labels by combining multiset label uncertainty with learner model uncertainty. This prevents acquiring labels either when the existing multiset of noisy labels is already statistically conclusive about the true label, or when the predictive model is already highly confident in its prediction for that instance.

    Let SLUS_{LU} denote the label uncertainty score of instance xx computed from the incomplete Beta cumulative distribution of the current label multiset, and let SMUS_{MU} denote the model uncertainty score of xx computed across an ensemble of mm learned models H1,…,HmH_1, \dots, H_m:

    SMU=0.5−∣1m∑i=1mPr⁡(+∣x,Hi)−0.5∣S_{MU} = 0.5 - \left| \frac{1}{m} \sum_{i=1}^m \Pr(+ \mid x, H_i) - 0.5 \right|

    The combined uncertainty score SLMUS_{LMU} is defined as the geometric mean of the two individual uncertainty scores:

    SLMU=SMU⋅SLUS_{LMU} = \sqrt{S_{MU} \cdot S_{LU}}

    Instances with the highest SLMUS_{LMU} values are prioritized for repeated label acquisition.

  2. Knowl 2 — Bayesian Estimation of Label Uncertainty for Repeated Labeling

    model/method

    To measure label uncertainty without assuming known individual labeler accuracies, the Label Uncertainty (LU) method uses a Bayesian formulation over the true class probability. Assuming a uniform prior Beta(1,1)\text{Beta}(1, 1) on the probability of a positive label, observing LposL_{\text{pos}} positive labels and LnegL_{\text{neg}} negative labels for an instance yields a posterior distribution Beta(Lpos+1,Lneg+1)\text{Beta}(L_{\text{pos}} + 1, L_{\text{neg}} + 1).

    The label uncertainty score SLUS_{LU} is the tail probability that the majority label assignment is incorrect, evaluated at the classification decision threshold x=0.5x = 0.5 using the regularized incomplete beta function Ix(α,β)I_x(\alpha, \beta):

    Ix(α,β)=∑j=αα+β−1(α+β−1)!j!(α+β−1−j)!xj(1−x)α+β−1−jI_x(\alpha, \beta) = \sum_{j=\alpha}^{\alpha+\beta-1} \frac{(\alpha+\beta-1)!}{j!(\alpha+\beta-1-j)!} x^j (1-x)^{\alpha+\beta-1-j}

    With parameters α=Lpos+1\alpha = L_{\text{pos}} + 1 and β=Lneg+1\beta = L_{\text{neg}} + 1, the uncertainty score is:

    SLU=min⁡{I0.5(Lpos+1,Lneg+1),1−I0.5(Lpos+1,Lneg+1)}S_{LU} = \min\left\{ I_{0.5}(L_{\text{pos}} + 1, L_{\text{neg}} + 1), 1 - I_{0.5}(L_{\text{pos}} + 1, L_{\text{neg}} + 1) \right\}

    This metric penalizes small sample sizes (where unanimity does not imply high certainty) and assigns near-zero uncertainty to large multisets with strong majority support.

  3. Knowl 3 — Benchmark Comparison of Selective Repeated-Labeling Strategies

    data/table

    The table below presents the average classification accuracy (in percent) over acquisition iterations for four repeated-labeling strategies across 12 binary classification datasets under high label noise (p=0.6p = 0.6). All methods began with 3 initial labels per training instance and acquired additional labels in increments of 2. Classifiers were induced using decision trees (C4.5 / Weka J48) and evaluated on a held-out 30% test set with true ground-truth labels across 10 random runs.

    Data Set GRR MU LU LMU
    bmg 62.97 71.90 64.82 68.93
    expedia 80.61 84.72 81.72 85.01
    kr-vs-kp 76.75 76.71 81.25 82.55
    mushroom 89.07 94.17 92.56 95.52
    qvc 64.67 76.12 66.88 74.54
    sick 88.50 93.72 91.06 93.75
    spambase 72.79 79.52 77.04 80.69
    splice 69.76 68.16 73.23 73.06
    thyroid 89.54 93.59 92.12 93.97
    tic-tac-toe 59.59 62.87 61.96 62.91
    travelocity 64.29 73.94 67.18 72.31
    waveform 65.34 69.88 66.36 70.24
    Average 73.65 78.77 76.35 79.46

    The results show that:

    1. Selective labeling guided by label uncertainty alone (LU) or model uncertainty alone (MU) outperforms generalized round-robin repeated labeling (GRR).
    2. Combining label and model uncertainty (LMU) achieves the highest grand average accuracy (79.46%) and consistently outperforms GRR on every individual dataset (statistically significant by a one-tailed sign test at p<0.1p < 0.1).
  4. Knowl 4 — Integrated Label Quality under Majority Voting with Homogeneous Labelers

    theoretical result

    For a binary classification task where 2N+12N + 1 independent labelers each assign a label with identical accuracy p=Pr⁡(yij=yi)p = \Pr(y_{ij} = y_i), the integrated label y^i\hat{y}_i obtained by majority voting has an integrated quality q=Pr⁡(y^i=yi)q = \Pr(\hat{y}_i = y_i) given by:

    q=∑i=0N(2N+1i)p2N+1−i(1−p)iq = \sum_{i=0}^N \binom{2N+1}{i} p^{2N+1-i} (1-p)^i

    where ii index the count of incorrect votes among the 2N+12N+1 labelers.

    From this formulation:

    • When individual labeler quality exceeds chance (p>0.5p > 0.5), majority voting strictly increases label quality (q>pq > p), though marginal gains diminish as NN grows.
    • When p<0.5p < 0.5 (an adversarial setting), majority voting reduces quality (q<pq < p), with quality degrading further as the number of labelers increases.
    • The marginal benefit of repeated labeling is largest for moderate noise levels (e.g., p=0.7p = 0.7), where increasing from 1 to 3 labelers improves integrated quality by approximately 0.10.1.
  5. Knowl 5 — Integrated Quality of Three Heterogeneous Labelers vs. Single Best Labeler

    theoretical result

    Consider three independent labelers with symmetric quality dispersion: a middle quality pp, a lowest quality p−dp - d, and a highest quality p+dp + d, where d≥0d \ge 0 represents the quality deviation. The integrated label quality qq under majority voting across all three labelers is:

    q=−2p3+2pd2+3p2−d2q = -2p^3 + 2pd^2 + 3p^2 - d^2

    Using majority voting across the triplet yields higher label quality than relying solely on the single best individual labeler (which achieves quality p+dp + d) if and only if:

    −2p3+2pd2+3p2−d2>p+d-2p^3 + 2pd^2 + 3p^2 - d^2 > p + d

    When the quality spread dd is small relative to pp, pooling multiple noisy labelers outperforms the best single expert; when dd exceeds a critical threshold for a given pp, using the single highest-quality labeler is strictly superior.

  6. Knowl 6 — Multiplied Examples Procedure for Uncertainty-Preserving Learning

    model/method

    The Multiplied Examples (ME) method enables standard supervised induction algorithms to learn from the soft label distributions of repeated noisy labeling without requiring native algorithmic support for probabilistic targets.

    For each unlabeled training instance xix_i with an associated multiset of acquired labels Li={yij}L_i = \{y_{ij}\}:

    1. ME creates one replica of xix_i for each unique class label present in LiL_i.
    2. For each replica with class label yy, ME assigns an instance weight equal to the relative frequency of that label in the multiset:

    wi,y=N(y,Li)∣Li∣w_{i, y} = \frac{N(y, L_i)}{|L_i|}

    where N(y,Li)N(y, L_i) is the number of times label yy appears in LiL_i, and ∣Li∣|L_i| is the total number of labels collected for instance xix_i.

    These weighted replicas can be consumed directly by cost-sensitive/weighted learning algorithms (e.g., weighted decision trees) or reduced to uniform-weighted datasets via importance weighting reductions.

  7. Knowl 7 — Failure of Multiset Purity and Entropy for Selective Re-labeling

    theoretical result

    Heuristic selection criteria that allocate repeated labels based on empirical multiset impurity—such as multiset entropy or proximity of majority label frequency to 0.5 (even with Laplace corrections)—fail under noisy labeling conditions and are outperformed in the long run by simple round-robin labeling.

    This pathology occurs because multiset purity does not reflect true posterior uncertainty under noise:

    1. False Certainty on Small Samples: A small multiset with unanimous votes (e.g., {+,+,+}\{+, +, +\}) appears perfectly pure (00 entropy), so the strategy never requests additional labels, even though under moderate noise (p=0.6p = 0.6) the true label remains highly uncertain.
    2. Resource Waste on Large Samples: An instance with many noisy labels (e.g., 600600 positive and 400400 negative votes at p=0.6p = 0.6) maintains an impure label mixture and continuously consumes labeling resources, despite the true class being statistically certain beyond reasonable doubt.
  8. Knowl 8 — Data Acquisition Cost Formulation for Repeated Labeling

    definition

    In supervised induction where acquiring the feature vector xix_i of an unlabeled instance incurs cost CUC_U and acquiring a noisy label yijy_{ij} incurs cost CLC_L, the total data acquisition cost CDC_D is:

    CD=CU⋅Tr+CL⋅NLC_D = C_U \cdot T_r + C_L \cdot N_L

    where TrT_r is the number of distinct training instances acquired and NLN_L is the total number of labels obtained across all instances.

    For single labeling, NL=TrN_L = T_r. For generalized round-robin repeated labeling where each acquired instance receives a fixed number kk labels, NL=k⋅TrN_L = k \cdot T_r. The cost regime is governed by the cost ratio ρ=CUCL\rho = \frac{C_U}{C_L}, yielding:

    CD=(ρ+k)⋅CL⋅Tr∝ρ+kC_D = (\rho + k) \cdot C_L \cdot T_r \propto \rho + k

    This formulation allows fair evaluation of generalization performance per unit cost between single-labeling and repeated-labeling schemes.

  9. Knowl 9 — Cost-Effectiveness Regimes for Repeated Labeling over Single Labeling

    empirical result

    Repeated noisy labeling outperforms single labeling of new training instances under two primary regimes:

    1. Non-negligible Feature Acquisition Cost (ρ=CU/CL>0\rho = C_U / C_L > 0): As the cost of obtaining unlabeled features CUC_U grows relative to label cost CLC_L, round-robin repeated labeling with majority voting consistently yields higher generalization accuracy per unit cost than acquiring single-labeled new instances.
    2. Flat Regions of the Learning Curve: Even when unlabeled instances are free (CU=0C_U = 0), repeated labeling of existing instances is preferable to single-labeling new instances once a sufficient baseline of training data has been acquired. In flat or plateauing regions of the learning curve, shifting the effective label quality qq upward via repeated labeling yields a larger marginal gain in test accuracy than increasing sample size with noisy labels.
  10. Knowl 10 — Noise-Dependent Advantage of Uncertainty-Preserving Labeling over Majority Voting

    empirical result

    The relative advantage of uncertainty-preserving Multiplied Examples (ME) over hard Majority Voting (MV) depends on the underlying labeler accuracy pp and the data acquisition budget:

    1. High-Noise Regime (p=0.6p = 0.6): ME consistently matches or outperforms MV across datasets, with its accuracy advantage widening as more training instances and labels are acquired. Explicitly representing the empirical label distribution prevents the loss of uncertainty information inherent in hard thresholding.
    2. Low-Noise Regime (p=0.8p = 0.8): The performance advantage of ME over MV disappears, and both methods achieve nearly identical generalization accuracy. When labeler quality is high, multiset uncertainty is minimal, making majority voting sufficient.

Coverage note — None was omitted. All primary models, algorithms (LMU, LU, MU, ME), analytical results (majority voting quality equations for uniform and heterogeneous labelers), data acquisition cost formulations, and experimental benchmark tables from the paper are represented.

References

  1. 1.Baram, Y., El-Yaniv, R., and Luz, K. Online choice of active learning algorithms. Journal of Machine Learning Research 5 (Mar. 2004), 255–291.
  2. 2.Blake, C. L., and Merz, C. J. UCI repository of machine learning databases. http://www.ics.uci.edu/~mlearn/MLRepository.html, 1998.
  3. 3.Boutell, M. R., Luo, J., Shen, X., and Brown, C. M. Learning multi-label scene classification. Pattern Recognition 37, 9 (Sept. 2004), 1757–1771.
  4. 4.Breiman, L. Random forests. Machine Learning 45, 1 (Oct. 2001), 5–32.
  5. 5.Cohn, D. A., Atlas, L. E., and Ladner, R. E. Improving generalization with active learning. Machine Learning 15, 2 (May 1994), 201–221.
  6. 6.Dawid, A. P., and Skene, A. M. Maximum likelihood estimation of observer error-rates using the EM algorithm. Applied Statistics 28, 1 (Sept. 1979), 20–28.
  7. 7.Domingos, P. MetaCost: A general method for making classifiers cost-sensitive. In KDD (1999), pp. 155–164.
  8. 8.Elkan, C. The foundations of cost-sensitive learning. In IJCAI (2001), pp. 973–978.
  9. 9.Gelman, A., Carlin, J. B., Stern, H. S., and Rubin, D. B. Bayesian Data Analysis, 2nd ed. Chapman and Hall/CRC, 2003.
  10. 10.Jin, R., and Ghahramani, Z. Learning with multiple labels. In NIPS (2002), pp. 897–904.
  11. 11.Kapoor, A., and Greiner, R. Learning and classifying under hard budgets. In ECML (2005), pp. 170–181.
  12. 12.Lizotte, D. J., Madani, O., and Greiner, R. Budgeted learning of naive-bayes classifiers. In UAI) (2003), pp. 378–385.
  13. 13.Lugosi, G. Learning with an unreliable teacher. Pattern Recognition 25, 1 (Jan. 1992), 79–87.
  14. 14.Margineantu, D. D. Active cost-sensitive learning. In IJCAI) (2005), pp. 1622–1613.
  15. 15.McCallum, A. Multi-label text classification with a mixture model trained by EM. In AAAI’99 Workshop on Text Learning (1999).
  16. 16.Melville, P., Provost, F. J., and Mooney, R. J. An expected utility approach to active feature-value acquisition. In ICDM (2005), pp. 745–748.
  17. 17.Melville, P., Saar-Tsechansky, M., Provost, F. J., and Mooney, R. J. Active feature-value acquisition for classifier induction. In ICDM (2004), pp. 483–486.
  18. 18.Morrison, C. T., and Cohen, P. R. Noisy information value in utility-based decision making. In UBDM’05: Proceedings of the First International Workshop on Utility-based Data Mining (2005), pp. 34–38.
  19. 19.Provost, F. Toward economic machine learning and utility-based data mining. In UBDM ’05: Proceedings of the 1st International Workshop on Utility-based Data Mining (2005), pp. 1–1.
  20. 20.Provost, F., and Danyluk, A. Learning from Bad Data. In Proceedings of the ML-95 Workshop on Applying Machine Learning in Practice (1995).
  21. 21.Quinlan, J. R. C4.5: Programs for Machine Learning. Morgan Kaufmann Publishers, Inc., 1992.
  22. 22.Saar-Tsechansky, M., Melville, P., and Provost, F. J. Active feature-value acquisition. Tech. Rep. IROM-08-06, University of Texas at Austin, McCombs Research Paper Series, Sept. 2007.
  23. 23.Saar-Tsechansky, M., and Provost, F. Active sampling for class probability estimation and ranking. Journal of Artificial Intelligence Research 54, 2 (2004), 153–178.
  24. 24.Silverman, B. W. Some asymptotic properties of the probabilistic teacher. IEEE Transactions on Information Theory 26, 2 (Mar. 1980), 246–249.
  25. 25.Smyth, P. Learning with probabilistic supervision. In Computational Learning Theory and Natural Learning Systems, Vol. III: Selecting Good Models, T. Petsche, Ed. MIT Press, Apr. 1995.
  26. 26.Smyth, P. Bounds on the mean classification error rate of multiple experts. Pattern Recognition Letters 17, 12 (May 1996).
  27. 27.Smyth, P., Burl, M. C., Fayyad, U. M., and Perona, P. Knowledge discovery in large image databases: Dealing with uncertainties in ground truth. In Knowledge Discovery in Databases: Papers from the 1994 AAAI Workshop (KDD-94) (1994), pp. 109–120.
  28. 28.Smyth, P., Fayyad, U. M., Burl, M. C., Perona, P., and Baldi, P. Inferring ground truth from subjective labelling of Venus images. In NIPS (1994), pp. 1085–1092.
  29. 29.Ting, K. M. An instance-weighting method to induce cost-sensitive trees. IEEE Transactions on Knowledge and Data Engineering 14, 3 (Mar. 2002), 659–665.
  30. 30.Turney, P. D. Cost-sensitive classification: Empirical evaluation of a hybrid genetic decision tree induction algorithm. Journal of Artificial Intelligence Research 2 (1995), 369–409.
  31. 31.Turney, P. D. Types of cost in inductive concept learning. In Proceedings of the ICML-2000 Workshop on Cost-Sensitive Learning (2000), pp. 15–21.
  32. 32.Weiss, G. M., and Provost, F. J. Learning when training data are costly: The effect of class distribution on tree induction. Journal of Artificial Intelligence Research 19 (2003), 315–354.
  33. 33.Whittle, P. Some general points in the theory of optimal experimental design. Journal of the Royal Statistical Society, Series B (Methodological) 35, 1 (1973), 123–130.
  34. 34.Witten, I. H., and Frank, E. Data Mining: Practical Machine Learning Tools and Techniques, 2nd ed. Morgan Kaufmann Publishing, June 2005.
  35. 35.Zadrozny, B., Langford, J., and Abe, N. Cost-sensitive learning by cost-proportionate example weighting. In ICDM (2003), pp. 435–442.
  36. 36.Zheng, Z., and Padmanabhan, B. Selectively acquiring customer information: A new data acquisition problem and an active learning-based solution. Management Science 52, 5 (May 2006), 697–712.
  37. 37.Zhu, X., and Wu, X. Cost-constrained data acquisition for intelligent data preparation. IEEE TKDE 17, 11 (Nov. 2005), 1542–1556.

Citation

MLA
Sheng, V. S., et al. “Get Another Label? Improving Data Quality and Data Mining Using Multiple, Noisy Labelers”. Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2008, pp. 614–22, https://doi.org/10.1145/1401890.1401965.
APA
Sheng, V. S., Provost, F., & Ipeirotis, P. G. (2008). Get another label? improving data quality and data mining using multiple, noisy labelers. Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 614–622. https://doi.org/10.1145/1401890.1401965
Chicago
Sheng, V. S., F. Provost, and P. G. Ipeirotis. 2008. “Get Another Label? Improving Data Quality and Data Mining Using Multiple, Noisy Labelers”. Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 614–22. https://doi.org/10.1145/1401890.1401965.
Harvard
Sheng, V.S., Provost, F. and Ipeirotis, P.G. (2008) “Get another label? improving data quality and data mining using multiple, noisy labelers”, Proceedings of the 14th ACM SIGKDD international conference on Knowledge discovery and data mining. ACM, pp. 614–622. Available at: https://doi.org/10.1145/1401890.1401965.
Vancouver
1. Sheng VS, Provost F, Ipeirotis PG (2008) Get another label? improving data quality and data mining using multiple, noisy labelers. In: Proceedings of the 14th ACM SIGKDD international conference on Knowledge discovery and data mining. ACM, pp 614–622

BibTeX

@inproceedings{Sheng_2008, series={KDD08}, title={Get another label? improving data quality and data mining using multiple, noisy labelers}, url={http://dx.doi.org/10.1145/1401890.1401965}, DOI={10.1145/1401890.1401965}, booktitle={Proceedings of the 14th ACM SIGKDD international conference on Knowledge discovery and data mining}, publisher={ACM}, author={Sheng, Victor S. and Provost, Foster and Ipeirotis, Panagiotis G.}, year={2008}, month=Aug, pages={614–622}, collection={KDD08} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF