Whose Vote Should Count More: Optimal Integration of Labels from Labelers of Unknown Expertise

Jacob WhitehillPaul RuvoloTingfan WuJacob BergsmaJavier Movellan

article2009NeurIPS1,347 citations

Introduces the GLAD probabilistic model to simultaneously infer true data labels, annotator expertise, and item difficulty from crowdsourced annotations, outperforming majority voting even in the presence of noisy or adversarial labelers.

Listen

Modern machine learning applications demand vast volumes of labeled data, creating a major operational bottleneck. While online crowdsourcing platforms offer access to distributed human workers at low cost, this labor pool presents significant reliability challenges: contributors possess unknown and widely varying skill levels, some act maliciously or consistently misunderstand tasks, and the difficulty of individual items varies. Standard aggregation techniques such as simple majority voting fail to account for these nuances and frequently misclassify difficult items or succumb to poor-quality labelers.

The article introduces and evaluates a probabilistic model called GLAD (Generative model of Labels, Abilities, and Difficulties). The primary objective is to simultaneously infer the true label of each item, the expertise of each labeler, and the difficulty of each item without requiring advance knowledge of worker abilities or an initial answer key.

To achieve this, the article utilizes a maximum likelihood estimation framework optimized via the Expectation-Maximization algorithm. The method was validated using synthetic benchmarks of up to 2,000 images and 50 simulated workers, as well as real-world crowdsourced datasets: 100 synthetic perceptual images (Greebles) and 160 facial images evaluated for subtle emotional expressions (Duchenne smiles) by 20 online workers yielding 3,572 annotations against certified expert benchmarks.

The findings establish clear advantages over existing aggregation methods. First, the model consistently outperforms majority voting across both synthetic and real-world experiments; in the facial expression study, it achieved 78.12% accuracy compared to 71.88% for majority voting, delivering a roughly 6% performance gain. Second, in controlled simulations modeling image difficulty, GLAD reduced classification error to 4.5%, compared to 11.2% for majority voting and 8.4% for models that account only for worker skill. Third, the model demonstrated resilience against corrupted inputs; it maintained stable accuracy when subjected to thousands of purely random labels and successfully recovered true classifications by automatically detecting and flipping the votes of adversarial labelers. Finally, the algorithm scaled linearly with data volume, processing one million image labels in approximately 10 minutes on a standard single-core processor.

These results demonstrate that organizations can lower data acquisition costs and improve training data quality by weighting inputs according to automatically estimated expertise rather than treating all votes equally. Because GLAD requires fewer total annotations per item to achieve high confidence, it reduces expenditure on crowdsourced labor and mitigates the risk of erroneous training data caused by unskilled or dishonest contributors.

Organizations managing large data annotation pipelines should adopt probabilistic aggregation models in place of simple majority heuristics. The article suggests integrating active sampling strategies next, using the model's confidence estimates to dynamically select which items need additional labeling and which workers require further evaluation. While the model is highly effective and computationally stable across varied initializations, current validations remain bounded within binary classification settings; practitioners applying the method to broader use cases should consider pilot testing as extensions to multi-class or continuous tasks develop.

Whitehill et al (2009).pdf
Cover for Whose Vote Should Count More: Optimal Integration of Labels from Labelers of Unknown Expertise

Abstract

Modern machine learning-based approaches to computer vision require very large databases of hand labeled images. Some contemporary vision systems already require on the order of millions of images for training (e.g., Omron face detector [9]). New Internet-based services allow for a large number of labelers to collaborate around the world at very low cost. However, using these services brings interesting theoretical and practical challenges: (1) The labelers may have wide ranging levels of expertise which are unknown a priori, and in some cases may be adversarial; (2) images may vary in their level of difficulty; and (3) multiple labels for the same image must be combined to provide an estimate of the actual label of the image. Probabilistic approaches provide a principled way to approach these problems. In this paper we present a probabilistic model and use it to simultaneously infer the label of each image, the expertise of each labeler, and the difficulty of each image. On both simulated and real data, we demonstrate that the model outperforms the commonly used “Majority Vote” heuristic for inferring image labels, and is robust to both noisy and adversarial labelers.

Table of Contents

  • 1 Introduction
  • 2 Modeling the Labeling Process
  • 3 Inference
  • 3.1 Priors on α, β
  • 3.2 Computational Complexity
  • 4 Simulations
  • 4.1 Stability of EM under Various Starting Points
  • 5 Empirical Study I: Greebles
  • Inferred Label Accuracy of Greeble Images
  • 6 Empirical Study II: Duchenne Smiles
  • Duchenne Smiles
  • Non-Duchenne Smiles
  • 7 Related Work
  • 8 Summary and Further Research
  • References

Knowls

  1. Knowl 1 — Generative Model of Labels, Abilities, and Difficulties (GLAD)

    model/method

    The Generative model of Labels, Abilities, and Difficulties (GLAD) is a probabilistic generative model designed to infer true unobserved binary labels Zj∈{0,1}Z_j \in \{0, 1\} for nn items j∈{1,…,n}j \in \{1, \dots, n\} from labels provided by mm labelers of unknown expertise.

    The generative process accounts for two continuous latent factors:

    • Labeler expertise αi∈(−∞,+∞)\alpha_i \in (-\infty, +\infty): where αi=+∞\alpha_i = +\infty denotes a perfectly accurate labeler, αi=0\alpha_i = 0 denotes an uninformative labeler guessing randomly, and αi<0\alpha_i < 0 denotes an adversarial or misunderstanding labeler who systematically inverts labels.
    • Item inverse-difficulty (ease) βj∈(0,+∞)\beta_j \in (0, +\infty): where 1/βj=01/\beta_j = 0 (i.e., βj→∞\beta_j \to \infty) represents an item so easy that any non-adversarial annotator labels it correctly, and 1/βj→∞1/\beta_j \to \infty (i.e., βj→0\beta_j \to 0) represents an ambiguous item where even expert labelers have a 50%50\% probability of correctness.

    Given true label ZjZ_j, labeler expertise αi\alpha_i, and item ease βj\beta_j, the probability that the observed label Lij∈{0,1}L_{ij} \in \{0, 1\} given by labeler ii to item jj matches the true label ZjZ_j is modeled as: p(Lij=Zj∣αi,βj)=11+e−αiβjp(L_{ij} = Z_j \mid \alpha_i, \beta_j) = \frac{1}{1 + e^{-\alpha_i \beta_j}} Assuming conditional independence of labels given Zj,αi,βjZ_j, \alpha_i, \beta_j, the likelihood of a specific observed label lij∈{0,1}l_{ij} \in \{0, 1\} is: p(Lij=lij∣zj,αi,βj)=(11+e−αiβj)I(lij=zj)(1−11+e−αiβj)I(lij≠zj)p(L_{ij} = l_{ij} \mid z_j, \alpha_i, \beta_j) = \left(\frac{1}{1 + e^{-\alpha_i \beta_j}}\right)^{\mathbb{I}(l_{ij} = z_j)} \left(1 - \frac{1}{1 + e^{-\alpha_i \beta_j}}\right)^{\mathbb{I}(l_{ij} \ne z_j)} where I(⋅)\mathbb{I}(\cdot) is the indicator function.

  2. Knowl 2 — Bilinear Log-Odds Formulation of Annotator Correctness

    equation

    In the GLAD model, the log-odds of an observed label Lij∈{0,1}L_{ij} \in \{0, 1\} matching the true binary class label Zj∈{0,1}Z_j \in \{0, 1\} is a bilinear function of the annotator's expertise αi∈(−∞,+∞)\alpha_i \in (-\infty, +\infty) and the item's inverse-difficulty βj∈(0,+∞)\beta_j \in (0, +\infty): log⁡(p(Lij=Zj∣αi,βj)1−p(Lij=Zj∣αi,βj))=αiβj\log \left( \frac{p(L_{ij} = Z_j \mid \alpha_i, \beta_j)}{1 - p(L_{ij} = Z_j \mid \alpha_i, \beta_j)} \right) = \alpha_i \beta_j Under this relationship:

    1. When expertise αi\alpha_i increases (for βj>0\beta_j > 0), the odds of a correct label increase exponentially.
    2. When either αi→0\alpha_i \to 0 or βj→0\beta_j \to 0 (infinite difficulty 1/βj1/\beta_j), the log-odds approach 00, yielding p(Lij=Zj)=0.5p(L_{ij} = Z_j) = 0.5.
    3. When αi<0\alpha_i < 0 (adversarial labeling), the log-odds become negative, making the annotator more likely to provide the incorrect label than the correct one.
  3. Knowl 3 — Expectation-Maximization Inference Algorithm for GLAD

    algorithm

    An Expectation-Maximization (EM) algorithm is used to simultaneously infer the latent true labels Z={Zj}j=1nZ = \{Z_j\}_{j=1}^n and estimate the maximum likelihood or MAP parameters α={αi}i=1m\alpha = \{\alpha_i\}_{i=1}^m and β={βj}j=1n\beta = \{\beta_j\}_{j=1}^n from sparse observed labels l={lij}l = \{l_{ij}\}.

    Input: Observed labels l={lij}l = \{l_{ij}\}, class priors p(Zj)p(Z_j), parameter priors p(αi)p(\alpha_i), p(βj)p(\beta_j)
    Output: Posterior label distributions p(Zj∣l,α^,β^)p(Z_j \mid l, \hat{\alpha}, \hat{\beta}) and parameter estimates α^,β^\hat{\alpha}, \hat{\beta}
    Initialize α\alpha and β\beta
    repeat
        // E-step: Compute posterior distribution of true label for each item
        for each item j∈{1,…,n}j \in \{1, \dots, n\} do
            for each class z∈{0,1}z \in \{0, 1\} do
                p(Zj=z∣lj,α,βj)∝p(Zj=z)∏i∈labelers(j)p(Lij=lij∣Zj=z,αi,βj)p(Z_j = z \mid l_j, \alpha, \beta_j) \propto p(Z_j = z) \prod_{i \in \text{labelers}(j)} p(L_{ij} = l_{ij} \mid Z_j = z, \alpha_i, \beta_j)
            Normalize p(Zj=1∣lj,α,βj)+p(Zj=0∣lj,α,βj)=1p(Z_j = 1 \mid l_j, \alpha, \beta_j) + p(Z_j = 0 \mid l_j, \alpha, \beta_j) = 1
        end for
        // M-step: Maximize the expected joint log-likelihood with log priors
        Define Q(α,β)=∑jEZj[ln⁡p(Zj)]+∑i,jEZj[ln⁡p(Lij=lij∣Zj,αi,βj)]+∑iln⁡p(αi)+∑jln⁡p(βj)Q(\alpha, \beta) = \sum_{j} \mathbb{E}_{Z_j}[\ln p(Z_j)] + \sum_{i, j} \mathbb{E}_{Z_j}[\ln p(L_{ij} = l_{ij} \mid Z_j, \alpha_i, \beta_j)] + \sum_i \ln p(\alpha_i) + \sum_j \ln p(\beta_j)
        Update α,β←arg⁡max⁡α,βQ(α,β)\alpha, \beta \leftarrow \arg\max_{\alpha, \beta} Q(\alpha, \beta) using gradient ascent / conjugate gradient
    until convergence
    return p(Z∣l,α^,β^),α^,β^p(Z \mid l, \hat{\alpha}, \hat{\beta}), \hat{\alpha}, \hat{\beta}

    The E-step requires O(n+∣L∣)O(n + |L|) operations, where ∣L∣|L| is the total count of observed labels. The M-step evaluates QQ and ∇Q\nabla Q in O(n+m+∣L∣)O(n + m + |L|) time per gradient iteration.

  4. Knowl 4 — Priors and Parameterizations for Labeler Expertise and Item Ease

    model/method

    In the Maximum A Posteriori (MAP) formulation of GLAD:

    • Labeler expertise αi\alpha_i: A Gaussian prior αi∼N(μ=1,σ=1)\alpha_i \sim \mathcal{N}(\mu = 1, \sigma = 1) is assigned to penalize negative values of αi\alpha_i, reflecting the prior assumption that most human labelers act cooperatively rather than adversarially.
    • Item ease βj\beta_j: To enforce the constraint βj>0\beta_j > 0 without constrained optimization, βj\beta_j is re-parameterized as βj=eβj′\beta_j = e^{\beta'_j}, and an unconstrained Gaussian prior βj′∼N(μ=1,σ=1)\beta'_j \sim \mathcal{N}(\mu = 1, \sigma = 1) is imposed on βj′\beta'_j.
    • Semi-supervised ground truth clamping: If true labels ZjZ_j are known a priori for a subset of gold-standard validation items, the prior p(Zj)p(Z_j) can be clamped to near 11 for the true class during the E-step, which anchors the latent space and improves parameter estimation of α\alpha and β\beta.
  5. Knowl 5 — Comparison of Label Inference Methods on Heterogeneous Item Difficulties

    data/table

    To evaluate the specific impact of modeling item difficulty, a simulation was performed on 1000 binary items (500 "hard" and 500 "easy") annotated by 50 simulated labelers with a 25:1 ratio of "good" to "bad" labelers. Correctness probabilities were configured as:

    Labeler Type Hard Item Accuracy Easy Item Accuracy
    Good 0.95 1.00
    Bad 0.54 1.00

    Label error rates averaged across 20 simulation trials were:

    Method Error Rate
    GLAD 4.5%
    Dawid Skene (1979) 8.4%
    Majority Vote 11.2%

    GLAD reduces error by more than half compared to Majority Vote and significantly outperforms the Dawid and Skene model, which accounts for labeler error rates via confusion matrices but does not model varying item difficulties.

  6. Knowl 6 — Resilience of GLAD to Random Noise and Adversarial Annotators

    empirical result

    When tested on crowdsourced Duchenne smile classification (160 images) with synthetic perturbations, GLAD demonstrates superior robustness compared to Majority Vote:

    • Random Noise: When adding between 0 and 5000 uniformly random (uninformative) labels, Majority Vote classification accuracy falls steadily from ≈72%\approx 72\% to ≈66%\approx 66\%. GLAD maintains stable accuracy between 77%77\% and 78%78\%, as it infers near-zero expertise (αi≈0\alpha_i \approx 0) for random labelers and effectively discounts their votes.
    • Adversarial Labelers: When adding between 0 and 750 systematically inverted (adversarial) labels, Majority Vote accuracy drops steeply below 45%45\%. In contrast, GLAD automatically infers negative abilities (αi<0\alpha_i < 0), inverts the supplied votes, and leverages the adversarial annotations to improve accuracy up toward nearly 100%100\%.
  7. Knowl 7 — Human Annotation Aggregation on Greeble Gender Classification

    empirical result

    In an experiment using 100 synthetic "Greeble" images (48x48 pixel images categorized as male or female based on horn orientation) labeled by 10 Amazon Mechanical Turk annotators per image, GLAD was compared against Majority Vote across varying numbers of sampled labels M∈[2,8]M \in [2, 8]:

    • For every value of MM, GLAD achieved significantly higher classification accuracy than Majority Vote (p<0.01p < 0.01) and lower variance across 100 experimental trials.
    • For even values of MM, Majority Vote suffered performance drops due to ties and the lack of an optimal tie-breaking mechanism, whereas GLAD avoided tie degradation by weighting votes by estimated labeler expertise αi\alpha_i and item difficulty βj\beta_j.
  8. Knowl 8 — Duchenne Smile Classification from Crowdsourced Annotations

    empirical result

    On a dataset of 160 face images (58 containing Duchenne "enjoyment" smiles and 102 Non-Duchenne "social" smiles as coded by certified Facial Action Coding System experts) labeled by 20 Mechanical Turk annotators yielding 3572 total annotations:

    • GLAD inferred the correct ground-truth label on 78.12%78.12\% of the images by thresholding the posterior class probability at 0.50.5.
    • The Majority Vote heuristic achieved 71.88%71.88\% accuracy.
    • GLAD provided an absolute accuracy gain of 6.24%6.24\% over Majority Vote on raw crowdsourced annotations without requiring known ground-truth labels during inference.
  9. Knowl 9 — Parameter Recovery and Classification Scaling with Annotator Count

    empirical result

    In synthetic experiments with 2000 binary items, annotator abilities drawn from αi∼N(1,1)\alpha_i \sim \mathcal{N}(1, 1), and item ease parameters drawn from βj=exp⁡(N(1,1))\beta_j = \exp(\mathcal{N}(1, 1)), parameter estimation and label inference were evaluated across 40 trials as the number of labelers varied from 4 to 20:

    • Parameter Recovery: As the number of labelers increased, parameter estimates converged toward true values, with Pearson correlation for α\alpha exceeding 0.950.95 and Spearman rank correlation for β\beta reaching ≈0.85\approx 0.85 at 20 labelers.
    • Classification Accuracy: GLAD's inferred label accuracy improved from ≈89%\approx 89\% with 4 labelers to ≈99%\approx 99\% with 20 labelers, consistently outperforming Majority Vote across all labeler counts, with the largest performance margins observed at low labeler counts (44 to 88 labelers).
  10. Knowl 10 — Robustness of EM Optimization to Random Initialization in GLAD

    empirical result

    The EM parameter estimation procedure for GLAD exhibits minimal sensitivity to initial parameter starting points. In a simulation involving 2000 images and 20 labelers with starting parameters randomly drawn from uniform distributions: αi∼U[0,4],ln⁡(βj)∼U[0,3]\alpha_i \sim \mathcal{U}[0, 4], \quad \ln(\beta_j) \sim \mathcal{U}[0, 3] Across 50 independent runs, the inferred label accuracy had a mean of 85.74%85.74\% and a standard deviation across trials of only 0.024%0.024\%, demonstrating consistent convergence across arbitrary positive starting configurations.

Coverage note — None omitted; all core contributions including the probabilistic model formulation, inference algorithm, prior parameterizations, synthetic benchmarks, real-world vision experiments, and noise/adversarial analyses are covered.

References

  1. 1.Amazon. Mechanical turk. http://www.mturk.com.
  2. 2.W. H. Batchelder and A. K. Romney. Test theory without an answer key. Psychometrika, 53(1):71–92, 1988.
  3. 3.A. Birnbaum. Some latent trait models and their use in inferring an examinee’s ability. Statistical theories of mental test scores, 1968.
  4. 4.N. Butko and J. Movellan. I-POMDP: An infomax model of eye movement. In Proceedings of the International Conference on Development and Learning, 2008.
  5. 5.A. Dawid and A. Skene. Maximum likelihood estimation of observer error-rates using the em algorithm. Applied Statistics, 28(1):20–28, 1979.
  6. 6.I. Gauthier and M. Tarr. Becoming a “greeble” expert: Exploring mechanisms for face recognition. Vision Research, 37(12), 1997.
  7. 7.V. Johnson. On bayesian analysis of multi-rater ordinal data: An application to automated essay grading. Journal of the American Statistical Association, 91:42–51, 1996.
  8. 8.G. Karabatsos and W. H. Batchelder. Markov chain estimation for test theory without an answer key. Psychometrika, 68(3):373–389, 2003.
  9. 9.Omron. OKAO vision brochure, July 2008.
  10. 10.G. Rasch. Probabilistic Models for Some Intelligence and Attainment Tests. Denmark, 1960.
  11. 11.S. Rogers, M. Girolami, and T. Polajnar. Semi-parametric analysis of multi-rater data. Statistics and Computing, 2009.
  12. 12.V. Sheng, F. Provost, and P. Ipeirotis. Get another label? improving data quality and data mining using multiple noisy labelers. In Knowledge Discovery and Data Mining, 2008.
  13. 13.P. Smyth, U. Fayyad, M. Burl, P. Perona, and P. Baldi. Inferring ground truth from subjective labelling of venus images. In Advances of Neural Information Processing Systems, 1994.
  14. 14.R. Snow, B. O’Connor, D. Jurafsky, and A. Y. Ng. Cheap and fast - but is it good? evaluating non-expert annotations for natural language tasks. In Proceedings of the 2008 Conference on Empirical Methods on Natural Language Processing, 2008.
  15. 15.S. Steinbach, V. Rabaud, and S. Belongie. Soylent grid: it’s made of people! In International Conference on Computer Vision, 2007.
  16. 16.D. Turnbull, R. Liu, L. Barrington, and G. Lanckriet. A Game-based Approach for Collecting Semantic Annotations of Music. In 8th International Conference on Music Information Retrieval (ISMIR), 2007.
  17. 17.L. von Ahn and L. Dabbish. Labeling Images with A Computer Game. In Proceedings of the SIGCHI conference on Human factors in computing systems, pages 319–326. ACM Press New York, NY, USA, 2004.
  18. 18.L. von Ahn, B. Maurer, C. McMillen, D. Abraham, and M. Blum. reCAPTCHA: Human-Based Character Recognition via Web Security Measures. Science, 321(5895):1465, 2008.

Citation

MLA
Whitehill, J., et al. “Whose Vote Should Count More: Optimal Integration of Labels from Labelers of Unknown Expertise”. Advances in Neural Information Processing Systems, vol. 22, 2009, https://proceedings.neurips.cc/paper_files/paper/2009/file/f899139df5e1059396431415e770c6dd-Paper.pdf.
APA
Whitehill, J., Wu, T.-. fan ., Bergsma, J., Movellan, J., & Ruvolo, P. (2009). Whose Vote Should Count More: Optimal Integration of Labels from Labelers of Unknown Expertise. Advances in Neural Information Processing Systems, 22. https://proceedings.neurips.cc/paper_files/paper/2009/file/f899139df5e1059396431415e770c6dd-Paper.pdf
Chicago
Whitehill, J., T.-. fan . Wu, J. Bergsma, J. Movellan, and P. Ruvolo. 2009. “Whose Vote Should Count More: Optimal Integration of Labels from Labelers of Unknown Expertise”. Advances in Neural Information Processing Systems 22. https://proceedings.neurips.cc/paper_files/paper/2009/file/f899139df5e1059396431415e770c6dd-Paper.pdf.
Harvard
Whitehill, J. et al. (2009) “Whose Vote Should Count More: Optimal Integration of Labels from Labelers of Unknown Expertise”, Advances in Neural Information Processing Systems. Curran Associates, Inc. Available at: https://proceedings.neurips.cc/paper_files/paper/2009/file/f899139df5e1059396431415e770c6dd-Paper.pdf.
Vancouver
1. Whitehill J, Wu T-fan, Bergsma J, Movellan J, Ruvolo P (2009) Whose Vote Should Count More: Optimal Integration of Labels from Labelers of Unknown Expertise. Advances in Neural Information Processing Systems 22:

BibTeX

@inproceedings{whitehill2009whose,
  title = {Whose Vote Should Count More: Optimal Integration of Labels from Labelers of Unknown Expertise},
  author = {Whitehill, Jacob and Wu, Ting-fan and Bergsma, Jacob and Movellan, Javier and Ruvolo, Paul},
  year = {2009},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {22},
  url = {https://proceedings.neurips.cc/paper_files/paper/2009/file/f899139df5e1059396431415e770c6dd-Paper.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors