Cheap and Fast – But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks

Rion SnowBrendan O'ConnorDaniel JurafskyAndrew Y. Ng

article2008EMNLP2,534 citations

Demonstrates that carefully designed non-expert crowdsourcing can produce reliable NLP labels at a fraction of expert annotation cost.

Listen

Natural language processing research depends on large annotated datasets, but creating them with expert labelers is costly and slow. The article examines whether Amazon Mechanical Turk can supply reliable annotations from non-experts at far lower expense.

The study set out to test non-expert performance on five established tasks: affect recognition in headlines, word similarity judgments, recognizing textual entailment, temporal ordering of events, and word sense disambiguation. Researchers collected ten independent labels per item via Mechanical Turk and compared the results directly to existing expert gold standards using agreement metrics such as Pearson correlation and majority-vote accuracy.

For every task, non-expert labels reached high agreement with experts once multiple annotations were averaged. Four non-expert labels per item typically matched the reliability of one expert; ten labels produced even stronger results. A bias-correction model that accounts for individual worker accuracy improved performance by roughly four percentage points on two categorical tasks. Classifiers trained on non-expert affect labels performed as well as or better than those trained on single-expert labels.

These outcomes show that crowdsourced annotations can replace or supplement expert work for many labeling projects, cutting costs to hundreds of labels per dollar while maintaining usable quality. The approach also surfaced and corrected an error in one expert gold set.

The findings support wider adoption of Mechanical Turk for dataset creation, provided projects collect redundant labels and apply simple quality controls. Additional testing on new tasks, larger scales, and more complex annotation schemes would strengthen before full-scale deployment.

Cover for Cheap and Fast – But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks

Abstract

Human linguistic annotation is crucial for many natural language processing tasks but can be expensive and time-consuming. We explore the use of Amazon’s Mechanical Turk system, a significantly cheaper and faster method for collecting annotations from a broad base of paid non-expert contributors over the Web. We investigate five tasks: affect recognition, word similarity, recognizing textual entailment, event temporal ordering, and word sense disambiguation. For all five, we show high agreement between Mechanical Turk non-expert annotations and existing gold standard labels provided by expert labelers. For the task of affect recognition, we also show that using non-expert labels for training machine learning algorithms can be as effective as using gold standard annotations from experts. We propose a technique for bias correction that significantly improves annotation quality on two tasks. We conclude that many large labeling tasks can be effectively designed and carried out in this method at a fraction of the usual expense.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Task Design
  • 3.1 Amazon Mechanical Turk
  • 3.2 Task Design
  • 4 Annotation Tasks
  • 4.1 Affective Text Analysis
  • 4.2 Word Similarity
  • 4.3 Recognizing Textual Entailment
  • 4.4 Event Annotation
  • 4.5 Word Sense Disambiguation
  • 4.6 Summary
  • 5 Bias correction for non-expert annotators
  • 5.1 Bias correction in categorical data
  • 5.1.1 Example tasks: RTE-1 and event annotation
  • 6 Training a system with non-expert annotations
  • 6.1 Experimental Design
  • 6.2 Experiments
  • 7 Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Bayesian Bias Correction for Categorical Crowdsourced Annotations

    model/method

    A probabilistic model is used to calibrate individual non-expert worker biases using a small pool of gold standard training labels. For an item ii with a true categorical label xi{Y,N}x_i \in \{Y, N\} and observed worker responses yi1,yi2,,yiWy_{i1}, y_{i2}, \dots, y_{iW}, the workers' judgments are modeled as conditionally independent given xix_i:

    P(yi1,,yiW,xi)=(wP(yiwxi))p(xi)P(y_{i1}, \dots, y_{iW}, x_i) = \left(\prod_{w} P(y_{iw} \mid x_i)\right) p(x_i)

    The worker-specific response likelihoods P(yiwxi)P(y_{iw} \mid x_i) are estimated from worker performance on gold-standard items using Laplace smoothing with 1 pseudocount.

    To infer the true label for a new item, the posterior log-odds is computed via Bayes' rule:

    logP(xi=Yyi1,,yiW)P(xi=Nyi1,,yiW)=wlogP(yiwxi=Y)P(yiwxi=N)+logP(xi=Y)P(xi=N)\log \frac{P(x_i = Y \mid y_{i1}, \dots, y_{iW})}{P(x_i = N \mid y_{i1}, \dots, y_{iW})} = \sum_{w} \log \frac{P(y_{iw} \mid x_i = Y)}{P(y_{iw} \mid x_i = N)} + \log \frac{P(x_i = Y)}{P(x_i = N)}

    Under Maximum A Posteriori (MAP) estimation, this formulation operates as a weighted voting rule where each worker's vote is weighted by their log-likelihood ratio for that response: workers with greater than 50% accuracy receive positive vote weights, uninformative workers receive zero weight, and negatively correlated workers receive negative weights.

  2. Knowl 2 — Affect Classifier Performance Trained on Crowdsourced vs Expert Annotations

    empirical result

    A supervised bag-of-words unigram affect recognition classifier was trained on 100 headline annotations and tested on 900 headlines from SemEval-2007 Task 14 across six emotional dimensions (anger, disgust, fear, joy, sadness, surprise) and valence. Models trained on annotations from individual expert annotators were compared against models trained on subsets of non-expert crowdsourced annotators. System quality is evaluated by Pearson correlation against the gold standard formed by averaging the remaining five expert annotators.

    Emotion 1-Expert 10-NE kk kk-NE
    Anger 0.084 0.233 1 0.172
    Disgust 0.130 0.231 1 0.185
    Fear 0.159 0.247 1 0.176
    Joy 0.130 0.125
    Sadness 0.127 0.174 1 0.141
    Surprise 0.060 0.101 1 0.061
    Valence 0.159 0.229 2 0.146
    Avg. Emo 0.116 0.185 1 0.135
    Avg. All 0.122 0.191 1 0.137

    In the table, "1-Expert" is the average Pearson correlation of classifiers trained on a single expert's labels; "10-NE" is the correlation when trained on the average of 10 non-expert annotations per headline; "kk" is the minimum number of non-expert labels per item required for the classifier to outperform the 1-Expert model; and "kk-NE" is the classifier correlation achieved with kk non-experts.

    Across all categories, a classifier trained on the annotations from an average of 10 non-experts achieves a correlation of 0.191 compared to 0.122 for a single expert. Furthermore, for 5 of the 7 individual tasks, training on a single set of non-expert annotations (k=1k=1, correlation 0.137 overall) outperforms training on a single expert annotator (0.122). This occurs because crowdsourced datasets aggregate judgments from multiple diverse individuals, which reduces idiosyncratic annotator bias relative to any single human annotator.

  3. Knowl 3 — Non-Expert Crowd Inter-Annotator Agreement on Affective Text Analysis

    empirical result

    Non-expert crowdsourced annotators on Amazon Mechanical Turk were evaluated against a 6-expert panel on a 100-headline subset of SemEval-2007 Task 14 (Affective Text). Annotators assigned numeric ratings in [0,100][0, 100] for six emotions (anger, disgust, fear, joy, sadness, surprise) and in [100,100][-100, 100] for emotional valence. Inter-annotator agreement (ITA) was measured via Pearson correlation.

    Emotion E vs. E E vs. All NE vs. E NE vs. All
    Anger 0.459 0.503 0.444 0.573
    Disgust 0.583 0.594 0.537 0.647
    Fear 0.711 0.683 0.418 0.498
    Joy 0.596 0.585 0.340 0.421
    Sadness 0.645 0.650 0.563 0.651
    Surprise 0.464 0.463 0.201 0.225
    Valence 0.759 0.767 0.530 0.554
    Avg. Emo 0.576 0.603 0.417 0.503
    Avg. All 0.580 0.607 0.433 0.510
    Emotion 1-Expert 10-NE kk kk-NE
    Anger 0.459 0.675 2 0.536
    Disgust 0.583 0.746 2 0.627
    Fear 0.711 0.689
    Joy 0.596 0.632 7 0.600
    Sadness 0.645 0.776 2 0.656
    Surprise 0.464 0.496 9 0.481
    Valence 0.759 0.844 5 0.803
    Avg. Emo. 0.576 0.669 4 0.589
    Avg. All 0.603 0.694 4 0.613

    Individual experts exhibit higher agreement with other experts (mean ITA 0.580) than individual non-experts do with experts (mean ITA 0.433). However, averaging subsets of nn non-expert responses forms a "meta-labeler" whose agreement increases monotonically with nn. On average across all affect dimensions, pooling k=4k=4 non-expert annotations per item achieves an agreement correlation of 0.613, surpassing the average single-expert agreement of 0.603. Averaging 10 non-experts achieves an ITA of 0.694.

  4. Knowl 4 — Cost and Throughput of Non-Expert Annotations on Amazon Mechanical Turk

    data/table

    Annotation experiments were conducted across five natural language processing tasks using Amazon Mechanical Turk with ten independent annotators per item. The table summarizes the scale, financial cost, completion time, and throughput rates for each task:

    Task Labels Cost (USD) Time (hrs) Labels per USD Labels per hr
    Affect 7000 $2.00 5.93 3500.0 1180.4
    WSim 300 $0.20 0.174 1500.0 1724.1
    RTE 8000 $8.00 89.3 1000.0 89.59
    Event 4620 $13.86 39.9 333.3 115.85
    WSD 1770 $1.76 8.59 1005.7 206.1
    Total 21690 $25.82 143.9 840.0 150.7

    Across all tasks (Affective Text Analysis, Word Similarity [WSim], Recognizing Textual Entailment [RTE], Event Temporal Ordering [Event], and Word Sense Disambiguation [WSD]), a total of 21,690 labels were obtained for $25.82 over 143.9 total wall-clock hours, averaging 840.0 labels per USD and 150.7 labels per hour. Tasks with shorter response requirements like word similarity achieved throughput as high as 1724.1 labels per hour and 1500 labels per USD.

  5. Knowl 5 — Unigram Bag-of-Words Model for Headline Affect Prediction

    model/method

    A unigram bag-of-words scoring model for headline affect prediction is defined without lemmatization or preprocessing by assigning emotion weights to vocabulary tokens from training headlines and scoring new headlines via token averaging.

    For each token tt in a training corpus, its weight for emotion dimension ee is computed as the mean score across all training headlines containing tt:

    Score(e,t)=HHtScore(e,H)Ht\text{Score}(e, t) = \frac{\sum_{H \in H_t} \text{Score}(e, H)}{|H_t|}

    where HtH_t is the set of headlines containing token tt, Score(e,H)\text{Score}(e, H) is the annotated numerical score of headline HH for emotion ee, and Ht|H_t| is the number of headlines containing tt.

    To score a novel headline HH for emotion ee, the predicted score is the mean weight of its constituent tokens that appeared in the training set:

    Score(e,H)=tHScore(e,t)H\text{Score}(e, H) = \frac{\sum_{t \in H} \text{Score}(e, t)}{|H|}

    where H|H| is the count of tokens in headline HH observed during training (tokens unobserved in the training set are omitted).

  6. Knowl 6 — Performance Gains from Gold-Calibrated Bias Correction on Non-Expert Annotations

    empirical result

    Using a Naive Bayes-style calibration model trained on a small gold standard subset via 20-fold cross-validation with Laplace smoothing of 1 pseudocount, non-expert worker responses were corrected for individual bias and noise on categorical and numeric tasks.

    On the PASCAL RTE-1 recognizing textual entailment binary classification task, gold-calibrated worker weighting improved accuracy by an average of +4.0%+4.0\% over naive 50% majority voting across crowds consisting of 2 to 10 annotators per item.

    On the TimeBank event temporal ordering binary classification task ("strictly before" vs. "strictly after"), gold-calibrated weighting produced an average accuracy gain of +3.4%+3.4\% over naive majority voting across 2 to 10 annotators per item.

    On the continuous SemEval affect recognition task, calibrating worker responses using a Gaussian noise model ywxN(x+μw,σw)y_w \mid x \sim \mathcal{N}(x + \mu_w, \sigma_w) yielded an average +0.6%+0.6\% gain in Pearson correlation across all numbers of annotators.

  7. Knowl 7 — Agreement of Crowdsourced Annotations on Word Similarity, RTE, Temporal Ordering, and WSD

    empirical result

    Non-expert crowdsourced annotators on Amazon Mechanical Turk were evaluated against expert gold standards across four core NLP tasks:

    1. Word Similarity: On 30 word pairs from the Miller and Charles (1991) dataset scored on a scale of [0,10][0, 10], averaging judgments from 10 non-expert annotators yielded a Pearson correlation of r=0.952r = 0.952 with gold scores, closely matching the expert panel correlation of r=0.958r = 0.958 reported by Resnik (1999).

    2. Recognizing Textual Entailment (RTE): On 800 sentence pairs from PASCAL RTE-1 with binary inference choices, majority voting over 10 non-experts achieved 89.7% accuracy against gold labels (compared to expert agreement ranges of 91% to 96%).

    3. Event Temporal Ordering: On 462 verb event pairs from TimeBank with binary labels ("strictly before" or "strictly after"), simple majority voting across 10 non-experts achieved 94.0% accuracy.

    4. Word Sense Disambiguation (WSD): On 177 instances of the noun "president" with three sense classes from SemEval-2007 Task 17, majority voting over 10 non-experts reached 99.4% agreement with the original gold standard, and 100% agreement after resolving a known error in the gold annotation.

  8. Knowl 8 — Identification of Expert Gold Standard Errors via Crowdsourced Consensus

    empirical result

    Crowdsourced non-expert consensus can detect and correct errors in established expert-annotated gold standard datasets. In an evaluation of 177 examples for the noun "president" from SemEval-2007 Task 17 Word Sense Disambiguation, non-expert annotators using majority voting achieved an initial accuracy of 99.4% against the gold standard with only a single disagreement.

    In that single instance (the sentence beginning "The Egyptian president said he would visit Libya today..."), the expert gold standard had marked the sense as "executive officer of a firm, corporation, or university". Non-expert annotators voted 9-to-1 in favor of "head of a country (other than the U.S.)". Inspection confirmed that the original expert gold label was erroneous, and correcting it brought non-expert consensus accuracy on the dataset to 100%.

Coverage note — No substantial contributed material was omitted; the extracted knowls comprehensively cover the five crowdsourced tasks, expert agreement comparisons, bias calibration formulations and evaluation gains, downstream classifier experiments, and cost-throughput analyses.

References

  1. 1.Paul S. Albert and Lori E. Dodd. 2004. A Cautionary Note on the Robustness of Latent Class Models for Estimating Diagnostic Error without a Gold Standard. Biometrics, Vol. 60 (2004), pp. 427-435.
  2. 2.Collin F. Baker, Charles J. Fillmore and John B. Lowe. 1998. The Berkeley FrameNet project. In Proc. of COLING-ACL 1998.
  3. 3.Michele Banko and Eric Brill. 2001. Scaling to Very Very Large Corpora for Natural Language Disambiguation. In Proc. of ACL-2001.
  4. 4.Junfu Cai, Wee Sun Lee and Yee Whye Teh. 2007. Improving Word Sense Disambiguation Using Topic Features. In Proc. of EMNLP-2007.
  5. 5.Timothy Chklovski and Rada Mihalcea. 2002. Building a sense tagged corpus with Open Mind Word Expert. In Proc. of the Workshop on "Word Sense Disambiguation: Recent Successes and Future Directions", ACL 2002.
  6. 6.Timothy Chklovski and Yolanda Gil. 2005. Towards Managing Knowledge Collection from Volunteer Contributors. Proceedings of AAAI Spring Symposium on Knowledge Collection from Volunteer Contributors (KCVC05).
  7. 7.Ido Dagan, Oren Glickman and Bernardo Magnini. 2006. The PASCAL Recognising Textual Entailment Challenge. Machine Learning Challenges. Lecture Notes in Computer Science, Vol. 3944, pp. 177-190, Springer, 2006.
  8. 8.Wisam Dakka and Panagiotis G. Ipeirotis. 2008. Automatic Extraction of Useful Facet Terms from Text Documents. In Proc. of ICDE-2008.
  9. 9.A. P. Dawid and A. M. Skene. 1979. Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm. Applied Statistics, Vol. 28, No. 1 (1979), pp. 20-28.
  10. 10.Michael Kaisser and John B. Lowe. 2008. A Research Collection of QuestionAnswer Sentence Pairs. In Proc. of LREC-2008.
  11. 11.Michael Kaisser, Marti Hearst, and John B. Lowe. 2008. Evidence for Varying Search Results Summary Lengths. In Proc. of ACL-2008.
  12. 12.Phil Katz, Matthew Singleton, Richard Wicentowski. 2007. SWAT-MP: The SemEval-2007 Systems for Task 5 and Task 14. In Proc. of SemEval-2007.
  13. 13.Aniket Kittur, Ed H. Chi, and Bongwon Suh. 2008. Crowdsourcing user studies with Mechanical Turk. In Proc. of CHI-2008.
  14. 14.Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini. 1993. Building a large annotated corpus of English: the Penn Treebank. Computational Linguistics 19:2, June 1993.
  15. 15.George A. Miller and William G. Charles. 1991. Contextual Correlates of Semantic Similarity. Language and Cognitive Processes, vol. 6, no. 1, pp. 1-28, 1991.
  16. 16.George A. Miller, Claudia Leacock, Randee Tengi, and Ross T. Bunke. 1993. A semantic concordance. In Proc. of HLT-1993.
  17. 17.Preslav Nakov. 2008. Paraphrasing Verbs for Noun Compound Interpretation. In Proc. of the Workshop on Multiword Expressions, LREC-2008.
  18. 18.Martha Palmer, Dan Gildea, and Paul Kingsbury. 2005. The Proposition Bank: A Corpus Annotated with Semantic Roles. Computational Linguistics, 31:1.
  19. 19.Sameer Pradhan, Edward Loper, Dmitriy Dligach and Martha Palmer. 2007. SemEval-2007 Task-17: English Lexical Sample, SRL and All Words. In Proc. of SemEval-2007.
  20. 20.James Pustejovsky, Patrick Hanks, Roser Saur, Andrew See, Robert Gaizauskas, Andrea Setzer, Dragomir Radev, Beth Sundheim, David Day, Lisa Ferro and Marcia Lazo. 2003. The TIMEBANK Corpus. In Proc. of Corpus Linguistics 2003, 647-656.
  21. 21.Philip Resnik. 1999. Semantic Similarity in a Taxonomy: An Information-Based Measure and its Application to Problems of Ambiguity in Natural Language. JAIR, Volume 11, pages 95-130.
  22. 22.Herbert Rubenstein and John B. Goodenough. 1965. Contextual Correlates of Synonymy. Communications of the ACM, 8(10):627–633.
  23. 23.Victor S. Sheng, Foster Provost, and Panagiotis G. Ipeirotis. 2008. Get Another Label? Improving Data Quality and Data Mining Using Multiple, Noisy Labelers. In Proc. of KDD-2008.
  24. 24.Push Singh. 2002. The public acquisition of commonsense knowledge. In Proc. of AAAI Spring Symposium on Acquiring (and Using) Linguistic (and World) Knowledge for Information Access, 2002.
  25. 25.Alexander Sorokin and David Forsyth. 2008. Utility data annotation with Amazon Mechanical Turk. To appear in Proc. of First IEEE Workshop on Internet Vision at CVPR, 2008. See also: http://vision.cs.uiuc.edu/annotation/
  26. 26.David G. Stork. 1999. The Open Mind Initiative. IEEE Expert Systems and Their Applications pp. 16-20, May/June 1999.
  27. 27.Carlo Strapparava and Rada Mihalcea. 2007. SemEval-2007 Task 14: Affective Text In Proc. of SemEval-2007.
  28. 28.Qi Su, Dmitry Pavlov, Jyh-Herng Chow, and Wendell C. Baker. 2007. Internet-Scale Collection of Human-Reviewed Data. In Proc. of WWW-2007.
  29. 29.Luis von Ahn and Laura Dabbish. 2004. Labeling Images with a Computer Game. In ACM Conference on Human Factors in Computing Systems, CHI 2004.
  30. 30.Luis von Ahn, Mihir Kedia and Manuel Blum. 2006. Verbosity: A Game for Collecting Common-Sense Knowledge. In ACM Conference on Human Factors in Computing Systems, CHI Notes 2006.
  31. 31.Ellen Voorhees and Hoa Trang Dang. 2006. Overview of the TREC 2005 question answering track. In Proc. of TREC-2005.
  32. 32.Janyce M. Wiebe, Rebecca F. Bruce and Thomas P. O’Hara. 1999. Development and use of a gold-standard data set for subjectivity classifications. In Proc. of ACL-1999.
  33. 33.Annie Zaenen. Submitted. Do give a penny for their thoughts. International Journal of Natural Language Engineering (submitted).

Citation

MLA
Snow, R., et al. “Cheap and Fast---but Is It Good?”. Proceedings of the Conference on Empirical Methods in Natural Language Processing - EMNLP '08, 2008, p. 254, https://doi.org/10.3115/1613715.1613751.
APA
Snow, R., O'Connor, B., Jurafsky, D., & Ng, A. Y. (2008). Cheap and fast---but is it good?. Proceedings of the Conference on Empirical Methods in Natural Language Processing - EMNLP '08, 254. https://doi.org/10.3115/1613715.1613751
Chicago
Snow, R., B. O'Connor, D. Jurafsky, and A. Y. Ng. 2008. “Cheap and Fast---but Is It Good?”. Proceedings of the Conference on Empirical Methods in Natural Language Processing - EMNLP '08, 254. https://doi.org/10.3115/1613715.1613751.
Harvard
Snow, R. et al. (2008) “Cheap and fast---but is it good?”, Proceedings of the Conference on Empirical Methods in Natural Language Processing - EMNLP '08. Association for Computational Linguistics, p. 254. Available at: https://doi.org/10.3115/1613715.1613751.
Vancouver
1. Snow R, O'Connor B, Jurafsky D, Ng AY (2008) Cheap and fast---but is it good?. In: Proceedings of the Conference on Empirical Methods in Natural Language Processing - EMNLP '08. Association for Computational Linguistics, p 254

BibTeX

@inproceedings{Snow_2008, series={EMNLP ’08}, title={Cheap and fast---but is it good?: evaluating non-expert annotations for natural language tasks}, url={http://dx.doi.org/10.3115/1613715.1613751}, DOI={10.3115/1613715.1613751}, booktitle={Proceedings of the Conference on Empirical Methods in Natural Language Processing - EMNLP ’08}, publisher={Association for Computational Linguistics}, author={Snow, Rion and O’Connor, Brendan and Jurafsky, Daniel and Ng, Andrew Y.}, year={2008}, pages={254}, collection={EMNLP ’08} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by-nc-sa/4.0/