Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI

Nick PangakisSam Wolken

article2025International Conference on Web and Social Media50 citations

Demonstrates through twenty-seven private social science tasks that large language model annotations vary unpredictably and diverge from human judgment, proving that human validation remains essential for automated research workflows.

Listen

Natural language processing tools, particularly generative large language models, are increasingly deployed to automate the manual annotation of text data in social science and policy research. While automated annotation offers substantial cost and speed advantages over human coders, previous evaluations often relied on public benchmark datasets susceptible to data leakage and memorization. Consequently, it remained uncertain whether artificial intelligence could reliably replace human judgment on novel, real-world annotation tasks without introducing severe biases into downstream analyses.

The article evaluates how accurately generative large language models approximate human annotations on non-public datasets and tests a human-centered evaluation framework for responsible automated annotation. To avoid pre-training contamination, the authors replicated 27 manual annotation tasks across 11 password-protected datasets from computational social science studies published in high-impact journals within the past three years. Using OpenAI's GPT-4, the analysis evaluated more than 75,000 text samples against human ground-truth labels and benchmarked performance against over 500 fine-tuned supervised baseline models.

The findings reveal that while GPT-4 achieves solid overall results—with a median accuracy of 85.0% and a median F1 score of 0.707—performance is highly inconsistent across tasks. For one-third of the evaluated tasks, the model produced either precision or recall below 0.50, meaning that false positives outnumbered true positives or more than half of the true instances were missed. The model exhibited higher recall than precision in 20 of the 27 tasks, with a median recall of 0.83 compared to a median precision of 0.65. Furthermore, common interventions such as prompt refinement and temperature tuning yielded only marginal gains, and traditional supervised classifiers outperformed GPT-4 when sufficient training data (1,000 samples) was available.

These results demonstrate that generative models cannot be trusted to operate as unvalidated replacements for human annotators, especially on nuanced tasks requiring cultural context or conversational inference. Deploying generative tools without validation introduces substantial risks of misclassification and biased insights. However, the model's high recall makes it highly effective as an initial filtering mechanism within multi-stage annotation workflows, where automated systems identify potential positive cases for subsequent human verification.

Organizations should adopt human-centered evaluation workflows, validating automated classifications against task-specific, human-labeled samples before deploying models into production pipelines. Practitioners can also exploit model uncertainty by assessing label consistency across repeated queries, which successfully identifies ambiguous edge cases that require manual review. Where significant quantities of human-annotated data already exist, teams should consider deploying dedicated supervised classifiers rather than few-shot generative models. Confidence in these findings is high due to the evaluation of non-contaminated datasets, though researchers should note that baseline labels rely on the quality of original human coder agreements.

  • Paper: ChatGPT outperforms crowd workers for text-annotation tasks, Fabrizio Gilardi et al. (2023). This foundational study demonstrates that zero-shot LLMs can outperform crowd workers on text annotation tasks, establishing the baseline paradigm of automated LLM annotation that the source paper critically evaluates and re-examines across protected datasets.
  • Paper: Is GPT-3 a Good Data Annotator?, Bosheng Ding et al. (2023). It systematically explores using GPT models as automated data annotators for supervised learning, providing the technical and empirical background for evaluating LLM label generation against human ground truth.
  • Paper: Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions, John Joon Young Chung et al. (2023). It provides necessary context on integrating human-in-the-loop verification strategies to maintain data accuracy when using LLMs for dataset construction and annotation.
  • Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). It introduces foundational methodologies and validation protocols for evaluating GPT-4's agreement with human judgment across complex evaluative tasks.
  • Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). This survey provides a comprehensive synthesis of LLM judgment capabilities, systematic biases, and alignment with human annotators, establishing key concepts scrutinized in the source.
  • Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). It investigates how prompt design and demonstrations impact in-context classification performance, informing the prompt optimization techniques analyzed in the source.
  • Paper: GPT-4 Technical Report, OpenAI (2023). It documents the architecture, capabilities, and baseline evaluation performance of GPT-4, the primary model evaluated for text annotation in the source study.
Cover for Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI

Abstract

Automated text annotation is a compelling use case for generative large language models (LLMs) in social media research. Recent work suggests that LLMs can achieve strong performance on annotation tasks; however, these studies evaluate LLMs on a small number of tasks and likely suffer from contamination due to a reliance on public benchmark datasets. Here, we test a human-centered framework for responsibly evaluating artificial intelligence tools used in automated annotation. We use GPT-4 to replicate 27 annotation tasks across 11 password-protected datasets from recently published computational social science articles in high-impact journals. For each task, we compare GPT-4 annotations against human-annotated ground-truth labels and against annotations from separate supervised classification models fine-tuned on human-generated labels. Although the quality of LLM labels is generally high, we find significant variation in LLM performance across tasks, even within datasets. Our findings underscore the importance of a human-centered workflow and careful evaluation standards: Automated annotations significantly diverge from human judgment in numerous scenarios, despite various optimization strategies such as prompt tuning. Grounding automated annotation in validation labels generated by humans is essential for responsible evaluation.

Table of Contents

  • Introduction
  • Data and Replication Procedures
  • Data
  • Replication Procedures
  • Results
  • Manual Prompt Optimization
  • Exploiting Uncertainty in LLM Labels
  • Manual Temperature Optimization
  • GPT-4 Performance Over Time
  • Comparing GPT-4 to supervised model baseline
  • Discussion
  • References
  • Limitations
  • Ethics Statement
  • A Appendix A: Replication data and tasks
  • B Appendix B: Human labels in replication data
  • C Appendix C: Methodological details
  • D Appendix D: Comparing to an open-source alternative
  • E Appendix E: Relationship between task difficulty and performance
  • F Appendix F: Miscellaneous additional information
  • G Checklist

Knowls

  1. Knowl 1 — Human-in-the-Loop LLM Annotation Workflow

    model/method

    A human-centered framework for automating text classification using large language models (LLMs) integrates human validation to ensure label validity across four sequential steps:

    1. Instruction Formulation: Construct a qualitative codebook containing explicit concept definitions and categorization rules that serves both as instructions for human annotators and as the system prompt for the LLM.
    2. Validation on Human-Labeled Subsample: Prompt the LLM to classify a randomly selected subset of text samples (n=250n = 250) already annotated by humans. Compute baseline performance metrics (accuracy, precision, recall, and F1 score) against the human ground-truth labels.
    3. Prompt Optimization: Inspect systematic error patterns and false predictions on the validation subset. Refine the prompt instructions by clarifying ambiguous edge cases and providing few-shot examples of prior misclassifications to align model decisions with human criteria.
    4. Held-Out Evaluation: Prompt the LLM with the updated instructions to label a separate held-out sample of human-annotated text (n=1,000n = 1{,}000) to measure out-of-sample classification performance prior to deploying the model on unannotated corpora.
  2. Knowl 2 — Contamination-Resistant Computational Social Science Annotation Benchmark

    experimental setup

    To evaluate large language model (LLM) text annotation while minimizing the risk of pretraining data contamination and memorization, 27 annotation tasks were compiled from 11 password-protected replication datasets of computational social science (CSS) papers published between 2019 and 2024 across fields including political science and psychology.

    Multiclass classification problems were decomposed into individual binary classification tasks to standardize the evaluation workflow and permit fine-grained error analysis. The benchmark properties include:

    • Annotation Sources: 40.8% of tasks were originally labeled by undergraduate research assistants, 37.0% by domain experts, and 22.2% by crowdsourced workers.
    • Class Imbalance: Positive class prevalence ranges from a minimum of 0.04% to a maximum of 70.7%, with a median of 24.5%.
    • Data Split: For each task, evaluation is performed on 1,000 held-out text instances, with a separate set of 250 instances allocated for validation and prompt refinement.
  3. Knowl 3 — Aggregate Performance Distribution of GPT-4 Across CSS Annotation Tasks

    data/table

    Evaluation of GPT-4 across 27 binary text annotation tasks from 11 computational social science datasets on held-out test sets of 1,000 human-annotated samples per task demonstrates significant performance heterogeneity:

    Metric Minimum Mean Median Maximum
    F1 Score 0.06 0.66 0.71 0.97
    Accuracy 0.67 0.86 0.85 0.98
    Precision 0.03 0.62 0.65 0.96
    Recall 0.25 0.75 0.83 0.98

    Although GPT-4 achieves high median accuracy (0.850) and median F1 (0.707), performance drops severely on a substantial subset: 9 of the 27 tasks (33.3%) exhibit precision or recall below 0.50, and 3 tasks fall below 0.50 on both metrics simultaneously. Performance also varies widely within individual datasets; for instance, on political speech data from Card et al. (2022), F1 scores across individual annotation subtasks range from 0.259 to 0.811 (a difference of 0.552).

  4. Knowl 4 — Recall-Skewed Precision-Recall Trade-Off in LLM Annotation

    empirical result

    In few-shot text annotation using GPT-4 across 27 social science tasks, model performance exhibits a pronounced asymmetry favoring recall over precision:

    • In 20 of the 27 tasks (74.1%), recall exceeds precision.
    • The cross-task median recall is 0.83, compared to a median precision of 0.65 (and a mean recall of 0.75 vs. a mean precision of 0.62).

    Because generative LLMs are prone to generating false positives rather than missing true positive instances in these domains, zero-shot and few-shot LLMs serve effectively as high-recall first-stage filters in multi-stage annotation pipelines, where positive classifications are subsequently refined by precision-optimized classifiers or human review.

  5. Knowl 5 — LLM Label Consistency Score for Uncertainty Estimation

    equation

    Given a sequence of ll categorical predictions A=(A1,A2,…,Al)\mathbf{A} = (A_1, A_2, \dots, A_l) produced by repeatedly querying a generative LLM on the same text input with non-zero sampling temperature T>0T > 0, the empirical label consistency score C(A)∈[1/l,1]C(\mathbf{A}) \in [1/l, 1] is defined as the proportion of annotations matching the empirical modal category AmodeA_{\text{mode}}:

    C(A)=1l∑i=1lI(Ai=Amode)C(\mathbf{A}) = \frac{1}{l} \sum_{i=1}^{l} \mathbb{I}(A_i = A_{\text{mode}})

    where I(⋅)\mathbb{I}(\cdot) denotes the indicator function:

    I(Ai=Amode)={1if Ai=Amode0if Ai≠Amode\mathbb{I}(A_i = A_{\text{mode}}) = \begin{cases} 1 & \text{if } A_i = A_{\text{mode}} \\ 0 & \text{if } A_i \ne A_{\text{mode}} \end{cases}

    and Amode=arg⁡max⁡c∑i=1lI(Ai=c)A_{\text{mode}} = \arg\max_{c} \sum_{i=1}^l \mathbb{I}(A_i = c). In repeated annotation protocols with l≥5l \ge 5 iterations at temperature T=0.7T = 0.7, C(A)C(\mathbf{A}) measures the LLM's classification stability under stochastic token sampling.

  6. Knowl 6 — Error Identification via Annotation Consistency Filtering

    empirical result

    Computing the consistency score C(A)C(\mathbf{A}) over 5 stochastic runs at temperature T=0.7T = 0.7 identifies low-confidence and error-prone GPT-4 classifications:

    • Sample Distribution: 85.1% of all evaluated annotations achieve perfect consistency (C=1.0C = 1.0).
    • Accuracy Differential: Annotations with C=1.0C = 1.0 attain an average accuracy of 87.5%, whereas annotations with C<1.0C < 1.0 are 19.4 percentage points less likely to be correct.
    • Diagnostic Rates: Annotations with C=1.0C = 1.0 show a 16.4 percentage point increase in true positive rate (TPR) and a 21.4 percentage point increase in true negative rate (TNR) relative to annotations exhibiting inconsistency (C<1.0C < 1.0).

    Filtering instances by C<1.0C < 1.0 provides an automated mechanism to select ambiguous text samples for human inspection.

  7. Knowl 7 — GPT-4 vs. Supervised Fine-Tuned BERT Across Training Sample Sizes

    empirical result

    Benchmarking few-shot GPT-4 against encoder-only BERT-base classifiers (110M parameters, fine-tuned via grid search over learning rates {1e-5,2e-5,5e-5}\{1\text{e-}5, 2\text{e-}5, 5\text{e-}5\}, batch sizes {8,16}\{8, 16\}, and epochs {2,4,6}\{2, 4, 6\}) shows that relative model superiority depends on training sample volume:

    • Minimal Data Regime (n=250n = 250): Fine-tuned BERT is constrained by limited class representations and class imbalance; GPT-4 outperforms BERT on 18 of the 27 tasks (66.7%).
    • Moderate Data Regime (n=1,000n = 1{,}000): When training data increases to 1,000 labeled samples, fine-tuned BERT surpasses GPT-4 performance on 10 of the 14 replicated tasks (71.4%).

    Few-shot generative LLMs outperform supervised encoders when training annotations are extremely scarce (n≤250n \le 250), but fine-tuned supervised models systematically exceed GPT-4 performance once moderate human-annotated data (n≥1,000n \ge 1{,}000) becomes available.

  8. Knowl 8 — Impact of Prompt Optimization and Temperature Tuning on Annotation Quality

    empirical result

    Standard post-hoc adjustment strategies yield only marginal changes in automated annotation performance:

    • Manual Prompt Refinement: One round of human-in-the-loop prompt optimization based on error analysis of 250 validation instances produces small metric gains across tasks: median accuracy improves by 0.024, median F1 score improves by 0.018, and median precision improves by 0.043, while median recall exhibits no change (mean recall decreases by 0.028).
    • Temperature Tuning: Evaluating classification across sampling temperatures T∈{0.0,0.3,0.6,0.9}T \in \{0.0, 0.3, 0.6, 0.9\} across 7 annotation tasks shows no systematic relationship with F1 performance.

    Subpar zero-shot and few-shot LLM annotation performance is largely structural and cannot be fully resolved by prompt tuning or sampling hyperparameter adjustments alone.

  9. Knowl 9 — Determinants of LLM Annotation Accuracy: Text Complexity vs. Intercoder Reliability

    empirical result

    Evaluating GPT-4 performance against task and text characteristics reveals distinct predictive relationships:

    • Text Readability and Length: Average character length per text passage and readability (measured via the Flesch-Kincaid Grade Level score) are uncorrelated with GPT-4 F1 classification scores.
    • Human Intercoder Reliability: Across 21 tasks reporting human intercoder agreement (16 using Krippendorff's alpha and 5 using Cohen's kappa), tasks with low intercoder reliability invariably result in low GPT-4 F1 scores. However, tasks with moderate to high intercoder agreement exhibit high variance in GPT-4 F1, demonstrating that high human agreement is a necessary but insufficient condition for strong LLM annotation performance.
    • Task Type: Tasks focused on factual extraction and summarization achieve high F1 scores, whereas tasks requiring complex cultural context (e.g., differentiating racial, national, and religious identities) or implicit conversational inference perform poorly.
  10. Knowl 10 — Performance Deficit of Open-Source Mistral-7B Relative to GPT-4 in Annotation

    empirical result

    Comparing few-shot annotations generated by the open-source model Mistral-7B against GPT-4 on identical prompt instructions and test sets across the 27 benchmark tasks reveals a substantial capability disparity:

    • Mistral-7B achieves an average F1 score that is 0.16 lower than GPT-4.
    • Mistral-7B exhibits lower median values across F1, precision, and recall distributions compared to GPT-4.

    Smaller open-source language models experience higher classification error rates than proprietary frontiers in few-shot social science annotation tasks.

  11. Knowl 11 — Temporal Stability of GPT-4 Annotation Performance

    empirical result

    Re-evaluating a subset of 14 computational social science annotation tasks with GPT-4 over a six-month interval (April 2023 versus November 2023) using identical prompt templates and test sets showed minimal capability drift:

    • Median accuracy changed by +0.007+0.007.
    • Median F1 score changed by +0.022+0.022.
    • Median precision changed by +0.033+0.033.
    • Median recall changed by −0.041-0.041.

    GPT-4 classification performance remained broadly consistent across this period without signs of significant model degradation.

Coverage note — Deliberately omitted individual prompt text strings and detailed bibliographical summaries of the 11 original source papers beyond their aggregated use in the benchmark setup.

References

  1. 1.Allwein, E. L.; Schapire, R. E.; and Singer, Y. 2000. Reducing multiclass to binary: A unifying approach for margin classifiers. Journal of machine learning research, 1(Dec): 113–141.
  2. 2.Artstein, R.; and Poesio, M. 2008. Inter-coder agreement for computational linguistics. Computational linguistics, 34(4): 555–596.
  3. 3.Ashwin, J.; Chhabra, A.; and Rao, V. 2023. Using Large Language Models for Qualitative Analysis can Introduce Serious Bias. arXiv preprint arXiv:2309.17147.
  4. 4.Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020. Language Models are Few-Shot Learners. In Larochelle, H.; Ranzato, M.; Hadsell, R.; Balcan, M.; and Lin, H., eds., Advances in Neural Information Processing Systems, volume 33, 1877–1901. Curran Associates, Inc.
  5. 5.Card, D.; Chang, S.; Becker, C.; Mendelsohn, J.; Voigt, R.; Boustan, L.; Abramitzky, R.; and Jurafsky, D. 2022. Computational analysis of 140 years of US political speeches reveals more positive but increasingly polarized framing of immigration. Proceedings of the National Academy of Sciences of the United States of America, 31.
  6. 6.Chen, L.; Zaharia, M.; and Zou, J. 2023. How is ChatGPT's behavior changing over time? arXiv:2307.09009.
  7. 7.Chiang, C.-H.; and Lee, H.-y. 2023. Can Large Language Models Be an Alternative to Human Evaluations? arXiv preprint arXiv:2305.01937.
  8. 8.Cusimano, C.; and Goodwin, G. P. 2020. People judge others to have more voluntary control over beliefs than they themselves do. Journal of Personality and Social Psychology, 119.
  9. 9.Davani, A. M.; Díaz, M.; and Prabhakaran, V. 2022. Dealing with disagreements: Looking beyond the majority vote in subjective annotations. Transactions of the Association for Computational Linguistics, 10: 92–110.
  10. 10.Devlin, J.; Chang, M.-W.; Le, K.; and Toutanova, K. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171–4186.
  11. 11.Ding, B.; Qin, C.; Liu, L.; Bing, L.; Joty, S.; and Li, B. 2022. Is GPT-3 a Good Data Annotator? arXiv:2212.10450.
  12. 12.Egami, N.; Fong, C. J.; Grimmer, J.; Roberts, M. E.; and Stewart, B. M. 2022. How to make causal inferences using texts. Science Advances, 8(42): eabg2652.
  13. 13.Espinosa, L.; and Salathé, M. 2024. Use of large language models as a scalable approach to understanding public health discourse. medRxiv, 2024–02.
  14. 14.Falco, G.; Shneiderman, B.; Badger, J.; Carrier, R.; Dahbura, A.; Danks, D.; Eling, M.; Goodloe, A.; Gupta, J.; Hart, C.; Jirotka, M.; Johnson, H.; LaPointe, C.; Llorens, A. J.; Mackworth, A. K.; Maple, C.; Pálsson, S. E.; Pasquale, F.; Winfield, A.; and Yeong, Z. K. 2021. Governing AI safety through independent audits. Nature Machine Intelligence, 7: 566–571.
  15. 15.Felkner, V.; Thompson, J.; and May, J. 2024. GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14104–14115. Bangkok, Thailand: Association for Computational Linguistics.
  16. 16.Gebru, T.; Morgenstern, J.; Vecchione, B.; Vaughan, J. W.; Wallach, H.; Iii, H. D.; and Crawford, K. 2021. Datasheets for datasets. Communications of the ACM, 64(12): 86–92.
  17. 17.Gilardi, F.; Alizadeh, M.; and Kubli, M. 2023. ChatGPT outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences, 120(30): e2305016120.
  18. 18.Grimmer, J.; Roberts, M. E.; and Stewart, B. 2022. Text as Data: A New Framework for Machine Learning and the Social Sciences. Princeton University Press.
  19. 19.Grimmer, J.; and Stewart, B. M. 2013. Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts. Political Analysis, 21(3): 267–297.
  20. 20.He, X.; Lin, Z.; Gong, Y.; Jin, A.-L.; Zhang, H.; Lin, C.; Jiao, J.; Yiu, S. M.; Duan, N.; and Chen, W. 2023. AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators. arXiv:2303.16854.
  21. 21.Hopkins, D. J.; Lelkes, Y.; and Wolken, S. 2024. The Rise of and Demand for Identity-Oriented Media Coverage. American Journal of Political Science.
  22. 22.Huang, F.; Kwak, H.; and An, J. 2023. Is chatgpt better than human annotators? potential and limitations of chatgpt in explaining implicit hate speech. arXiv:2302.07736.
  23. 23.Jacobs, A. Z.; and Wallach, H. 2021. Measurement and fairness. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, 375–385.
  24. 24.Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; de las Casas, D.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; Lavaud, L. R.; Lachaux, M.-A.; Stock, P.; Scao, T. L.; Lavril, T.; Wang, T.; Lacroix, T.; and Sayed, W. E. 2023. Mistral 7B. arXiv:2310.06825.
  25. 25.Joshi, B.; Amini, M. R.; Partalas, I.; Iutzeler, F.; and Maximov, Y. 2017. Aggressive sampling for multi-class to binary reduction with applications to text classification. Advances in Neural Information Processing Systems, 30.
  26. 26.Keith, K. A.; Teichmann, C.; O'Connor, B.; and Meij, E. 2020. Uncertainty over uncertainty: Investigating the assumptions, annotations, and text measurements of economic policy uncertainty. arXiv preprint arXiv:2010.04706.
  27. 27.Khattab, O.; Singhvi, A.; Maheshwari, P.; Zhang, Z.; Santhanam, K.; Vardhamanan, S.; Haq, S.; Sharma, A.; Joshi, T. T.; Moazam, H.; Miller, H.; Zaharia, M.; and Potts, C. 2023. DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines. arXiv:2310.03714.
  28. 28.Knox, D.; Lucas, C.; and Cho, W. K. T. 2022. Testing Causal Theories with Learned Proxies. Annual Review of Political Science, 25.
  29. 29.Krippendorff, K. 2018. Content Analysis: An Introduction to Its Methodology. Sage, 4 edition.
  30. 30.Kumarage, T.; Bhattacharjee, A.; and Garland, J. 2024. Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection. arXiv:2403.08035.
  31. 31.Li, C.; and Flanigan, J. 2023. Task Contamination: Language Models May Not Be Few-Shot Anymore. arXiv:2312.16337.
  32. 32.Li, M.; Shi, T.; Ziems, C.; Kan, M.-Y.; Chen, N. F.; Liu, Z.; and Yang, D. 2023. CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data Annotation. arXiv:2310.15638.
  33. 33.Lorena, A. C.; De Carvalho, A. C.; and Gama, J. M. 2008. A review on the combination of binary classifiers in multiclass problems. Artificial Intelligence Review, 30: 19–37.
  34. 34.Magar, I.; and Schwartz, R. 2022. Data contamination: From memorization to exploitation. arXiv:2203.08242.
  35. 35.McKinney, W. 2011. pandas: a foundational Python library for data analysis and statistics. Python for high performance and scientific computing, 14(9): 1–9.
  36. 36.Mellon, J.; Bailey, J.; Scott, R.; Breckwoldt, J.; and Miori, M. 2022. Does GPT-3 know what the Most Important Issue is? Using Large Language Models to Code Open-Text Social Survey Responses At Scale. Working paper.
  37. 37.Mikhaylov, S.; Laver, M.; and Benoit, K. R. 2012. Coder reliability and misclassification in the human coding of party manifestos. Political Analysis, 20(1): 78–91.
  38. 38.Müller, S. 2022. The temporal focus of campaign communication. The Journal of Politics, 84(1): 585–590.
  39. 39.Neuendorf, K. A. 2016. The Content Analysis Guidebook. Sage Publications.
  40. 40.Ollion, E.; Shen, R.; Macanovic, A.; and Chatelain, A. 2023. ChatGPT for Text Annotation? Mind the Hype! arXiv:https://osf.io/preprints/socarxiv/x58kn.
  41. 41.Pan, A.; Chan, J. S.; Zou, A.; Li, N.; Basart, S.; Woodside, T.; Zhang, H.; Emmons, S.; and Hendrycks, D. 2023. Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark. Proceedings of the 40th International Conference on Machine Learning, 26837–26867.
  42. 42.Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; and et al., T. K. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems.
  43. 43.Peng, H.; Romero, D. M.; and Horvat, E.-A. 2022. Dynamics of cross-platform attention to retracted papers. Proceedings of the National Academy of Sciences, 119(25): 585–590.
  44. 44.Reiss, M. V. 2023. Testing the Reliability of ChatGPT for Text Annotation and Classification: A Cautionary Remark. arXiv:2304.11085.
  45. 45.Roberts, M. E.; Stewart, B. M.; and Nielsen, R. A. 2020. Adjusting for Confounding with Text Matching. American Journal of Political Science, 64(4): 887–903.
  46. 46.Schub, R. 2022. Informing the Leader: Bureaucracies and International Crises. American Political Science Review, 116.
  47. 47.Shneiderman, B. 2020. Human-centered artificial intelligence: Reliable, safe and trustworthy. International Journal of Human–Computer Interaction, 6: 495–504.
  48. 48.Spirling, A. 2023. Why open-source generative AI models are an ethical way forward for science. Nature, 413.
  49. 49.Thapa, S.; Naseem, U.; and Nasim, M. 2023. From humans to machines: can ChatGPT-like LLMs effectively replace human annotators in NLP tasks. Workshop Proceedings of the 17th International AAAI Conference on Web and Social Media.
  50. 50.Törnberg, P. 2023a. ChatGPT-4 Outperforms Experts and Crowd Workers in Annotating Political Twitter Messages with Zero-Shot Learning. arXiv:2304.06588.
  51. 51.Törnberg, P. 2023b. How to use LLMs for Text Analysis. arXiv:2307.13106.
  52. 52.Wang, S.; Liu, Y.; Xu, Y.; Zhu, C.; and Zeng, M. 2021. Want to reduce labeling cost? GPT-3 can help. arXiv:2108.13487.
  53. 53.Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; and Sam Shleifer, J. D.; von Platen, P.; Ma, C.; Jernite, Y.; Plu, J.; Xu, C.; Scao, T. L.; Gugger, S.; Drame, M.; Lhoest, Q.; and Rush, A. M. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of EMNLP.
  54. 54.Yang, K.-C.; and Menczer, F. 2023. Large language models can rate news outlet credibility. arXiv:2304.00228.
  55. 55.Zack, T.; Lehman, E.; Suzgun, M.; Rodriguez, J. A.; Celi, L. A.; Gichoya, J.; Jurafsky, D.; Szolovits, P.; Bates, D. W.; Abdulnour, R.-E. E.; Butte, A. J.; and Alsentzer, E. 2024. Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study. The Lancet Digital Health, 6(1): e12–e22.
  56. 56.Zhu, Y.; Zhang, P.; Haq, E.-U.; Hui, P.; and Tyson, G. 2023. Can ChatGPT Reproduce Human-Generated Labels? A Study of Social Computing Tasks. arXiv:2304.10145.
  57. 57.Ziems, C.; Held, W.; Shaikh, O.; Chen, J.; Zhang, Z.; and Yang, D. 2023. Can Large Language Models Transform Computational Social Science? Working paper.

Citation

MLA
Pangakis, N., and S. Wolken. “Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI”. Proceedings of the International AAAI Conference on Web and Social Media, vol. 19, 2025, pp. 1471–92, https://doi.org/10.1609/icwsm.v19i1.35883.
APA
Pangakis, N., & Wolken, S. (2025). Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI. Proceedings of the International AAAI Conference on Web and Social Media, 19, 1471–1492. https://doi.org/10.1609/icwsm.v19i1.35883
Chicago
Pangakis, N., and S. Wolken. 2025. “Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI”. Proceedings of the International AAAI Conference on Web and Social Media 19: 1471–92. https://doi.org/10.1609/icwsm.v19i1.35883.
Harvard
Pangakis, N. and Wolken, S. (2025) “Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI”, Proceedings of the International AAAI Conference on Web and Social Media, 19, pp. 1471–1492. Available at: https://doi.org/10.1609/icwsm.v19i1.35883.
Vancouver
1. Pangakis N, Wolken S (2025) Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI. Proceedings of the International AAAI Conference on Web and Social Media 19:1471–1492

BibTeX

@article{Pangakis_2025, title={Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI}, volume={19}, ISSN={2162-3449}, url={http://dx.doi.org/10.1609/icwsm.v19i1.35883}, DOI={10.1609/icwsm.v19i1.35883}, journal={Proceedings of the International AAAI Conference on Web and Social Media}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Pangakis, Nick and Wolken, Sam}, year={2025}, month=June, pages={1471–1492} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/