Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI
Nick PangakisSam Wolken
Demonstrates through twenty-seven private social science tasks that large language model annotations vary unpredictably and diverge from human judgment, proving that human validation remains essential for automated research workflows.
Natural language processing tools, particularly generative large language models, are increasingly deployed to automate the manual annotation of text data in social science and policy research. While automated annotation offers substantial cost and speed advantages over human coders, previous evaluations often relied on public benchmark datasets susceptible to data leakage and memorization. Consequently, it remained uncertain whether artificial intelligence could reliably replace human judgment on novel, real-world annotation tasks without introducing severe biases into downstream analyses.
The article evaluates how accurately generative large language models approximate human annotations on non-public datasets and tests a human-centered evaluation framework for responsible automated annotation. To avoid pre-training contamination, the authors replicated 27 manual annotation tasks across 11 password-protected datasets from computational social science studies published in high-impact journals within the past three years. Using OpenAI's GPT-4, the analysis evaluated more than 75,000 text samples against human ground-truth labels and benchmarked performance against over 500 fine-tuned supervised baseline models.
The findings reveal that while GPT-4 achieves solid overall results—with a median accuracy of 85.0% and a median F1 score of 0.707—performance is highly inconsistent across tasks. For one-third of the evaluated tasks, the model produced either precision or recall below 0.50, meaning that false positives outnumbered true positives or more than half of the true instances were missed. The model exhibited higher recall than precision in 20 of the 27 tasks, with a median recall of 0.83 compared to a median precision of 0.65. Furthermore, common interventions such as prompt refinement and temperature tuning yielded only marginal gains, and traditional supervised classifiers outperformed GPT-4 when sufficient training data (1,000 samples) was available.
These results demonstrate that generative models cannot be trusted to operate as unvalidated replacements for human annotators, especially on nuanced tasks requiring cultural context or conversational inference. Deploying generative tools without validation introduces substantial risks of misclassification and biased insights. However, the model's high recall makes it highly effective as an initial filtering mechanism within multi-stage annotation workflows, where automated systems identify potential positive cases for subsequent human verification.
Organizations should adopt human-centered evaluation workflows, validating automated classifications against task-specific, human-labeled samples before deploying models into production pipelines. Practitioners can also exploit model uncertainty by assessing label consistency across repeated queries, which successfully identifies ambiguous edge cases that require manual review. Where significant quantities of human-annotated data already exist, teams should consider deploying dedicated supervised classifiers rather than few-shot generative models. Confidence in these findings is high due to the evaluation of non-contaminated datasets, though researchers should note that baseline labels rely on the quality of original human coder agreements.
- Paper: ChatGPT outperforms crowd workers for text-annotation tasks, Fabrizio Gilardi et al. (2023). This foundational study demonstrates that zero-shot LLMs can outperform crowd workers on text annotation tasks, establishing the baseline paradigm of automated LLM annotation that the source paper critically evaluates and re-examines across protected datasets.
- Paper: Is GPT-3 a Good Data Annotator?, Bosheng Ding et al. (2023). It systematically explores using GPT models as automated data annotators for supervised learning, providing the technical and empirical background for evaluating LLM label generation against human ground truth.
- Paper: Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions, John Joon Young Chung et al. (2023). It provides necessary context on integrating human-in-the-loop verification strategies to maintain data accuracy when using LLMs for dataset construction and annotation.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). It introduces foundational methodologies and validation protocols for evaluating GPT-4's agreement with human judgment across complex evaluative tasks.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). This survey provides a comprehensive synthesis of LLM judgment capabilities, systematic biases, and alignment with human annotators, establishing key concepts scrutinized in the source.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). It investigates how prompt design and demonstrations impact in-context classification performance, informing the prompt optimization techniques analyzed in the source.
- Paper: GPT-4 Technical Report, OpenAI (2023). It documents the architecture, capabilities, and baseline evaluation performance of GPT-4, the primary model evaluated for text annotation in the source study.
- Paper: LLMs Corrupt Your Documents When You Delegate, Philippe Laban et al. (2026). It extends the source's findings on LLM judgment instability by assessing how repeated, multi-step document delegation leads to severe corruption and drift in real-world environments without human oversight.
- Paper: LLM4Eval@WSDM 2025: Large Language Model for Evaluation in Information Retrieval, Hossein A. Rahmani et al. (2025). It applies and expands the source's evaluation principles to benchmark the reliability, robustness, and human-alignment of LLMs when generating labels in information retrieval.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). It synthesizes broader methodological paradigms for structuring and testing LLM judgment systems, directly building upon the practical limitations and validation requirements highlighted in the source.
