Ghostbuster: Detecting Text Ghostwritten by Large Language Models
Vivek VermaEve FleisigNicholas TomlinDan Klein
Introduces Ghostbuster, a black-box AI text detector that achieves state-of-the-art accuracy across diverse writing domains and unseen generation models by combining token probabilities from smaller reference language models through structured feature search.
The rapid adoption of advanced artificial intelligence models capable of writing fluent text has raised urgent concerns regarding misinformation, media trustworthiness, and academic integrity. Existing automated detection systems often fail to generalize when encountering unfamiliar writing domains, novel prompts, or newer models, and simple metric approaches risk disproportionately misclassifying authentic writing, including essays by non-native English speakers. In response, the article introduces Ghostbuster, a detection system designed to accurately identify machine-generated text without requiring internal token probabilities from the generating model, making it effective for black-box and unknown architectures.
Ghostbuster relies on a three-stage framework that avoids the brittleness of pure perplexity thresholds and the overfitting common in deep neural classifiers. The system first scores candidate text using several weaker language models, including basic n-gram models and early GPT-3 variants. It then executes a structured algorithmic search to generate and select combinations of token probability functions across mathematical operations. Finally, a standard linear classifier evaluates these chosen features alongside select heuristic metrics, such as word length and probability outliers, to classify documents as either human- or AI-generated. The evaluation utilized three newly curated benchmark datasets across student essays, news articles, and creative writing, complemented by robustness tests and evaluations on human writing by non-native English speakers.
Across extensive testing, Ghostbuster demonstrated state-of-the-art performance. For in-domain evaluation, it achieved a 99.0 F1 score, outperforming DetectGPT by 41.6 points and GPTZero by 5.9 points. When tested across out-of-domain datasets, it maintained a strong 97.0 average F1 score, exceeding existing systems by 7.5 to 39.6 points and significantly outperforming deep neural baselines. In tests across unseen prompt styles, Ghostbuster sustained a 99.5 F1 score, while on text produced by an entirely unseen model (Claude), it led all baselines with a 92.2 F1 score. Extensive perturbation testing showed that the model remains robust against minor spelling, spacing, and sentence-ordering edits, though extensive paraphrasing tools can reduce recall.
These findings indicate that structured combinations of probabilities from weaker open models can reliably capture statistical signatures of machine-generated text across varied writing styles. This offers organizations a cost-effective, high-accuracy alternative to closed-source commercial detectors or computationally heavy neural classifiers. However, the results also demonstrate that detection accuracy decreases noticeably on short texts below 100 words and when classifying short essays by non-native English speakers, where Ghostbuster achieved 74.7% accuracy on a legacy short-essay benchmark compared to over 95% on longer samples.
Given the risks associated with false positives—such as wrongfully penalizing students—the article recommends that Ghostbuster should not be integrated into automated disciplinary pipelines without human supervision. Instead, stakeholders can deploy the system immediately for lower-risk tasks, such as filtering AI content from model training corpuses or verifying web source integrity. Future research should prioritize enhancing detection accuracy on short paragraphs, expanding multilingual and multi-dialect training coverage, and developing explainable outputs to support human reviewers.
- Paper: DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature, Eric Mitchell et al. (2023). Ghostbuster compares itself with DetectGPT, so this earlier probability-curvature detector clarifies the zero-shot baseline its feature-based approach improves upon.
No sufficiently relevant recommendations were found.
