Cross-Domain Detection of GPT-2-Generated Technical Text
Juan Diego RodriguezTodd HayDavid GrosZain ShamsiRavi Srinivasan
Demonstrates that machine-generated technical text detectors can successfully transfer across distinct scientific domains with only a few hundred labeled examples, enabling effective identification of synthetic content and paragraph tampering in full-length research papers.
Recent advances in large language models make it increasingly easy to generate convincing synthetic text, threatening information integrity and the peer-reviewed scientific record. While human readers and non-expert evaluators struggle to distinguish synthetic technical text from genuine research, manual vetting by subject matter experts is too slow and expensive to scale.
The article evaluates whether automated machine learning models trained on one scientific domain can reliably detect synthetic text in another discipline, and assesses how well these systems identify partially manipulated, full-length research papers.
The authors simulated realistic adversarial scenarios using the GPT-2 language model to generate technical abstracts and conditioned body paragraphs. They evaluated detectors built on the RoBERTa transformer architecture across physics and biomedical datasets, testing a two-stage training strategy where a detector is first trained on abundant out-of-domain proxy text and then refined with varying amounts of expert-labeled in-domain samples.
The analysis revealed that accurate cross-domain detection requires only a small investment in expert annotation. Adapting a physics-based detector to biomedical abstracts achieved approximately 90% accuracy with as few as 100 to 500 expert-labeled samples, whereas relying entirely on out-of-domain proxy data without expert labels capped accuracy below 70%. Additionally, pre-training the base detector on broad scientific and technical corpora consistently improved accuracy and increased resilience against domain shifts. When applied to full-length documents, paragraph-level classifiers reliably detected documents with extensive alterations, but struggled when only a single paragraph was replaced. Applying a length filter to remove short paragraphs below 500 to 1,000 characters reduced false alarms on authentic text, but simultaneously reduced detection of documents containing sparse synthetic insertions.
These findings demonstrate that organizations do not need vast target-domain datasets or exact knowledge of an attacker's generation setup to defend scientific literature. Instead, combining automated proxy data generation with modest, high-quality expert annotation offers an efficient, cost-effective detection pipeline. However, detecting subtle, low-volume tampering remains a persistent operational risk due to elevated false positive rates in short text segments.
Organizations should adopt staged training pipelines that leverage broad scientific pre-training and reserve subject matter expert effort for annotating compact calibration datasets of a few hundred samples. Because the primary limitation of this approach is vulnerability to short-text errors and subtle paragraph replacements, stakeholders should exercise caution when screening documents with minimal synthetic content, and future work must investigate detection against newer generation models, noisy expert labels, and fine-grained phrase substitutions.
- Paper: The Curious Case of Neural Text Degeneration, Ari Holtzman et al. (2020). This foundational study analyzes GPT-2's decoding mechanics and text generation artifacts, providing essential background on the distributional properties that machine-text detectors target.
- Paper: Extracting Training Data from Large Language Models, Nicholas Carlini et al. (2020). Understanding how statistical metrics reveal memorized and synthetic token sequences from GPT-2 offers fundamental context for building text forensic and detection classifiers.
- Paper: Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods, Nicholas Carlini et al. (2017). This seminal work establishes the foundational threat models and evasion challenges inherent in designing secondary detection mechanisms for machine learning systems.
- Paper: Can AI-Generated Text be Reliably Detected?, Vinu Sankar Sadasivan et al. (2026). Extends the study of AI text detection by systematically evaluating theoretical limits and the robustness of detectors against deliberate recursive paraphrasing attacks.
- Paper: The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing, Michał Brzozowski et al. (2026). Investigates downstream contamination in real-world academic publishing by detecting distinct generative statistical artifacts across major language model families.
- Paper: Scalable watermarking for identifying large language model outputs, Sumanth Dathathri et al. (2024). Explores an alternative, generative watermarking paradigm for reliably identifying large language model outputs at scale without relying solely on post-hoc classifiers.
- Paper: Towards Automating Scientific Review with Google's Paper Assistant Tool, Rajesh Jayaram et al. (2026). Addresses the broader challenge of safeguarding scientific rigor against automated and error-prone text by implementing automated verification tools for academic manuscripts.
