Cross-Domain Detection of GPT-2-Generated Technical Text

Juan Diego RodriguezTodd HayDavid GrosZain ShamsiRavi Srinivasan

article2022NAACL81 citations

Demonstrates that machine-generated technical text detectors can successfully transfer across distinct scientific domains with only a few hundred labeled examples, enabling effective identification of synthetic content and paragraph tampering in full-length research papers.

Listen

Recent advances in large language models make it increasingly easy to generate convincing synthetic text, threatening information integrity and the peer-reviewed scientific record. While human readers and non-expert evaluators struggle to distinguish synthetic technical text from genuine research, manual vetting by subject matter experts is too slow and expensive to scale.

The article evaluates whether automated machine learning models trained on one scientific domain can reliably detect synthetic text in another discipline, and assesses how well these systems identify partially manipulated, full-length research papers.

The authors simulated realistic adversarial scenarios using the GPT-2 language model to generate technical abstracts and conditioned body paragraphs. They evaluated detectors built on the RoBERTa transformer architecture across physics and biomedical datasets, testing a two-stage training strategy where a detector is first trained on abundant out-of-domain proxy text and then refined with varying amounts of expert-labeled in-domain samples.

The analysis revealed that accurate cross-domain detection requires only a small investment in expert annotation. Adapting a physics-based detector to biomedical abstracts achieved approximately 90% accuracy with as few as 100 to 500 expert-labeled samples, whereas relying entirely on out-of-domain proxy data without expert labels capped accuracy below 70%. Additionally, pre-training the base detector on broad scientific and technical corpora consistently improved accuracy and increased resilience against domain shifts. When applied to full-length documents, paragraph-level classifiers reliably detected documents with extensive alterations, but struggled when only a single paragraph was replaced. Applying a length filter to remove short paragraphs below 500 to 1,000 characters reduced false alarms on authentic text, but simultaneously reduced detection of documents containing sparse synthetic insertions.

These findings demonstrate that organizations do not need vast target-domain datasets or exact knowledge of an attacker's generation setup to defend scientific literature. Instead, combining automated proxy data generation with modest, high-quality expert annotation offers an efficient, cost-effective detection pipeline. However, detecting subtle, low-volume tampering remains a persistent operational risk due to elevated false positive rates in short text segments.

Organizations should adopt staged training pipelines that leverage broad scientific pre-training and reserve subject matter expert effort for annotating compact calibration datasets of a few hundred samples. Because the primary limitation of this approach is vulnerability to short-text errors and subtle paragraph replacements, stakeholders should exercise caution when screening documents with minimal synthetic content, and future work must investigate detection against newer generation models, noisy expert labels, and fine-grained phrase substitutions.

Cover for Cross-Domain Detection of GPT-2-Generated Technical Text

Abstract

Machine-generated text presents a potential threat not only to the public sphere, but also to the scientific enterprise, whereby genuine research is undermined by convincing, synthetic text. In this paper we examine the problem of detecting GPT-2-generated technical research text. We first consider the realistic scenario where the defender does not have full information about the adversary's text generation pipeline, but is able to label small amounts of in-domain genuine and synthetic text in order to adapt to the target distribution. Even in the extreme scenario of adapting a physics-domain detector to a biomedical detector, we find that only a few hundred labels are sufficient for good performance. Finally, we show that paragraph-level detectors can be used to detect the tampering of full-length documents under a variety of threat models.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Threat model and defender capabilities
  • 3.1 Threat Model
  • 3.2 Defender Capabilities
  • Defender's access level
  • 4 Automated Detection
  • 4.1 Detecting Generated Abstracts
  • 4.2 Detecting Tampered Documents
  • 5 Datasets
  • 5.1 Real and Synthetic Abstracts
  • 5.1.1 Test Data
  • 5.1.2 Training Data
  • 5.2 Real and Tampered Documents
  • 5.2.1 Test Data
  • 5.2.2 Training Data
  • 6 Results
  • 6.1 Detection of Generated Abstracts
  • 6.2 Detection of Tampered Documents
  • 7 Conclusion
  • Acknowledgements
  • Ethical Considerations
  • References
  • A Appendix
  • A.1 Fine-tuning hyperparameters
  • A.2 Comparison of classifiers on in-domain detection
  • A.3 Full cross-domain results
  • A.4 Effect of paragraph length on detection performance
  • A.5 Effect of training on paragraphs generated without conditioning
  • A.6 STEM subset of the CORE dataset
  • A.7 Examples of generated technical text
  • A.8 Human-detected errors in generated biomedical abstracts

Knowls

  1. Knowl 1 — Two-Stage Sequential Task-Tuning for Cross-Domain Text Detection

    model/method

    To detect synthetic technical text when the adversary's generation pipeline and specific target domain are not known a priori, a defender can train a neural discriminator using a two-stage sequential task-tuning framework:

    1. Proxy Task-Tuning: A pretrained language model classifier (specifically RoBERTa) is first fine-tuned on a balanced dataset of real and synthetic text from an accessible proxy domain (such as physics or open-access scientific corpora) generated by a proxy generator (e.g., GPT-2 Medium fine-tuned on proxy data).
    2. Target Domain Task-Tuning: The intermediate detector checkpoint from step 1 is then fine-tuned a second time on a small set of labeled target-domain examples provided by a subject-matter expert (SME), consisting of balanced genuine human-written text and adversary-generated text.

    This sequential tuning approach enables the classifier to learn generic statistical artifacts of machine generation during proxy training while adapting its decision boundary to domain-specific features using minimal target-domain human labeling effort.

  2. Knowl 2 — Document-Level Tampering Aggregation Functions

    equation

    Given a document consisting of NN paragraphs, let si∈[0,1]s_i \in [0, 1] denote the predicted probability from a paragraph-level classifier that paragraph pip_i (for i∈{1,…,N}i \in \{1, \dots, N\}) is machine-generated. To evaluate document tampering under different defense objectives, three aggregation scoring functions are defined:

    1. Tampering Probability (S1): The probability that a document contains at least one machine-generated paragraph, assuming paragraph predictions are independent: P=1−∏i=1N(1−si)P = 1 - \prod_{i=1}^N (1 - s_i)

    2. Fraction of Generated Content (S2): The estimated proportion of generated paragraphs across the document: F=1N∑i=1N1si>0.5F = \frac{1}{N} \sum_{i=1}^N \mathbf{1}_{s_i > 0.5} where 1⋅\mathbf{1}_{\cdot} is the indicator function.

    3. Length-Thresholded Tampering Probability (S1-T): To reduce false positives caused by short paragraphs, only paragraphs exceeding a character length threshold TT are evaluated: P=1−∏i=1N(1−si⋅1len(pi)>T)P = 1 - \prod_{i=1}^N \left(1 - s_i \cdot \mathbf{1}_{\text{len}(p_i) > T}\right) where len(pi)\text{len}(p_i) denotes the number of characters in paragraph pip_i. If all paragraphs in a document are shorter than TT, the score is computed over the three longest paragraphs.

  3. Knowl 3 — Sample Efficiency and Domain Sensitivity in Cross-Domain Abstract Detection

    data/table

    Cross-domain detection accuracy of a RoBERTa-large classifier evaluated on a test set of 1,000 real and 1,000 GPT-2-generated biomedical abstracts from the Semantic Scholar corpus demonstrates the interplay between proxy sample count, target SME label count, and domain shift.

    SME Samples Biomedical Proxy Samples
    0 100 500 1k 10k 100k
    0 - 0.59 0.58 0.54 0.67 0.65
    100 0.68 0.67 0.89 0.89 0.81 0.82
    500 0.69 0.68 0.82 0.75 0.88 0.84
    1k 0.80 0.90 0.84 0.89 0.90 0.84
    10k 0.92 0.95 0.94 0.94 0.95 0.92
    Proxy Samples (nn) Biomedical Proxy Biomed+Physics Proxy Physics Proxy
    100 0.67 0.87 0.78
    500 0.89 0.82 0.83
    1k 0.89 0.85 0.84
    10k 0.81 0.69 0.62
    100k 0.82 0.70 0.66

    The second table shows accuracy when task-tuning on nn proxy samples followed by 100 SME target samples. Key observations:

    • Without SME labels (0 SME samples), proxy-only training achieves at most 0.67 accuracy.
    • Adding 100 in-domain SME samples boosts accuracy to 0.89 with 500-1k biomedical proxy samples.
    • Scaling proxy samples beyond 1,000 when SME samples are scarce (le100\\le 100) degrades performance due to negative transfer (accuracy drops from 0.89 to 0.81 for biomedical and from 0.84 to 0.62 for physics).
    • When plentiful target labels are available (1k or 10k SME samples), the classifier achieves 0.89-0.96 accuracy regardless of proxy domain.
  4. Knowl 4 — Domain-Adaptive Pretraining on STEM Text to Mitigate Domain Shift

    empirical result

    Performing domain-adaptive self-supervised pretraining of RoBERTa-large on a corpus of 916,074 scientific abstracts across science, technology, engineering, and mathematics (STEM) fields yields a model termed RoBERTa-large-STEM.

    Evaluating this model on cross-domain detection of generated technical abstracts shows:

    • Under matched proxy and target domains (biomedical proxy data), RoBERTa-large-STEM consistently outperforms standard RoBERTa-large by 1 to 5 points in accuracy across most sample sizes.
    • Under severe domain shift (such as training on 10,000 physics proxy abstracts and adapting with only 100 target biomedical SME samples), RoBERTa-large-STEM provides gains of over 20 percentage points in accuracy relative to RoBERTa-large.
    • With RoBERTa-large-STEM, 100 target SME samples achieve 0.91 accuracy using biomedical proxy data, while 500 SME samples are sufficient to reach 0.91 accuracy even when using out-of-domain physics proxy data.
  5. Knowl 5 — Document Tampering Detection via Length Thresholding

    data/table

    When detecting tampered biomedical documents where paragraphs are selectively replaced by GPT-2-generated paragraphs, performance depends heavily on the fraction of generated content and the paragraph length threshold TT used in length-thresholded scoring (S1-T). Paragraph classifiers were trained on 10,000 conditioned proxy paragraphs and 100 SME target paragraphs. Evaluations were conducted on test sets of 500 genuine human documents and 500 tampered documents across different tampering densities:

    Test Set T=500T = 500 (Accuracy [Prec, Recall]) T=1000T = 1000 (Accuracy [Prec, Recall])
    test-all-fake 0.87 [0.79, 1.00] 0.98 [0.96, 1.00]
    test-0.9 0.87 [0.79, 1.00] 0.97 [0.96, 0.99]
    test-0.5 0.86 [0.79, 0.99] 0.96 [0.96, 0.96]
    test-0.1 0.81 [0.77, 0.89] 0.80 [0.94, 0.63]
    test-1-fake 0.76 [0.74, 0.78] 0.67 [0.91, 0.37]
    Test Set P@100 P@250 P@500
    test-all-fake 1.00 1.00 0.99
    test-0.9 1.00 1.00 0.99
    test-0.5 0.98 0.97 0.95
    test-0.1 0.70 0.74 0.69
    test-1-fake 0.57 0.60 0.59

    Key takeaways:

    • Unthresholded aggregation (S1) yields high recall (0.99-1.00) but chance-level precision (0.55-0.56) due to false positives on short paragraphs.
    • Setting T=1000T = 1000 characters increases precision to 0.91-0.96 across all sets, achieving 0.96-0.98 accuracy on documents with ≥50%\ge 50\% replaced content. However, recall degrades severely on sparsely tampered documents (0.37 on test-1-fake) because the isolated synthetic paragraph is often shorter than 1000 characters and filtered out (filtering removed synthetic content in 64% of test-1-fake documents).
    • Ranking documents by estimated generated fraction (S2) retrieves tampered documents with precision up to 1.00 for dense tampering, but drops to 0.57-0.74 for sparse tampering.
  6. Knowl 6 — Paragraph-Level Tampering Threat Model with Context Conditioning

    experimental setup

    The experimental document tampering setup simulates an adversary replacing paragraphs within full-length scientific research articles (derived from S2ORC biomedical papers).

    • Adversary Model: The adversary fine-tunes GPT-2 Medium (355M parameters) on 890,000 biomedical paragraphs formatted as consecutive pairs: the last 256 tokens of paragraph AA followed by paragraph BB. During generation, the adversary conditions paragraph generation on the last 256 tokens of the preceding genuine paragraph, decoding with nucleus sampling (p∈[0.9,1.0]p \in [0.9, 1.0]) at temperature 1.01.0.
    • Defender Proxy Model: The defender constructs an independent proxy generator by fine-tuning GPT-2 Medium on a disjoint set of 890,000 biomedical paragraphs, generating synthetic training paragraphs conditioned on preceding paragraphs using nucleus sampling with temperature 0.80.8.
    • Evaluation Datasets: Five test suites (each containing 500 genuine and 500 tampered papers) represent varying levels of tampering: test-1-fake (exactly one paragraph replaced), test-0.1, test-0.5, test-0.9 (each paragraph replaced with probability 0.1, 0.5, and 0.9, respectively), and test-all-fake (all paragraphs generated sequentially conditioned on prior generated text).
  7. Knowl 7 — Text Length Sensitivity and Robustness to Paragraph Splitting Errors

    empirical result

    Paragraph-level synthetic text detection exhibits strong sensitivity to text length and paragraph segmentation boundaries:

    • Length Effects: When evaluating human and machine-generated paragraphs binned by character length, false positive rates and false negative rates increase dramatically for text segments containing fewer than 500 characters.
    • Paragraph Splitting Errors: When realistic parsing noise is introduced by segmenting text into fixed 5-sentence windows (using scispaCy) rather than true paragraph breaks, average segment length increases, raising precision under unthresholded scoring (S1) from 0.55 to 0.61-0.67. However, recall drops substantially on sparsely tampered sets (from 0.99 to 0.87 for test-0.1, and from 0.99 to 0.79 for test-1-fake).
    • Mixed-Content Degradation: Detection recall deteriorates when a segmented window contains a mixture of genuine and synthetic text. For 5-sentence windows containing 5, 4, and 3 synthetic sentences, detection recall drops from 0.95 to 0.66 and 0.33, respectively.
  8. Knowl 8 — Comparative Accuracy of Neural Classifiers for In-Domain Synthetic Text Detection

    data/table

    Evaluation of multiple discriminative architectures trained and tested on 10,000 in-domain scientific research abstracts (with human text from Semantic Scholar and synthetic text from an in-domain domain-tuned GPT-2 Medium generator):

    Discriminator Architecture Classification Accuracy
    Logistic Regression (Bag of Words) 0.64
    LSTM (GloVe embeddings) 0.67
    HAN (Hierarchical Attention Network, GloVe) 0.72
    BERT-base 0.86
    BERT-large 0.90
    XLNet-base 0.89
    XLNet-large 0.95
    RoBERTa-base 0.93
    RoBERTa-large 0.95

    RoBERTa-large and XLNet-large achieve the highest performance (0.95 accuracy). RoBERTa-large is preferred for downstream cross-domain and few-shot adaptation tasks due to known optimization instabilities of XLNet when fine-tuned on small datasets.

  9. Knowl 9 — Detection Invariance to Generation Conditioning in Proxy Data

    empirical result

    When building proxy training sets, the defender may generate synthetic paragraphs either conditionally (conditioned on the preceding paragraph prompt) or unconditionally (unprompted generation).

    • When no target SME labels are provided (0 SME samples), training RoBERTa-large on conditionally generated proxy paragraphs outperforms unconditionally generated proxy data by up to 0.10 in accuracy (0.93 vs. 0.83 at 10,000 proxy samples), driven primarily by an increase in recall of up to 0.21.
    • However, performing a second round of task-tuning with as few as 100 in-domain SME target samples closes this gap, reducing the accuracy difference between conditioned and unconditioned proxy training to only 0.02 (0.94 vs. 0.93 at 10,000 proxy samples). Thus, the defender does not strictly need to know whether the adversary conditioned on previous paragraphs if a small number of target SME labels are available.
  10. Knowl 10 — Typology of Artifacts and Errors in GPT-2 Generated Technical Text

    definition

    Subject-matter expert evaluation of GPT-2 generations fine-tuned on scientific technical corpora identifies eight characteristic categories of human-detectable errors:

    1. Non-words: Invented lexical items not present in standard or scientific vocabulary (e.g., ProBNER, gravidum, halliopeusing).
    2. Incorrect Acronyms: Misaligned abbreviation expansions or introduced abbreviations that are never defined or referenced properly (e.g., left middle cerebral artery (LMCMA), In Situ Analysis (SIA)).
    3. Coherence Breakdowns: Abrupt introduction of disjoint topics, mismatched conclusions, or unrelated domain concepts within a single abstract.
    4. Non-Existent Entities: Fictitious chemical formulas, physical models, or biological terms presented as factual (e.g., 4-OHDA instead of 6-OHDA, Nevographic Origin of Caustic Cygnosis).
    5. Domain/Context Inconsistencies: Impossible cross-discipline pairings or flawed empirical assertions (e.g., nonlinear RC receiver in a hydraulic grade, declaring a pathogen to be a molecule).
    6. Illogical or Contradictory Statements: Mathematical or semantic contradictions (e.g., claiming 31 out of 66 occlusions represents 90%, listing two rules after introducing 'three inference rules').
    7. Grammatical and Determiner Anomalies: Systematic omissions of articles, missing noun complements, or awkward phrasing around technical constructs.
    8. Lexical and Structural Repetition: Immediate re-use of identical noun phrases or clauses within consecutive sentences (e.g., STM based STM system, data from data).

Coverage note — All substantial methodological and empirical contributions regarding cross-domain detection of generated abstracts and tampered full documents are covered; specific listings of individual open-access CORE data providers (Table 10) and secondary hyperparameter sweeps were omitted.

References

  1. 1.David Ifeoluwa Adelani, Haotian Mai, Fuming Fang, Huy H. Nguyen, Junichi Yamagishi, and Isao Echizen. 2020. Generating sentiment-preserving fake online reviews using neural language models and their human- and machine-based detection. In Advanced Information Networking and Applications - Proceedings of the 34th International Conference on Advanced Information Networking and Applications, AINA-2020, Caserta, Italy, 15-17 April, volume 1151 of Advances in Intelligent Systems and Computing, pages 1341–1354. Springer.
  2. 2.Nadisha-Marie Aliman and Leon Kester. 2021. Epistemic defenses against scientific and empirical adversarial AI attacks. In CEUR Workshop Proceedings, 2021 Workshop on Artificial Intelligence Safety, AISafety 2021, 19 August 2021 through 20 August 2021. CEUR-WS.
  3. 3.Waleed Ammar, Dirk Groeneveld, Chandra Bhagavatula, Iz Beltagy, Miles Crawford, Doug Downey, Jason Dunkelberger, Ahmed Elgohary, Sergey Feldman, Vu Ha, Rodney Kinney, Sebastian Kohlmeier, Kyle Lo, Tyler Murray, Hsu-Han Ooi, Matthew Peters, Joanna Power, Sam Skjonsberg, Lucy Wang, Chris Wilhelm, Zheng Yuan, Madeleine van Zuylen, and Oren Etzioni. 2018. Construction of the literature graph in semantic scholar. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 3 (Industry Papers), pages 84–91, New Orleans - Louisiana. Association for Computational Linguistics.
  4. 4.Anton Bakhtin, Sam Gross, Myle Ott, Yuntian Deng, Marc’Aurelio Ranzato, and Arthur Szlam. 2019. Real or fake? Learning to discriminate machine from human generated text. arXiv preprint arXiv:1906.03351.
  5. 5.Alberto Bartoli and Eric Medvet. 2020. Exploring the potential of GPT-2 for generating fake reviews of research papers. In Fuzzy Systems and Data Mining VI: Proceedings of FSDM 2020, volume 331, pages 390–396. IOS Press.
  6. 6.Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. SciBERT: A pretrained language model for scientific text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3615–3620, Hong Kong, China. Association for Computational Linguistics.
  7. 7.Anja Belz. 2019. Fully automatic journalism: We need to talk about nonfake news generation. In Proceedings of the 2019 Truth and Trust Online Conference (TTO 2019), London, UK, October 4-5, 2019.
  8. 8.Meghana Moorthy Bhat and Srinivasan Parthasarathy. 2020. How effectively can machines defend against machine-generated fake news? an empirical study. In Proceedings of the First Workshop on Insights from Negative Results in NLP, pages 48–53, Online. Association for Computational Linguistics.
  9. 9.Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
  10. 10.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901.
  11. 11.Guillaume Cabanac and Cyril Labbé. 2021. Prevalence of nonsensical algorithmically generated papers in the scientific literature. Journal of the Association for Information Science and Technology.
  12. 12.Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. 2021. All that’s ‘human’ is not gold: Evaluating human evaluation of generated text. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7282–7296, Online. Association for Computational Linguistics.
  13. 13.Evan Crothers, Nathalie Japkowicz, Herna Viktor, and Paula Branco. 2022. Adversarial robustness of neural-statistical features in detection of generative transformers. arXiv preprint arXiv:2203.07983.
  14. 14.Tomasz Darmetko. 2021. Fake or not? generating adversarial examples from language models. Undergraduate thesis, Maastricht University.
  15. 15.Chris Donahue, Mina Lee, and Percy Liang. 2020. Enabling language models to fill in the blanks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2492–2501, Online. Association for Computational Linguistics.
  16. 16.Yao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah A Smith, and Yejin Choi. 2021. Scarecrow: A framework for scrutinizing machine text. arXiv preprint arXiv:2107.01294.
  17. 17.Tiziano Fagni, Fabrizio Falchi, Margherita Gambini, Antonio Martella, and Maurizio Tesconi. 2021. TweepFake: About detecting deepfake tweets. PLoS ONE, 16(5):e0251415.
  18. 18.Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 889–898, Melbourne, Australia. Association for Computational Linguistics.
  19. 19.Leon Fröhling and Arkaitz Zubiaga. 2021. Feature-based detection of automated language models: tackling GPT-2, GPT-3 and Grover. PeerJ Computer Science, 7:e443.
  20. 20.Sebastian Gehrmann, Hendrik Strobelt, and Alexander Rush. 2019. GLTR: Statistical detection and visualization of generated text. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 111–116, Florence, Italy. Association for Computational Linguistics.
  21. 21.Vivian Emily Gunser, Steffen Gottschling, Birgit Brucker, Sandra Richter, and Peter Gerjets. 2021. Can users distinguish narrative texts written by an artificial intelligence writing tool from purely human text? In International Conference on Human-Computer Interaction, pages 520–527. Springer.
  22. 22.Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360, Online. Association for Computational Linguistics.
  23. 23.Xiaochuang Han and Jacob Eisenstein. 2019. Unsupervised domain adaptation of contextualized embeddings for sequence labeling. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4238–4248, Hong Kong, China. Association for Computational Linguistics.
  24. 24.Fouzi Harrag, Maria Dabbah, Kareem Darwish, and Ahmed Abdelali. 2020. Bert transformer model for detecting Arabic GPT2 auto-generated tweets. In Proceedings of the Fifth Arabic Natural Language Processing Workshop, pages 207–214, Barcelona, Spain (Online). Association for Computational Linguistics.
  25. 25.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  26. 26.Dirk Hovy. 2016. The enemy in your own camp: How well can we detect statistically-generated fake reviews – an adversarial study. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 351–356, Berlin, Germany. Association for Computational Linguistics.
  27. 27.Daphne Ippolito, Daniel Duckworth, Chris Callison-Burch, and Douglas Eck. 2020. Automatic detection of generated text is easiest when humans are fooled. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1808–1822, Online. Association for Computational Linguistics.
  28. 28.Ganesh Jawahar, Muhammad Abdul-Mageed, and Laks Lakshmanan, V.S. 2020. Automatic detection of machine generated text: A critical survey. In Proceedings of the 28th International Conference on Computational Linguistics, pages 2296–2309, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  29. 29.Petr Knoth and Zdenek Zdrahal. 2012. CORE: three access levels to underpin open access. D-Lib Magazine, 18(11/12):1–13.
  30. 30.Sarah Kreps, R Miles McCain, and Miles Brundage. 2020. All the news that’s fit to fabricate: AI-generated text as a tool of media misinformation. Journal of Experimental Political Science.
  31. 31.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.
  32. 32.Vijini Liyanage, Davide Buscaldi, and Adeline Nazarenko. 2022. A benchmark corpus for the detection of automatically generated text in academic publications. arXiv preprint arXiv:2202.02013.
  33. 33.Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld. 2020. S2ORC: The semantic scholar open research corpus. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4969–4983, Online. Association for Computational Linguistics.
  34. 34.Patrice Lopez. 2009. GROBID: Combining automatic bibliographic data recognition and term extraction for scholarship publications. In International conference on theory and practice of digital libraries, pages 473–474. Springer.
  35. 35.Kaixin Ma, Jonathan Francis, Quanyang Lu, Eric Nyberg, and Alessandro Oltramari. 2019. Towards generalizable neuro-symbolic systems for commonsense question answering. In Proceedings of the First Workshop on Commonsense Inference in Natural Language Processing, pages 22–32, Hong Kong, China. Association for Computational Linguistics.
  36. 36.Anita Makri. 2017. Give the public the tools to trust scientists. Nature News, 541(7637):261.
  37. 37.Kris McGuffie and Alex Newhouse. 2020. The radicalization risks of GPT-3 and advanced neural language models. Technical report, Center on Terrorism, Extremism, and Counterterrorism, Middlebury Institute of International Studies at Monterrey.
  38. 38.Shaoor Munir, Brishna Batool, Zubair Shafiq, Padmini Srinivasan, and Fareed Zaffar. 2021. Through the looking glass: Learning to attribute synthetic text generated by language models. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1811–1822, Online. Association for Computational Linguistics.
  39. 39.Mark Neumann, Daniel King, Iz Beltagy, and Waleed Ammar. 2019. ScispaCy: Fast and robust models for biomedical natural language processing. In Proceedings of the 18th BioNLP Workshop and Shared Task, pages 319–327, Florence, Italy. Association for Computational Linguistics.
  40. 40.Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543, Doha, Qatar. Association for Computational Linguistics.
  41. 41.Jason Phang, Thibault Févry, and Samuel R Bowman. 2018. Sentence encoders on STILTs: Supplementary training on intermediate labeled-data tasks. arXiv preprint arXiv:1811.01088.
  42. 42.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  43. 43.Priyanka Ranade, Aritran Piplai, Sudip Mittal, Anupam Joshi, and Tim Finin. 2021. Generating fake cyber threat intelligence using transformer-based models. In International Joint Conference on Neural Networks, IJCNN 2021, Shenzhen, China, July 18-22, 2021, pages 1–9. IEEE.
  44. 44.Roei Schuster, Congzheng Song, Eran Tromer, and Vitaly Shmatikov. 2021. You autocomplete me: Poisoning vulnerabilities in neural code completion. In 30th USENIX Security Symposium (USENIX Security 21), pages 1559–1575. USENIX Association.
  45. 45.Tal Schuster, Roei Schuster, Darsh J. Shah, and Regina Barzilay. 2020. The limitations of stylometry for detecting machine-generated fake news. Computational Linguistics, 46(2):499–510.
  46. 46.Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881–7892, Online. Association for Computational Linguistics.
  47. 47.Yaman Kumar Singla, Swapnil Parekh, Somesh Singh, Junyi Jessy Li, Rajiv Ratn Shah, and Changyou Chen. 2021. AES are both overstable and oversensitive: Explaining why and proposing defenses. arXiv preprint arXiv:2109.11728.
  48. 48.Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al. 2019. Release strategies and the social impacts of language models. arXiv preprint arXiv:1908.09203.
  49. 49.Harald Stiff and Fredrik Johansson. 2021. Detecting computer-generated disinformation. International Journal of Data Science and Analytics, pages 1–21.
  50. 50.Yi Tay, Dara Bahri, Che Zheng, Clifford Brunk, Donald Metzler, and Andrew Tomkins. 2020. Reverse engineering configurations of neural text generation models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 275–279, Online. Association for Computational Linguistics.
  51. 51.Adaku Uchendu, Thai Le, Kai Shu, and Dongwon Lee. 2020. Authorship attribution for neural text generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8384–8395, Online. Association for Computational Linguistics.
  52. 52.Max Weiss. 2019. Deepfake bot submissions to federal public comment websites cannot be distinguished from human submissions. Technology Science.
  53. 53.Max Wolff. 2020. Attacking neural text detectors. In Towards Trustworthy ML: Rethinking Security and Privacy for ML Workshop, Eighth International Conference on Learning Representations, 2020.
  54. 54.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019. XLNet: Generalized autoregressive pretraining for language understanding. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  55. 55.Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. 2016. Hierarchical attention networks for document classification. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1480–1489, San Diego, California. Association for Computational Linguistics.
  56. 56.Yuanshun Yao, Bimal Viswanath, Jenna Cryan, Haitao Zheng, and Ben Y. Zhao. 2017. Automated crowdturfing attacks and defenses in online review systems. In Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, CCS 2017, Dallas, TX, USA, October 30 - November 03, 2017, pages 1143–1158. ACM.
  57. 57.Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019. Defending against neural fake news. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada.
  58. 58.Wanjun Zhong, Duyu Tang, Zenan Xu, Ruize Wang, Nan Duan, Ming Zhou, Jiahai Wang, and Jian Yin. 2020. Neural deepfake detection with factual structure of text. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2461–2470, Online. Association for Computational Linguistics.

Citation

MLA
Rodriguez, J. D., et al. “Cross-Domain Detection of GPT-2-Generated Technical Text”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 1213–33, https://doi.org/10.18653/v1/2022.naacl-main.88.
APA
Rodriguez, J. D., Hay, T., Gros, D., Shamsi, Z., & Srinivasan, R. (2022). Cross-Domain Detection of GPT-2-Generated Technical Text. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 1213–1233. https://doi.org/10.18653/v1/2022.naacl-main.88
Chicago
Rodriguez, J. D., T. Hay, D. Gros, Z. Shamsi, and R. Srinivasan. 2022. “Cross-Domain Detection of GPT-2-Generated Technical Text”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 1213–33. https://doi.org/10.18653/v1/2022.naacl-main.88.
Harvard
Rodriguez, J.D. et al. (2022) “Cross-Domain Detection of GPT-2-Generated Technical Text”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 1213–1233. Available at: https://doi.org/10.18653/v1/2022.naacl-main.88.
Vancouver
1. Rodriguez JD, Hay T, Gros D, Shamsi Z, Srinivasan R (2022) Cross-Domain Detection of GPT-2-Generated Technical Text. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 1213–1233

BibTeX

@inproceedings{rodriguez-etal-2022-cross,
    title = "Cross-Domain Detection of {GPT}-2-Generated Technical Text",
    author = "Rodriguez, Juan Diego  and
      Hay, Todd  and
      Gros, David  and
      Shamsi, Zain  and
      Srinivasan, Ravi",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.88/",
    doi = "10.18653/v1/2022.naacl-main.88",
    pages = "1213--1233"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/