Ghostbuster: Detecting Text Ghostwritten by Large Language Models

Vivek VermaEve FleisigNicholas TomlinDan Klein

article2024NAACL163 citations

Introduces Ghostbuster, a black-box AI text detector that achieves state-of-the-art accuracy across diverse writing domains and unseen generation models by combining token probabilities from smaller reference language models through structured feature search.

Listen

The rapid adoption of advanced artificial intelligence models capable of writing fluent text has raised urgent concerns regarding misinformation, media trustworthiness, and academic integrity. Existing automated detection systems often fail to generalize when encountering unfamiliar writing domains, novel prompts, or newer models, and simple metric approaches risk disproportionately misclassifying authentic writing, including essays by non-native English speakers. In response, the article introduces Ghostbuster, a detection system designed to accurately identify machine-generated text without requiring internal token probabilities from the generating model, making it effective for black-box and unknown architectures.

Ghostbuster relies on a three-stage framework that avoids the brittleness of pure perplexity thresholds and the overfitting common in deep neural classifiers. The system first scores candidate text using several weaker language models, including basic n-gram models and early GPT-3 variants. It then executes a structured algorithmic search to generate and select combinations of token probability functions across mathematical operations. Finally, a standard linear classifier evaluates these chosen features alongside select heuristic metrics, such as word length and probability outliers, to classify documents as either human- or AI-generated. The evaluation utilized three newly curated benchmark datasets across student essays, news articles, and creative writing, complemented by robustness tests and evaluations on human writing by non-native English speakers.

Across extensive testing, Ghostbuster demonstrated state-of-the-art performance. For in-domain evaluation, it achieved a 99.0 F1 score, outperforming DetectGPT by 41.6 points and GPTZero by 5.9 points. When tested across out-of-domain datasets, it maintained a strong 97.0 average F1 score, exceeding existing systems by 7.5 to 39.6 points and significantly outperforming deep neural baselines. In tests across unseen prompt styles, Ghostbuster sustained a 99.5 F1 score, while on text produced by an entirely unseen model (Claude), it led all baselines with a 92.2 F1 score. Extensive perturbation testing showed that the model remains robust against minor spelling, spacing, and sentence-ordering edits, though extensive paraphrasing tools can reduce recall.

These findings indicate that structured combinations of probabilities from weaker open models can reliably capture statistical signatures of machine-generated text across varied writing styles. This offers organizations a cost-effective, high-accuracy alternative to closed-source commercial detectors or computationally heavy neural classifiers. However, the results also demonstrate that detection accuracy decreases noticeably on short texts below 100 words and when classifying short essays by non-native English speakers, where Ghostbuster achieved 74.7% accuracy on a legacy short-essay benchmark compared to over 95% on longer samples.

Given the risks associated with false positives—such as wrongfully penalizing students—the article recommends that Ghostbuster should not be integrated into automated disciplinary pipelines without human supervision. Instead, stakeholders can deploy the system immediately for lower-risk tasks, such as filtering AI content from model training corpuses or verifying web source integrity. Future research should prioritize enhancing detection accuracy on short paragraphs, expanding multilingual and multi-dialect training coverage, and developing explainable outputs to support human reviewers.

arXiv: 2305.15047vivek3141/ghostbuster

No sufficiently relevant recommendations were found.

Cover for Ghostbuster: Detecting Text Ghostwritten by Large Language Models

Abstract

We introduce Ghostbuster, a state-of-the-art system for detecting AI-generated text. Our method works by passing documents through a series of weaker language models, running a structured search over possible combinations of their features, and then training a classifier on the selected features to predict whether documents are AI-generated. Crucially, Ghostbuster does not require access to token probabilities from the target model, making it useful for detecting text generated by black-box or unknown models. In conjunction with our model, we release three new datasets of human- and AI-generated text as detection benchmarks in the domains of student essays, creative writing, and news articles. We compare Ghostbuster to several existing detectors, including DetectGPT and GPTZero, as well as a new RoBERTa baseline. Ghostbuster achieves 99.0 F1 when evaluated across domains, which is 5.9 F1 higher than the best preexisting model. It also outperforms all previous approaches in generalization across writing domains (+7.5 F1), prompting strategies (+2.1 F1), and language models (+4.4 F1). We also analyze our system’s robustness to a variety of perturbations and paraphrasing attacks, and evaluate its performance on documents by non-native English speakers.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Datasets
  • 4 Model
  • 4.1 Probability Computation
  • 4.2 Feature Selection
  • 4.3 Classifier Training
  • 5 Baselines
  • 6 Results
  • 6.1 In-domain Classification
  • 6.2 Generalization Across Domains
  • 6.3 Generalization Across Prompts
  • 6.4 Generalization Across Models
  • 7 Analysis
  • 7.1 Ablations
  • 7.2 Robustness
  • 7.3 Non-Native English Speaker Data
  • 7.4 Performance Across Document Lengths
  • 8 Conclusion
  • 9 Ethics and Limitations
  • Acknowledgments
  • References
  • A Prompting and Dataset Details
  • B Non-Native English Speaker Data Descriptions
  • C Additional Implementation Details
  • D Additional Features
  • E Best Features
  • F Additional Benchmarks
  • G Qualitative Analysis of Trends in Token Probabilities

Knowls

  1. Knowl 1 — Ghostbuster detects generated text with structured probability features

    model/method

    Ghostbuster classifies a document using token probabilities from weaker language models rather than requiring access to probabilities from the model that generated the text. It computes one probability vector per document from four models: a unigram fertility model, a Kneser–Ney trigram model, and the non-instruction-tuned GPT-3 models ada and davinci. The unigram and trigram models use Brown Corpus counts and the GPT-3 tokenizer vocabulary; the trigram model uses a discount factor of 0.90.9.

    The system searches for scalar features by combining probability vectors with six vector operations—addition, subtraction, multiplication, division, and elementwise greater-than and less-than indicators—and reducing the resulting vector with one of seven scalar operations: maximum, minimum, mean, mean of the 25 lowest values, vector length, L2L_2 norm, or population variance. The search starts from each model’s probability vector, uses a maximum depth of 3, and prunes repeated applications of the same function and duplicate combinations of commutative operations. This yields 2,534 candidate features at depth 3 (322 at depth 2). Forward feature selection is performed separately for each training dataset.

    The final classifier is logistic regression with L2L_2 regularization and C=1C=1. It uses the selected probability-based features plus seven handcrafted features measuring unusually large token probabilities, averages among the largest token probabilities (including differences between davinci and ada), and lengths of the longest words in tokens. The design combines a constrained, interpretable feature search with a linear classifier, and can be applied when the target generator is a black box.

  2. Knowl 2 — Three paired benchmarks cover essays, news, and creative writing

    data/table

    The paper introduces three datasets pairing human-authored documents with generated documents in student essays, news, and creative writing. Each domain contains 7,000 documents: 1,000 human documents, 5,000 ChatGPT documents, and 1,000 Claude documents. Of the ChatGPT documents, 1,000 use the original prompt and are eligible for training, validation, and test splits alongside the human documents; the remaining 4,000 use different prompts and are reserved for generalization evaluation. Claude documents are also evaluation-only. ChatGPT training examples were generated with gpt-3.5-turbo, and generation lengths were set to approximately match the paired human documents.

    DomainHuman sourceMedian words: humanMedian words: ChatGPTMedian words: Claude
    Student essaysIvyPanda529559442
    News articlesReuters 50-50498510384
    Creative writingr/WritingPrompts455512384

    For creative writing, ChatGPT was given the original writing prompt. The news and student-essay collections lacked original prompts or headlines, so ChatGPT first generated a corresponding headline or prompt from each human document, then generated a new document from it. This construction provides paired examples while limiting simple prompt-content and length cues.

  3. Knowl 3 — Ghostbuster achieves strong in-domain and cross-domain F1

    empirical result

    On document classification across the three domains, Ghostbuster achieves 99.0 F1 when trained and evaluated across all domains together. When trained and tested within individual domains, its F1 is 99.5 on news, 98.4 on creative writing, and 99.5 on student essays. For out-of-domain evaluation, the model is trained on two domains and tested on the held-out third domain; its F1 is 97.9 on news, 95.3 on creative writing, and 97.7 on student essays, averaging 97.0.

    ModelAll domains, in-domainNews, in-domainCreative writing, in-domainStudent essays, in-domainNews, out-of-domainCreative writing, out-of-domainStudent essays, out-of-domain
    Perplexity only81.582.284.192.171.949.093.4
    DetectGPT57.456.648.267.356.648.267.3
    GPTZero93.191.593.183.991.593.183.9
    RoBERTa98.199.497.697.488.395.771.4
    Ghostbuster99.099.598.499.597.995.397.7

    All values are F1. The comparison shows that Ghostbuster’s held-out-domain performance is higher than GPTZero’s by 7.5 F1 on average and is more consistent than the fine-tuned RoBERTa baseline, whose out-of-domain student-essay score is 71.4. DetectGPT and GPTZero are unsupervised, so their scores are the same across the in-domain and out-of-domain columns.

  4. Knowl 4 — Performance transfers across prompting strategies and to Claude text

    empirical result

    Ghostbuster was evaluated on ChatGPT-generated documents produced using prompting strategies not used for training, including requests for particular roles or styles and requests for short sentences. Across prompt variants, Ghostbuster reaches 99.5 F1, compared with 97.4 for RoBERTa and 96.1 for GPTZero. On documents generated by Claude, which was not used to train Ghostbuster, it reaches 92.2 F1; RoBERTa reaches 87.8 and GPTZero 75.6.

    ModelPrompt generalization F1Claude generalization F1
    Perplexity only85.384.1
    DetectGPT70.864.2
    GPTZero96.175.6
    RoBERTa97.487.8
    Ghostbuster99.592.2

    The supervised Ghostbuster and RoBERTa models were trained on human and ChatGPT text from all three domains. The results indicate strong transfer to prompt changes and better transfer to a new generator than the tested baselines, but the 6.8-point decrease from Ghostbuster’s 99.0 all-domain in-domain F1 to 92.2 on Claude shows that generalization across generator models remains harder.

  5. Knowl 5 — Ablations show the value of depth, neural-model probabilities, and structured search

    empirical result

    Ablations indicate that Ghostbuster’s feature search and its neural-language-model probability inputs are important, especially for generalization. With only handcrafted features, F1 is 80.5 across all in-domain data and 75.8, 77.2, and 77.2 on out-of-domain news, creative writing, and student essays, respectively; the full model scores 99.0 and 97.9, 95.3, and 97.7 on those conditions. Limiting search depth to 1 gives 93.7 all-domain in-domain F1, while depth 2 gives 98.3; increasing depth to 4 also gives 98.3, consistent with depth 3 being a useful compromise rather than a guarantee that deeper search helps.

    Removing both GPT-3 probability sources (ada and davinci) reduces all-domain in-domain F1 to 88.2 and out-of-domain F1 to 70.1 on news, 78.5 on creative writing, and 75.5 on student essays. Removing davinci alone has a smaller in-domain effect (98.8 all-domain F1) but lowers out-of-domain scores to 97.3, 90.3, and 91.9, respectively. Removing handcrafted features leaves all-domain in-domain F1 at 98.9 but lowers out-of-domain creative-writing F1 from 95.3 to 93.4. Thus, the structured search accounts for most of the performance, while davinci probabilities and handcrafted features contribute to generalization.

  6. Knowl 6 — Global reordering is robust, but repeated paraphrasing and local edits can evade detection

    empirical result

    The robustness evaluation modified essays with character-level edits (adding, deleting, or swapping characters), spacing changes, random capitalization or lowercasing, adjacent-word swaps, and synonym replacements, as well as sentence- or paragraph-level swaps and paraphrases. Ghostbuster’s F1 generally decreased gradually as local edits accumulated, and the authors report that numerous local edits were typically needed to cause false negatives. Swapping adjacent sentences or paragraphs had negligible impact, while repeated paraphrasing degraded performance more substantially.

    In a separate test, 100 AI-generated essays from each of the news, creative-writing, and student-essay domains were passed through the commercial evasion service Undetectable AI. Ghostbuster’s recall on AI-generated text fell from 99% to 62% after this processing. The results show that simple global reordering is not an effective evasion in this setup, but targeted rewriting and paraphrasing remain meaningful weaknesses.

  7. Knowl 7 — Short documents are less reliable, including in non-native English evaluations

    empirical result

    Ghostbuster’s performance decreases on shorter documents. When evaluated on documents trimmed to different lengths, the authors found substantial degradation at 100 tokens or fewer and performance that largely levels off at 500 tokens or more; they caution that the detector may be unreliable below 100 tokens. This length effect appeared both in-domain and under domain shifts.

    On human-only datasets by non-native English speakers, Ghostbuster’s accuracy was 95.5% on Lang8, 99.9% on TOEFL 11, and 74.7% on the 91 TOEFL essays collected by Liang et al. (2023). Because these evaluations contain no corresponding AI-generated documents, accuracy is equivalent to precision in this setting.

    DatasetMedian words per documentGhostbuster accuracy
    Lang87795.5%
    TOEFL 1131599.9%
    91 TOEFL essays10474.7%

    The 91 TOEFL essays and Lang8 documents are much shorter than TOEFL 11 or the paper’s primary datasets. The authors argue that length can explain much, but not necessarily all, of the lower results on these two datasets.

  8. Knowl 8 — Human annotators perform only modestly above chance on the benchmarks

    empirical result

    Human evaluation supports the claim that distinguishing the paired human and AI documents is difficult. In an initial study, six undergraduate and PhD students with experience using text-generation models each labeled a random set of 50 documents, balanced between human and AI text. Their average accuracy was 59%, with a maximum of 80% and a minimum of 34%. In a larger web-based evaluation, 233 people participated; among the 17 participants who made at least 25 guesses, average balanced binary-classification accuracy was 58.1% with a standard deviation of 11.1%, ranging from 39.7% to 82.0%.

  9. Knowl 9 — Token-probability trends differ between human and ChatGPT documents

    empirical result

    The paper analyzes average token probabilities across document positions for human and ChatGPT documents in the creative-writing, news, and student-essay domains, using GPT-3 ada and davinci. For a set of documents WW, the plotted value at token position ii is the average conditional log probability of the token at that position:

    f(i)=1∣W∣∑w∈Wlog⁡Pθ(wi∣w1,…,wi−1),f(i)=\frac{1}{|W|}\sum_{w\in W}\log P_{\theta}(w_i\mid w_1,\ldots,w_{i-1}),

    where ww is a document, wiw_i is its token at position ii, and PθP_{\theta} is the language model’s conditional token probability. The authors call the position-wise trend an entropy rate, while noting that their definition differs from the classical information-theoretic one. In all three domains, the curves fall sharply near the beginning and then plateau or decline more gradually. ChatGPT documents are generally more predictable than human documents, and the difference becomes more pronounced toward the end. The paper presents this as an example of a distributional signal that Ghostbuster’s features may exploit.

  10. Knowl 10 — The authors caution against using Ghostbuster for automatic punishment

    limitation

    Ghostbuster was trained and evaluated on three domains, but those datasets do not represent all writing styles or topics and predominantly contain British and American English. The authors identify likely failure risks for short documents, domains distant from the training data, English varieties outside standard American or British English, non-English text, writing by non-native English speakers, and AI text that has been edited or paraphrased. The paper’s experiments also focus on whole-document detection, not identifying AI-written passages within mixed-authorship documents.

    Because these limitations can make incorrect predictions more likely, the authors strongly discourage using Ghostbuster in systems that automatically penalize students or other writers. They recommend human supervision and consideration of additional evidence whenever an AI-generated classification could harm a person. They describe lower-risk uses, such as filtering generated text from language-model training data or flagging potentially generated online content, as more appropriate applications.

Coverage note — The paper’s proposed author-identification and paragraph-level benchmark designs are omitted because they are released as future-work testbeds rather than evaluated contributions; detailed prompt templates and the full per-domain feature lists are also omitted because the benchmark construction and feature-selection method capture their main contribution without reproducing implementation inventories.

References

  1. 1.Scott Aaronson. 2023. Watermarking of large language models. Workshop on Large Language Models and Transformers, Simons Institute, UC Berkeley.
  2. 2.Anton Bakhtin, Sam Gross, Myle Ott, Yuntian Deng, Marc’Aurelio Ranzato, and Arthur Szlam. 2019. Real or fake? learning to discriminate machine from human generated text. arXiv preprint arXiv:1906.03351.
  3. 3.Amrita Bhattacharjee, Tharindu Kumarage, Raha Moraffah, and Huan Liu. 2023. ConDA: Contrastive domain adaptation for AI-generated text detection.
  4. 4.Daniel Blanchard, Joel Tetreault, Derrick Higgins, Aoife Cahill, and Martin Chodorow. 2013. TOEFL11: A CORPUS OF NON-NATIVE ENGLISH. ETS Research Report Series, 2013(2):i–15.
  5. 5.Souradip Chakraborty, Amrit Singh Bedi, Sicheng Zhu, Bang An, Dinesh Manocha, and Furong Huang. 2023. On the possibilities of ai-generated text detection. arXiv preprint arXiv:2304.04736.
  6. 6.Yutian Chen, Hao Kang, Vivian Zhai, Liangze Li, Rita Singh, and Bhiksha Raj. 2023. GPT-Sentinel: Distinguishing human and ChatGPT generated content.
  7. 7.Yao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah A. Smith, and Yejin Choi. 2022. Is GPT-3 text indistinguishable from human text? scarecrow: A framework for scrutinizing machine text.
  8. 8.W Nelson Francis and Henry Kucera. 1979. Brown Corpus manual. Letters to the Editor, 5(2):7.
  9. 9.Sebastian Gehrmann, Hendrik Strobelt, and Alexander M. Rush. 2019. GLTR: Statistical detection and visualization of generated text.
  10. 10.Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu. 2023. How close is ChatGPT to human experts? comparison corpus, evaluation, and detection.
  11. 11.Will Heaven. 2023. ChatGPT is going to change education, not destroy it. MIT Technology Review.
  12. 12.John Houvardas and Efstathios Stamatatos. 2006. N-gram feature selection for authorship identification. In Artificial Intelligence: Methodology, Systems, Applications.
  13. 13.Daphne Ippolito, Daniel Duckworth, Chris Callison-Burch, and Douglas Eck. 2020. Automatic detection of generated text is easiest when humans are fooled.
  14. 14.IvyPanda. IvyPanda essay dataset.
  15. 15.Ganesh Jawahar, Muhammad Abdul-Mageed, and Laks V. S. Lakshmanan. 2020. Automatic detection of machine generated text: A critical survey.
  16. 16.Nurul Shamimi Kamaruddin, Amirrudin Kamsin, Lip Yee Por, and Hameedur Rahman. 2018. A review of text watermarking: Theory, methods, and applications. IEEE Access, 6:8011–8028.
  17. 17.John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. 2023. A watermark for large language models. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 17061–17084. PMLR.
  18. 18.Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, and James Zou. 2023. Gpt detectors are biased against non-native english writers. Patterns, 4(7):100779.
  19. 19.Clara Meister and Ryan Cotterell. 2021. Language model evaluation beyond perplexity.
  20. 20.Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning, and Chelsea Finn. 2023. DetectGPT: Zero-shot machine-generated text detection using probability curvature.
  21. 21.Tomoya Mizumoto, Mamoru Komachi, Masaaki Nagata, and Yuji Matsumoto. 2011. Mining revision log of language learning SNS for automated Japanese error correction of second language learners. In Proceedings of 5th International Joint Conference on Natural Language Processing, pages 147–155, Chiang Mai, Thailand. Asian Federation of Natural Language Processing.
  22. 22.OpenAI. 2019. GPT-2: 1.5b release.
  23. 23.Jiameng Pu, Zain Sarwar, Sifat Muhammad Abdullah, Abdullah Rehman, Yoonjin Kim, Parantapa Bhattacharya, Mobin Javed, and Bimal Viswanath. 2023. Deepfake text detection: Limitations and opportunities. In 2023 IEEE Symposium on Security and Privacy (SP), pages 1613–1630. IEEE.
  24. 24.Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang, and Soheil Feizi. 2023. Can AI-generated text be reliably detected?
  25. 25.Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, Miles McCain, Alex Newhouse, Jason Blazakis, Kris McGuffie, and Jasmine Wang. 2019. Release strategies and the social impacts of language models.
  26. 26.Edward Tian. 2023. GPTZero: Home.
  27. 27.Adaku Uchendu, Thai Le, Kai Shu, and Dongwon Lee. 2020. Authorship attribution for neural text generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8384–8395, Online. Association for Computational Linguistics.
  28. 28.Vivek Verma, Nicholas Tomlin, and Dan Klein. 2023. Revisiting entropy rate constancy in text. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 15537–15549, Singapore. Association for Computational Linguistics.
  29. 29.Xuandong Zhao, Prabhanjan Ananth, Lei Li, and Yu-Xiang Wang. 2023. Provable robust watermarking for AI-generated text. arXiv preprint arXiv:2306.17439.

Citation

MLA
Verma, V., et al. “Ghostbuster: Detecting Text Ghostwritten by Large Language Models”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 1702–17, https://doi.org/10.18653/v1/2024.naacl-long.95.
APA
Verma, V., Fleisig, E., Tomlin, N., & Klein, D. (2024). Ghostbuster: Detecting Text Ghostwritten by Large Language Models. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 1702–1717. https://doi.org/10.18653/v1/2024.naacl-long.95
Chicago
Verma, V., E. Fleisig, N. Tomlin, and D. Klein. 2024. “Ghostbuster: Detecting Text Ghostwritten by Large Language Models”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 1702–17. https://doi.org/10.18653/v1/2024.naacl-long.95.
Harvard
Verma, V. et al. (2024) “Ghostbuster: Detecting Text Ghostwritten by Large Language Models”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1702–1717. Available at: https://doi.org/10.18653/v1/2024.naacl-long.95.
Vancouver
1. Verma V, Fleisig E, Tomlin N, Klein D (2024) Ghostbuster: Detecting Text Ghostwritten by Large Language Models. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 1702–1717

BibTeX

@inproceedings{verma-etal-2024-ghostbuster,
    title = "Ghostbuster: Detecting Text Ghostwritten by Large Language Models",
    author = "Verma, Vivek  and
      Fleisig, Eve  and
      Tomlin, Nicholas  and
      Klein, Dan",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.95/",
    doi = "10.18653/v1/2024.naacl-long.95",
    pages = "1702--1717"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/