Built independently by an author, for readers. Read the story and support ChapterPal

keyword

text perturbations

Text perturbations refer to modifications applied to written text, such as character alterations, synonym substitutions, word additions or deletions, and paraphrasing, typically designed to change the surface structure while retaining the underlying semantic meaning. In natural language processing and machine learning, these changes are commonly employed to evaluate model robustness, perform sensitivity analysis, generate adversarial examples, and augment training datasets. In applications such as content classification and artificial intelligence text detection, text perturbations serve as a standard method to measure how well analytical models withstand deliberate evasion strategies, noise, and stylistic variations.

1 item

Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods

Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods

Kathleen C. Fraser, Hillary Dawkins, Svetlana Kiritchenko

OrganizationsNational Research Council Canada

Why you should read this

Presents a comprehensive review of state-of-the-art AI-generated text detection methods, datasets, and practical factors that govern how reliably machine-written content can be identified across real-world scenarios.

Large language models (LLMs) have advanced to a point that even humans have difficulty discerning whether a text was generated by another human, or by a computer. However, knowing whether a text was produced by human or artificial intelligence (AI) is important to determining its trustworthiness, and has applications in many domains including detecting fraud and academic dishonesty, as well as combating the spread of misinformation and political propaganda. The task of AI-generated text (AIGT) detection is therefore both very challenging, and highly critical. In this survey, we summarize state-of-the art approaches to AIGT detection, including watermarking, statistical and stylistic analysis, and machine learning classification. We also provide information about existing datasets for this task. Synthesizing the research findings, we aim to provide insight into the salient factors that combine to determine how “detectable” AIGT text is under different scenarios, and to make practical recommendations for future work towards this significant technical and societal challenge.

Added

2026-09-26