Built independently by an author, for readers. Read the story and support ChapterPal

keyword

language model-based detectors

Language model-based detectors are automated computational systems designed to determine whether a given text was produced by artificial intelligence or composed by a human by leveraging the architecture, representations, or statistical properties of language models. These systems generally operate either as fine-tuned neural classifiers trained on labeled datasets of human and machine-generated writing or by evaluating predictive statistical signals such as token probabilities, perplexity, and structural regularity. They are commonly applied to maintain academic integrity, evaluate content authenticity, and curb the spread of automated misinformation, though their accuracy can vary depending on text length, stylistic modifications, domain shifts, and the rapid evolution of generative models.

1 item

Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods

Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods

Kathleen C. Fraser, Hillary Dawkins, Svetlana Kiritchenko

OrganizationsNational Research Council Canada

Why you should read this

Presents a comprehensive review of state-of-the-art AI-generated text detection methods, datasets, and practical factors that govern how reliably machine-written content can be identified across real-world scenarios.

Large language models (LLMs) have advanced to a point that even humans have difficulty discerning whether a text was generated by another human, or by a computer. However, knowing whether a text was produced by human or artificial intelligence (AI) is important to determining its trustworthiness, and has applications in many domains including detecting fraud and academic dishonesty, as well as combating the spread of misinformation and political propaganda. The task of AI-generated text (AIGT) detection is therefore both very challenging, and highly critical. In this survey, we summarize state-of-the art approaches to AIGT detection, including watermarking, statistical and stylistic analysis, and machine learning classification. We also provide information about existing datasets for this task. Synthesizing the research findings, we aim to provide insight into the salient factors that combine to determine how “detectable” AIGT text is under different scenarios, and to make practical recommendations for future work towards this significant technical and societal challenge.

Added

2026-09-26