Built independently by an author, for readers. Read the story and support ChapterPal

keyword

false positive rates

The false positive rate is the proportion of actual negative instances that are incorrectly identified or classified as positive by a binary classification model, diagnostic tool, or statistical test. Calculated mathematically as the number of false positives divided by the total number of actual negatives, which is the sum of false positives and true negatives, it represents the probability of committing a Type I error or generating a false alarm. In evaluation frameworks, hypothesis testing, and machine learning, this metric is critical for assessing model accuracy, establishing decision thresholds, and analyzing performance disparities across different groups or conditions to ensure reliable and equitable predictive outcomes.

3 items

Scalable Membership Inference Attacks via Quantile Regression

Scalable Membership Inference Attacks via Quantile Regression

Martín Bertrán, Shuai Tang, Aaron Roth, Michael Kearns, Jamie Morgenstern, Steven Wu

OrganizationsAmazon Web ServicesCarnegie Mellon UniversityUniversity of PennsylvaniaUniversity of Washington

Why you should read this

Proposes a quantile regression framework that executes highly effective black-box membership inference attacks by training only a single model without requiring any knowledge of the target architecture.

Membership inference attacks are designed to determine, using black box access to trained models, whether a particular example was used in training or not. Membership inference can be formalized as a hypothesis testing problem. The most effective existing attacks estimate the distribution of some test statistic (usually the model’s confidence on the true label) on points that were (and were not) used in training by training many shadow models—i.e. models of the same architecture as the model being attacked, trained on a random subsample of data. While effective, these attacks are extremely computationally expensive, especially when the model under attack is large. We introduce a new class of attacks based on performing quantile regression on the distribution of confidence scores induced by the model under attack on points that are not used in training. We show that our method is competitive with state-of-the-art shadow model attacks, while requiring substantially less compute because our attack requires training only a single model. Moreover, unlike shadow model attacks, our proposed attack does not require any knowledge of the architecture of the model under attack and is therefore truly “black-box”. We show the efficacy of this approach in an extensive series of experiments on various datasets and model architectures. Our code is available at github.com/amazon-science/quantile-mia.

Added

2026-09-26

RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors

RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors

Liam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu, Josh Magnus Ludan, Hainiu Xu, Daphne Ippolito, Chris Callison-Burch

OrganizationsCarnegie Mellon UniversityKing's College LondonUniversity College LondonUniversity of Pennsylvania

Why you should read this

Introduces a six-million-generation benchmark across eleven language models, eight domains, and eleven adversarial attacks, revealing that top AI text detectors easily fail when faced with minor sampling changes, repetition penalties, or unseen models.

Many commercial and open-source models claim to detect machine-generated text with extremely high accuracy (99% or more). However, very few of these detectors are evaluated on shared benchmark datasets and even when they are, the datasets used for evaluation are insufficiently challenging—lacking variations in sampling strategy, adversarial attacks, and open-source generative models. In this work we present RAID: the largest and most challenging benchmark dataset for machine-generated text detection. RAID includes over 6 million generations spanning 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies. Using RAID, we evaluate the out-of-domain and adversarial robustness of 8 open- and 4 closed-source detectors and find that current detectors are easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models. We release our data1 along with a leaderboard2 to encourage future research.

Added

2026-09-26