Scalable Membership Inference Attacks via Quantile Regression
Martín BertránShuai TangAaron RothMichael KearnsJamie MorgensternSteven Wu
Proposes a quantile regression framework that executes highly effective black-box membership inference attacks by training only a single model without requiring any knowledge of the target architecture.
Machine learning models often leak sensitive information about the records used to train them. Membership inference attacks aim to determine whether a specific data record was part of a target model's private training dataset. Existing state-of-the-art attacks typically train dozens or hundreds of "shadow models" to simulate target model behaviors. However, this shadow-model approach requires massive compute budgets and detailed knowledge of the target model's internal architecture, making it impractical against large, proprietary commercial systems.
The article introduces a scalable, black-box membership inference attack that requires training only a single quantile regression model. The main objective is to demonstrate that directly estimating sample-dependent confidence thresholds via pinball loss minimization achieves competitive or superior inference accuracy compared to expensive shadow-model baselines without requiring any knowledge of the target model's architecture.
The authors evaluate this method across standard image classification benchmarks (CIFAR-10, CIFAR-100, CINIC-10, and ImageNet-1k) and tabular datasets (OpenML and US Census ACS data). In the image domain, the target models included various ResNet architectures, while the single attack model utilized a fixed, off-the-shelf ConvNeXt architecture querying the target via black-box confidence scores. Tabular evaluations used gradient-boosted decision trees. Performance was measured by precision and true positive rates across fixed, low false positive rate thresholds.
The findings establish that the proposed quantile regression approach achieves state-of-the-art results on complex tasks while drastically cutting computational overhead. On ImageNet-1k, the single-model attack outperformed shadow-model baselines across all error rates, achieving 97.45% precision at a 1% false positive rate and 99.64% precision at a 0.1% false positive rate. On tabular benchmarks, the attack matched the accuracy of shadow-model methods requiring 16 to 64 models. Across all experiments, minimizing pinball loss directly corresponded to maximized membership inference accuracy. On smaller image datasets with limited data (CIFAR), the method outperformed standard marginal baselines but trailed shadow-model techniques due to sample scarcity.
These results demonstrate that proprietary commercial systems face severe membership privacy risks from adversaries with modest computational resources and no internal model access. For organizations deploying machine learning models, this shifts the privacy threat model by showing that architecture obfuscation provides no defense. Crucially, the same technique enables organizations to conduct fast, low-cost privacy auditing and compliance checks on their own models prior to deployment.
Organizations should incorporate quantile regression auditing into pre-deployment security workflows to detect training data leakage. Security teams can establish reliable auditing pipelines without the massive cost of training shadow ensembles. However, decision-makers should note that the attack's effectiveness depends on the availability of a representative public data sample and scales best with large datasets. When auditing smaller datasets, standard shadow models or parameterized distributions may still provide higher precision.
- Paper: Membership Inference Attacks Against Machine Learning Models, Reza Shokri et al. (2016). This seminal paper introduced membership inference attacks and the foundational shadow-model methodology that the source directly targets and aims to replace with an efficient quantile regression approach.
- Paper: LLM Dataset Inference: Did you train on my dataset?, Pratyush Maini et al. (2024). This work critiques and builds upon standard individual-record membership inference attacks by demonstrating their failures in large language models and extending the paradigm to aggregate dataset inference.
