Built independently by an author, for readers. Read the story and support ChapterPal

keyword

backdoor detection

Backdoor detection is the process of identifying hidden malicious manipulations, vulnerabilities, or triggers embedded within machine learning models and their associated data. In deep learning systems, backdoor attacks compromise a model during training or fine-tuning through data poisoning, direct weight modification, or supply-chain tampering, causing the model to maintain standard accuracy on benign inputs while producing adversary-specified outputs when presented with a specific trigger. Backdoor detection defenses operate across various phases of the machine learning pipeline by auditing training datasets for poisoned samples, inspecting pre-trained encoders and neural network representations to determine if a model has been trojaned, or evaluating inputs during the inference stage to intercept and filter out active trigger samples before they can cause misclassification.

7 items

Detecting Backdoors in Pre-trained Encoders

Detecting Backdoors in Pre-trained Encoders

Shiwei Feng, Guanhong Tao, Siyuan Cheng, Guangyu Shen, Xiangzhe Xu, Yingqi Liu, Kaiyuan Zhang, Shiqing Ma, Xiangyu Zhang

OrganizationsPurdue UniversityRutgers University

Why you should read this

Proposes DECREE, the first label-free backdoor detection framework designed to directly inspect self-supervised pre-trained encoders without relying on downstream classifier headers or full pre-training data.

Self-supervised learning in computer vision trains on unlabeled data, such as images (or image, text) pairs, to obtain an image encoder that learns high-quality embeddings for input data. Emerging backdoor attacks towards encoders expose critical vulnerabilities of self-supervised learning, since downstream classifiers (even further trained on clean data) may inherit backdoor behaviors from encoders. Existing backdoor detection methods mainly focus on supervised learning settings and cannot handle pre-trained encoders especially when input labels are not available. In this paper, we propose DECREE, the first backdoor detection approach for pre-trained encoders, requiring neither classifier headers nor input labels. We evaluate DECREE on over 400 encoders trojaned under 3 paradigms. We show the effectiveness of our method on image encoders pre-trained on ImageNet and OpenAI’s CLIP 400 million image-text pairs. Our method consistently has a high detection accuracy even if we have only limited or no access to the pre-training dataset. Code is available at https://github.com/GiantSeaweed/DECREE.

Added

2026-09-26

You Are Catching My Attention: Are Vision Transformers Bad Learners under Backdoor Attacks?

You Are Catching My Attention: Are Vision Transformers Bad Learners under Backdoor Attacks?

Zenghui Yuan, Pan Zhou, Kai Zou, Yu Cheng

OrganizationsHuazhong University of Science and TechnologyHubei Engineering Research Center on Big Data SecurityHubei Key Laboratory of Distributed System SecurityMicrosoftProtago Labs

Why you should read this

Reveals how the self-attention mechanism makes Vision Transformers uniquely vulnerable to patch-based backdoor attacks and introduces BadViT, an efficient attack framework that manipulates attention maps to implant stealthy triggers with minimal poisoning.

Vision Transformers (ViTs), which made a splash in the field of computer vision (CV), have shaken the dominance of convolutional neural networks (CNNs). However, in the process of industrializing ViTs, backdoor attacks have brought severe challenges to security. The success of ViTs benefits from the self-attention mechanism. However, compared with CNNs, we find that this mechanism of capturing global information within patches makes ViTs more sensitive to patch-wise triggers. Under such observations, we delicately design a novel backdoor attack framework for ViTs, dubbed BadViT, which utilizes a universal patch-wise trigger to catch the model's attention from patches beneficial for classification to those with triggers, thereby manipulating the mechanism on which ViTs survive to confuse itself. Furthermore, we propose invisible variants of BadViT to increase the stealth of the attack by limiting the strength of the trigger perturbation. Through a large number of experiments, it is proved that BadViT is an efficient backdoor attack method against ViTs, which is less dependent on the number of poisons, with satisfactory convergence, and is transferable for downstream tasks. Furthermore, the risks inside of ViTs to backdoor attacks are also explored from the perspective of existing advanced defense schemes.

Added

2026-09-26

Detecting Backdoors During the Inference Stage Based on Corruption Robustness Consistency

Detecting Backdoors During the Inference Stage Based on Corruption Robustness Consistency

Xiaogeng Liu, Minghui Li, Haoyu Wang, Shengshan Hu, Dengpan Ye, Hai Jin, Libing Wu, Chaowei Xiao

OrganizationsArizona State UniversityCluster and Grid Computing LabHuazhong University of Science and TechnologyHubei Engineering Research Center on Big Data SecurityHubei Key Laboratory of Distributed System SecurityNational Engineering Research Center for Big Data Technology and SystemServices Computing Technology and System LabWuhan University

Why you should read this

Presents TeCo, a practical test-time backdoor detection method that identifies trigger samples in strictly black-box, hard-label settings without requiring clean data or assumptions about trigger appearance by measuring prediction stability across diverse image corruptions.

Deep neural networks are proven to be vulnerable to backdoor attacks. Detecting the trigger samples during the inference stage, i.e., the test-time trigger sample detection, can prevent the backdoor from being triggered. However, existing detection methods often require the defenders to have high accessibility to victim models, extra clean data, or knowledge about the appearance of backdoor triggers, limiting their practicality. In this paper, we propose the test-time corruption robustness consistency evaluation (TeCo)¹, a novel test-time trigger sample detection method that only needs the hard-label outputs of the victim models without any extra information. Our journey begins with the intriguing observation that the backdoor-infected models have similar performance across different image corruptions for the clean images, but perform discrepantly for the trigger samples. Based on this phenomenon, we design TeCo to evaluate test-time robustness consistency by calculating the deviation of severity that leads to predictions' transition across different corruptions. Extensive experiments demonstrate that compared with state-of-the-art defenses, which even require either certain information about the trigger types or accessibility of clean data, TeCo outperforms them on different backdoor attacks, datasets, and model architectures, enjoying a higher AUROC by 10% and 5 times of stability.

Added

2026-09-26

BackdoorBench: A Comprehensive Benchmark of Backdoor Learning

BackdoorBench: A Comprehensive Benchmark of Backdoor Learning

Baoyuan Wu, Hongrui Chen, Mingda Zhang, Zihao Zhu, Shaokui Wei, Danni Yuan, Chao Shen

OrganizationsShenzhen Research Institute of Big DataThe Chinese University of Hong KongXi'an Jiaotong University

Why you should read this

Presents an extensible, modular benchmarking framework alongside 8,000 standardized evaluations across diverse attacks, defenses, datasets, and models to enable fair, reproducible comparisons in deep neural network backdoor learning.

Backdoor learning is an emerging and vital topic for studying deep neural networks' vulnerability (DNNs). Many pioneering backdoor attack and defense methods are being proposed, successively or concurrently, in the status of a rapid arms race. However, we find that the evaluations of new methods are often unthorough to verify their claims and accurate performance, mainly due to the rapid development, diverse settings, and the difficulties of implementation and reproducibility. Without thorough evaluations and comparisons, it is not easy to track the current progress and design the future development roadmap of the literature. To alleviate this dilemma, we build a comprehensive benchmark of backdoor learning called BackdoorBench. It consists of an extensible modular-based codebase (currently including implementations of 8 state-of-the-art (SOTA) attacks and 9 SOTA defense algorithms) and a standardized protocol of complete backdoor learning. We also provide comprehensive evaluations of every pair of 8 attacks against 9 defenses, with 5 poisoning ratios, based on 5 models and 4 datasets, thus 8,000 pairs of evaluations in total. We present abundant analysis from different perspectives about these 8,000 evaluations, studying the effects of different factors in backdoor learning. All codes and evaluations of BackdoorBench are publicly available at https://backdoorbench.github.io.

Added

2026-09-26

BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive Learning

BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive Learning

Siyuan Liang, Mingli Zhu, Aishan Liu, Baoyuan Wu, Xiaochun Cao, Ee-Chien Chang

OrganizationsBeihang UniversityNational University of SingaporeSun Yat-sen UniversityThe Chinese University of Hong Kong

Why you should read this

Proposes BadCLIP, a dual-embedding guided backdoor attack on multimodal contrastive learning models that aligns visual trigger patterns with target text and image representations to evade state-of-the-art detection and persist through clean fine-tuning.

While existing backdoor attacks have successfully infected multimodal contrastive learning models such as CLIP, they can be easily countered by specialized backdoor defenses for MCL models. This paper reveals the threats in this practical scenario and introduces the BadCLIP attack, which is resistant to backdoor detection and model fine-tuning defenses. To achieve this, we draw motivations from the perspective of the Bayesian rule and propose a dual-embedding guided framework for backdoor attacks. Specifically, we ensure that visual trigger patterns approximate the textual target semantics in the embedding space, making it challenging to detect the subtle parameter variations induced by backdoor learning on such natural trigger patterns. Additionally, we optimize the visual trigger patterns to align the poisoned samples with target vision features in order to hinder backdoor unlearning through clean fine-tuning. Our experiments show a significant improvement in attack success rate (+45.3% ASR) over current leading methods, even against state-of-the-art backdoor defenses, highlighting our attack’s effectiveness in various scenarios, including downstream tasks. Our codes can be found at https://github.com/LiangSiyuan21/BadCLIP.

Added

2026-09-26

BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain

BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain

Tianyu Gu, Brendan Dolan-Gavitt, Siddharth Garg

OrganizationsNew York University

Why you should read this

Exposes critical security vulnerabilities in outsourced machine learning by demonstrating that poisoned neural networks can achieve standard performance on normal tasks while triggering targeted misclassifications on attacker-manipulated inputs that persist through retraining.

Deep learning-based techniques have achieved state-of-the-art performance on a wide variety of recognition and classification tasks. However, these networks are typically computationally expensive to train, requiring weeks of computation on many GPUs; as a result, many users outsource the training procedure to the cloud or rely on pre-trained models that are then fine-tuned for a specific task. In this paper we show that outsourced training introduces new security risks: an adversary can create a maliciously trained network (a backdoored neural network, or a \emph{BadNet}) that has state-of-the-art performance on the user's training and validation samples, but behaves badly on specific attacker-chosen inputs. We first explore the properties of BadNets in a toy example, by creating a backdoored handwritten digit classifier. Next, we demonstrate backdoors in a more realistic scenario by creating a U.S. street sign classifier that identifies stop signs as speed limits when a special sticker is added to the stop sign; we then show in addition that the backdoor in our US street sign detector can persist even if the network is later retrained for another task and cause a drop in accuracy of {25}\% on average when the backdoor trigger is present. These results demonstrate that backdoors in neural networks are both powerful and---because the behavior of neural networks is difficult to explicate---stealthy. This work provides motivation for further research into techniques for verifying and inspecting neural networks, just as we have developed tools for verifying and debugging software.

Added

2026-09-14