Built independently by an author, for readers. Read the story and support ChapterPal

keyword

adversarial perturbation

An adversarial perturbation is a deliberate and carefully crafted modification applied to an input data point or model parameter designed to mislead a machine learning model or maximize its prediction error. These perturbations are typically computed via mathematical optimization, such as gradient-guided ascent on the loss function, to expose vulnerabilities where small, bounded, and often human-imperceptible changes cause dramatic shifts in a model output. They can manifest across various modalities, including pixel adjustments in computer vision, token substitutions in natural language processing, or physical disturbances in sensor and trajectory data. In addition to serving as a benchmark for evaluating model vulnerabilities during security attacks, adversarial perturbations are frequently utilized during training as regularization mechanisms to smooth decision boundaries, enhance robustness, and improve model generalization.

7 items

Towards Efficient and Scalable Sharpness-Aware Minimization

Towards Efficient and Scalable Sharpness-Aware Minimization

Yong Liu, Siqi Mai, Xiangning Chen, Cho-Jui Hsieh, Yang You

OrganizationsNational University of SingaporeUniversity of California, Los Angeles

Why you should read this

Proposes LookSAM and a layer-wise variant that dramatically accelerate Sharpness-Aware Minimization to first-order optimizer speeds while enabling vision transformer training at batch sizes up to 64k in under an hour without sacrificing generalization performance.

Recently, Sharpness-Aware Minimization (SAM), which connects the geometry of the loss landscape and generalization, has demonstrated a significant performance boost on training large-scale models such as vision transformers. However, the update rule of SAM requires two sequential (non-parallelizable) gradient computations at each step, which can double the computational overhead. In this paper, we propose a novel algorithm LookSAM - that only periodically calculates the inner gradient ascent, to significantly reduce the additional training cost of SAM. The empirical results illustrate that LookSAM achieves similar accuracy gains to SAM while being tremendously faster - it enjoys comparable computational complexity with first-order optimizers such as SGD or Adam. To further evaluate the performance and scalability of LookSAM, we incorporate a layer-wise modification and perform experiments in the large-batch training scenario, which is more prone to converge to sharp local minima. Equipped with the proposed algorithms, we are the first to successfully scale up the batch size when training Vision Transformers (ViTs). With a 64k batch size, we are able to train ViTs from scratch in minutes while maintaining competitive performance. The code is available here: https://github.com/yong-6/LookSAM

Added

2026-10-05

On Adversarial Robustness of Trajectory Prediction for Autonomous Vehicles

On Adversarial Robustness of Trajectory Prediction for Autonomous Vehicles

Qingzhao Zhang, Shengtuo Hu, Jiachen Sun, Qi Alfred Chen, Z. Morley Mao

OrganizationsUniversity of California, IrvineUniversity of Michigan

Why you should read this

Reveals critical safety vulnerabilities in autonomous vehicle trajectory prediction by developing physically plausible adversarial driving trajectories that induce severe planning errors, alongside practical defense strategies using data augmentation and smoothing.

Trajectory prediction is a critical component for autonomous vehicles (AVs) to perform safe planning and navigation. However, few studies have analyzed the adversarial robustness of trajectory prediction or investigated whether the worst-case prediction can still lead to safe planning. To bridge this gap, we study the adversarial robustness of trajectory prediction models by proposing a new adversarial attack that perturbs normal vehicle trajectories to maximize the prediction error. Our experiments on three models and three datasets show that the adversarial prediction increases the prediction error by more than 150%. Our case studies show that if an adversary drives a vehicle close to the target AV following the adversarial trajectory, the AV may make an inaccurate prediction and even make unsafe driving decisions. We also explore possible mitigation techniques via data augmentation and trajectory smoothing.

Added

2026-09-26

Pre-Trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness

Pre-Trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness

Sibo Wang, Jie Zhang, Zheng Yuan, Shiguang Shan

OrganizationsInstitute of Computing Technology, Chinese Academy of SciencesUniversity of Chinese Academy of Sciences

Why you should read this

Proposes a fine-tuning method that aligns adversarial features with representations from the original pre-trained model to prevent overfitting and boost CLIP's zero-shot defense performance on unseen datasets.

Large-scale pre-trained vision-language models like CLIP have demonstrated impressive performance across various tasks, and exhibit remarkable zero-shot generalization capability, while they are also vulnerable to imperceptible adversarial examples. Existing works typically employ adversarial training (fine-tuning) as a defense method against adversarial examples. However, direct application to the CLIP model may result in overfitting, compromising the model's capacity for generalization. In this paper, we propose Pre-trained Model Guided Adversarial Fine-Tuning (PMG-AFT) method, which leverages supervision from the original pre-trained model by carefully designing an auxiliary branch, to enhance the model's zero-shot adversarial robustness. Specifically, PMG-AFT minimizes the distance between the features of adversarial examples in the target model and those in the pre-trained model, aiming to preserve the generalization features already captured by the pre-trained model. Extensive Experiments on 15 zero-shot datasets demonstrate that PMG-AFT significantly outperforms the state-of-the-art method, improving the top-1 robust accuracy by an average of 4.99%. Furthermore, our approach consistently improves clean accuracy by an average of 8.72%. Our code is available at here.1

Added

2026-09-26

Shadows can be Dangerous: Stealthy and Effective Physical-world Adversarial Attack by Natural Phenomenon

Shadows can be Dangerous: Stealthy and Effective Physical-world Adversarial Attack by Natural Phenomenon

Yiqi Zhong, Xianming Liu, Deming Zhai, Junjun Jiang, Xiangyang Ji

OrganizationsHarbin Institute of TechnologyPeng Cheng LaboratoryTsinghua University

Why you should read this

Demonstrates that casting simple, natural shadows onto traffic signs can deceive vision models in physical-world black-box settings with success rates exceeding 90%, exposing a critical real-world vulnerability in autonomous systems.

Estimating the risk level of adversarial examples is essential for safely deploying machine learning models in the real world. One popular approach for physical-world attacks is to adopt the “sticker-pasting” strategy, which however suffers from some limitations, including difficulties in access to the target or printing by valid colors. A new type of non-invasive attacks emerged recently, which attempt to cast perturbation onto the target by optics based tools, such as laser beam and projector. However, the added optical patterns are artificial but not natural. Thus, they are still conspicuous and attention-grabbed, and can be easily noticed by humans. In this paper, we study a new type of optical adversarial examples, in which the perturbations are generated by a very common natural phenomenon, shadow, to achieve naturalistic and stealthy physical-world adversarial attack under the black-box setting. We extensively evaluate the effectiveness of this new attack on both simulated and real-world environments. Experimental results on traffic sign recognition demonstrate that our algorithm can generate adversarial examples effectively, reaching 98.23% and 90.47% success rates on LISA and GTSRB test sets respectively, while continuously misleading a moving camera over 95% of the time in real-world scenarios. We also offer discussions about the limitations and the defense mechanism of this attack1.

Added

2026-09-26

Virtual Adversarial Training: A Regularization Method for Supervised and Semi-Supervised Learning

Virtual Adversarial Training: A Regularization Method for Supervised and Semi-Supervised Learning

Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, Shin Ishii

OrganizationsATR Cognitive Mechanisms LaboratoriesKyoto UniversityPreferred Networks, Inc.Ritsumeikan University

Why you should read this

Introduces Virtual Adversarial Training, a computationally efficient regularization technique that enforces prediction smoothness against label-free perturbations, successfully extending adversarial training to unlabeled data to improve semi-supervised classification performance.

We propose a new regularization method based on virtual adversarial loss: a new measure of local smoothness of the conditional label distribution given input. Virtual adversarial loss is defined as the robustness of the conditional label distribution around each input data point against local perturbation. Unlike adversarial training, our method defines the adversarial direction without label information and is hence applicable to semi-supervised learning. Because the directions in which we smooth the model are only "virtually" adversarial, we call our method virtual adversarial training (VAT). The computational cost of VAT is relatively low. For neural networks, the approximated gradient of virtual adversarial loss can be computed with no more than two pairs of forward- and back-propagations. In our experiments, we applied VAT to supervised and semi-supervised learning tasks on multiple benchmark datasets. With a simple enhancement of the algorithm based on the entropy minimization principle, our VAT achieves state-of-the-art performance for semi-supervised learning tasks on SVHN and CIFAR-10.

Added

2026-09-12

How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models

How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models

Dun Li Chan, Emily Liu, Niyathi Allu, Christian Hoang

OrganizationsFPT UniversityINTI International College Penang

Why you should read this

Demonstrates that standard output metrics fail to capture internal representational disruptions by tracing how text perturbations propagate through hidden-state geometry and attention-head mechanics across decoder-only language models.

Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is usually evaluated only through output behavior. We study how six naturalistic and synthetic input perturbations propagate through decoder-only language models at three levels: output behavior, hidden-state geometry, and attention-head function. We evaluate behavioral effects across four GPT-2 and two Qwen2.5 checkpoints by analyzing layerwise geometry using centered kernel alignment and intrinsic dimension, and examine attention-head responses in GPT-2. Perturbation types produce distinguishable metric profiles that are not fully captured by output measures and are only partly consistent across the tested checkpoints. Copying scores are especially associated with activation-patching recovery under token substitution and shuffling. Gradient-guided HotFlip perturbations also cause stronger behavioral and representational disruption than rate-matched random token substitutions in GPT-2; their behavioral effects are consistent across all six tested checkpoints. Our results show that robustness claims based on a single behavioral or representational metric can be misleading, and motivate multi-level evaluation of how perturbations alter language-model computation.

Added

2026-09-05