Built independently by an author, for readers. Read the story and support ChapterPal

keyword

model robustness

Model robustness is the ability of a machine learning model to maintain consistent, reliable performance and make accurate predictions when exposed to unfamiliar conditions, distribution shifts, input noise, or deliberate perturbations. While standard models often perform effectively on data that closely mirrors their training distribution, a robust model avoids severe degradation when encountering out-of-distribution samples, environmental corruptions, irrelevant contextual distractions, spurious correlations, or adversarial attacks designed to deceive the system. Achieving robustness is essential for the safe and dependable deployment of artificial intelligence, as it ensures that a system relies on meaningful, invariant patterns rather than superficial shortcuts or fragile statistical artifacts.

10 items

Open Problems in Machine Unlearning for AI Safety

Open Problems in Machine Unlearning for AI Safety

Fazl Barez, Tingchen Fu, Ameya Prabhu, Stephen Casper, Amartya Sanyal, Adel Bibi, Aidan O'Gara, Robert Kirk, Ben Bucknall, Tim Fist, Luke Ong, Philip Torr, Kwok-Yan Lam, Robert Trager, David Krueger, Sören Mindermann, José Hernandez-Orallo, Mor Geva, Yarin Gal

OrganizationsLeverhulme Centre for the Future of IntelligenceMassachusetts Institute of TechnologyMila – Québec Artificial Intelligence InstituteNanyang Technological UniversitySingapore AI Safety InstituteTangenticTel Aviv UniversityUK AI Safety InstituteUniversitat Politècnica de ValènciaUniversity of CopenhagenUniversity of OxfordUniversity of Tübingen

Why you should read this

Exposes critical limitations of using machine unlearning for AI safety, showing why attempting to erase hazardous dual-use knowledge in domains like cybersecurity and biosecurity degrades beneficial capabilities and conflicts with existing safety mechanisms.

As AI systems become more capable, widely deployed, and increasingly autonomous in critical areas such as cybersecurity, biological research, and healthcare, ensuring their safety and alignment with human values is paramount. Machine unlearning -- the ability to selectively forget or suppress specific types of knowledge -- has shown promise for privacy and data removal tasks, which has been the primary focus of existing research. More recently, its potential application to AI safety has gained attention. In this paper, we identify key limitations that prevent unlearning from serving as a comprehensive solution for AI safety, particularly in managing dual-use knowledge in sensitive domains like cybersecurity and chemical, biological, radiological, and nuclear (CBRN) safety. In these contexts, information can be both beneficial and harmful, and models may combine seemingly harmless information for harmful purposes -- unlearning this information could strongly affect beneficial uses. We provide an overview of inherent constraints and open problems, including the broader side effects of unlearning dangerous knowledge, as well as previously unexplored tensions between unlearning and existing safety mechanisms. Finally, we investigate challenges related to evaluation, robustness, and the preservation of safety features during unlearning. By mapping these limitations and open challenges, we aim to guide future research toward realistic applications of unlearning within a broader AI safety framework, acknowledging its limitations and highlighting areas where alternative approaches may be required.

Added

2026-10-04

Correct-N-Contrast: a Contrastive Approach for Improving Robustness to Spurious Correlations

Correct-N-Contrast: a Contrastive Approach for Improving Robustness to Spurious Correlations

Michael Zhang, Nimit Sharad Sohoni, Hongyang R. Zhang, Chelsea Finn, Christopher Ré

OrganizationsNortheastern UniversityStanford University

Why you should read this

Presents a two-stage contrastive learning method that uses standard empirical risk minimization predictions to sample hard pairs and align representations across spurious attributes, achieving state-of-the-art worst-group accuracy without requiring group annotations during training.

Spurious correlations pose a major challenge for robust machine learning. Models trained with empirical risk minimization (ERM) may learn to rely on correlations between class labels and spurious attributes, leading to poor performance on data groups without these correlations. This is challenging to address when the spurious attribute labels are unavailable. To improve worst-group performance on spuriously correlated data without training attribute labels, we propose Correct-N-Contrast (CnC), a contrastive approach to directly learn representations robust to spurious correlations. As ERM models can be good spurious attribute predictors, CnC works by (1) using a trained ERM model’s outputs to identify samples with the same class but dissimilar spurious features, and (2) training a robust model with contrastive learning to learn similar representations for these samples. To support CnC, we introduce new connections between worst-group error and a representation alignment loss that CnC aims to minimize. We empirically observe that worst-group error closely tracks with alignment loss, and prove that the alignment loss over a class helps upper-bound the class’s worst-group vs. average error gap. On popular benchmarks, CnC reduces alignment loss drastically, and achieves state-of-the-art worst-group accuracy by 3.6% average absolute lift. CnC is also competitive with oracle methods that require group labels.

Added

2026-09-26

Shadows can be Dangerous: Stealthy and Effective Physical-world Adversarial Attack by Natural Phenomenon

Shadows can be Dangerous: Stealthy and Effective Physical-world Adversarial Attack by Natural Phenomenon

Yiqi Zhong, Xianming Liu, Deming Zhai, Junjun Jiang, Xiangyang Ji

OrganizationsHarbin Institute of TechnologyPeng Cheng LaboratoryTsinghua University

Why you should read this

Demonstrates that casting simple, natural shadows onto traffic signs can deceive vision models in physical-world black-box settings with success rates exceeding 90%, exposing a critical real-world vulnerability in autonomous systems.

Estimating the risk level of adversarial examples is essential for safely deploying machine learning models in the real world. One popular approach for physical-world attacks is to adopt the “sticker-pasting” strategy, which however suffers from some limitations, including difficulties in access to the target or printing by valid colors. A new type of non-invasive attacks emerged recently, which attempt to cast perturbation onto the target by optics based tools, such as laser beam and projector. However, the added optical patterns are artificial but not natural. Thus, they are still conspicuous and attention-grabbed, and can be easily noticed by humans. In this paper, we study a new type of optical adversarial examples, in which the perturbations are generated by a very common natural phenomenon, shadow, to achieve naturalistic and stealthy physical-world adversarial attack under the black-box setting. We extensively evaluate the effectiveness of this new attack on both simulated and real-world environments. Experimental results on traffic sign recognition demonstrate that our algorithm can generate adversarial examples effectively, reaching 98.23% and 90.47% success rates on LISA and GTSRB test sets respectively, while continuously misleading a moving camera over 95% of the time in real-world scenarios. We also offer discussions about the limitations and the defense mechanism of this attack1.

Added

2026-09-26

Understanding The Robustness in Vision Transformers

Understanding The Robustness in Vision Transformers

Daquan Zhou, Zhiding Yu, Enze Xie, Chaowei Xiao, Animashree Anandkumar, Jiashi Feng, José M. Álvarez

OrganizationsArizona State UniversityByteDanceCalifornia Institute of TechnologyNational University of SingaporeNVIDIAUniversity of Hong Kong

Why you should read this

Explains how visual grouping in self-attention reduces corruption sensitivity through an information bottleneck perspective, introducing fully attentional networks that integrate dynamic channel selection to achieve superior corruption error on ImageNet-C.

Recent studies show that Vision Transformers (ViTs) exhibit strong robustness against various corruptions. Although this property is partly attributed to the self-attention mechanism, there is still a lack of systematic understanding. In this paper, we examine the role of self-attention in learning robust representations. Our study is motivated by the intriguing properties of the emerging visual grouping in Vision Transformers, which indicates that self-attention may promote robustness through improved mid-level representations. We further propose a family of fully attentional networks (FANs) that strengthen this capability by incorporating an attentional channel processing design. We validate the design comprehensively on various hierarchical backbones. Our model achieves a state-of-the-art 87.1% accuracy and 35.8% mCE on ImageNet-1k and ImageNet-C with 76.8M parameters. We also demonstrate state-of-the-art accuracy and robustness in two downstream tasks: semantic segmentation and object detection. Code will be available at https://github.com/NVlabs/FAN.

Added

2026-09-26

Large Language Models Can Be Easily Distracted by Irrelevant Context

Large Language Models Can Be Easily Distracted by Irrelevant Context

Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H. Chi, Nathanael Schärli, Denny Zhou

OrganizationsGooglePurdue UniversityToyota Technological Institute at Chicago

Why you should read this

Demonstrates that irrelevant context causes severe performance drops in large language model reasoning on the GSM-IC benchmark, while providing concrete prompting and decoding methods to mitigate this distractibility.

Large language models have achieved impressive performance on various natural language processing tasks. However, so far they have been evaluated primarily on benchmarks where all information in the input context is relevant for solving the task. In this work, we investigate the distractibility of large language models, i.e., how the model problem-solving accuracy can be influenced by irrelevant context. In particular, we introduce Grade-School Math with Irrelevant Context (GSM-IC), an arithmetic reasoning dataset with irrelevant information in the problem description. We use this benchmark to measure the distractibility of cutting-edge prompting techniques for large language models, and find that the model performance is dramatically decreased when irrelevant information is included. We also identify several approaches for mitigating this deficiency, such as decoding with self-consistency and adding to the prompt an instruction that tells the language model to ignore the irrelevant information.

Added

2026-09-26

Decision-Based Adversarial Attacks: Reliable Attacks Against Black-Box Machine Learning Models

Decision-Based Adversarial Attacks: Reliable Attacks Against Black-Box Machine Learning Models

Wieland Brendel, Jonas Rauber, Matthias Bethge

OrganizationsUniversity of Tübingen

Why you should read this

Introduces the Boundary Attack, a method that generates adversarial perturbations using only the final class label of a black-box model and achieves performance competitive with gradient-based attacks on complex datasets like ImageNet.

Many machine learning algorithms are vulnerable to almost imperceptible perturbations of their inputs. So far it was unclear how much risk adversarial perturbations carry for the safety of real-world machine learning applications because most methods used to generate such perturbations rely either on detailed model information (gradient-based attacks) or on confidence scores such as class probabilities (score-based attacks), neither of which are available in most real-world scenarios. In many such cases one currently needs to retreat to transfer-based attacks which rely on cumbersome substitute models, need access to the training data and can be defended against. Here we emphasise the importance of attacks which solely rely on the final model decision. Such decision-based attacks are (1) applicable to real-world black-box models such as autonomous cars, (2) need less knowledge and are easier to apply than transfer-based attacks and (3) are more robust to simple defences than gradient- or score-based attacks. Previous attacks in this category were limited to simple models or simple datasets. Here we introduce the Boundary Attack, a decision-based attack that starts from a large adversarial perturbation and then seeks to reduce the perturbation while staying adversarial. The attack is conceptually simple, requires close to no hyperparameter tuning, does not rely on substitute models and is competitive with the best gradient-based attacks in standard computer vision tasks like ImageNet. We apply the attack on two black-box algorithms from this http URL. The Boundary Attack in particular and the class of decision-based attacks in general open new avenues to study the robustness of machine learning models and raise new questions regarding the safety of deployed machine learning systems. An implementation of the attack is available as part of Foolbox at this https URL .

Added

2026-09-24

Natural Adversarial Examples

Natural Adversarial Examples

Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, Dawn Song

OrganizationsUniversity of California BerkeleyUniversity of ChicagoUniversity of Washington

Why you should read this

Introduces ImageNet-A and ImageNet-O, two benchmarks of unmodified real-world images that expose shared blind spots in computer vision models by causing severe performance drops without synthetic pixel perturbations.

We introduce two challenging datasets that reliably cause machine learning model performance to substantially degrade. The datasets are collected with a simple adversarial filtration technique to create datasets with limited spurious cues. Our datasets' real-world, unmodified examples transfer to various unseen models reliably, demonstrating that computer vision models have shared weaknesses. The first dataset is called ImageNet-A and is like the ImageNet test set, but it is far more challenging for existing models. We also curate an adversarial out-of-distribution detection dataset called ImageNet-O, which is the first out-of-distribution detection dataset created for ImageNet models. On ImageNet-A a DenseNet-121 obtains around 2% accuracy, an accuracy drop of approximately 90%, and its out-of-distribution detection performance on ImageNet-O is near random chance levels. We find that existing data augmentation techniques hardly boost performance, and using other public training datasets provides improvements that are limited. However, we find that improvements to computer vision architectures provide a promising path towards robust models.

Added

2026-09-17