keyword
universal adversarial perturbations
Universal adversarial perturbations are input-agnostic modifications or noise patterns that, when applied across a wide variety of different inputs, cause machine learning models to fail or produce erroneous outputs. Unlike standard adversarial attacks that must be custom-crafted for each individual sample, a universal perturbation is a single, fixed modification capable of misleading a model across a large distribution of data points. In computer vision, these perturbations often consist of subtle, human-imperceptible pixel alterations that consistently trigger misclassifications across diverse images, while in natural language processing, they can take the form of transferable text patterns or prompt suffixes that systematically bypass safety alignments. Because they exploit shared geometric vulnerabilities within the decision boundaries of deep neural networks, universal perturbations frequently generalize across distinct network architectures, making them a fundamental focus in the study of artificial intelligence security, robustness, and model auditing.
3 items

Fingerprinting Deep Neural Networks Globally via Universal Adversarial Perturbations
Zirui Peng, Shaofeng Li, Guoxing Chen, Cheng Zhang, Haojin Zhu, Minhui Xue
Why you should read this
Proposes a practical deep neural network fingerprinting framework that leverages universal adversarial perturbations to capture global decision boundary geometry, enabling model owners to detect intellectual property theft through black-box queries with over 99.99% confidence.
In this paper, we propose a novel and practical mechanism to enable the service provider to verify whether a suspect model is stolen from the victim model via model extraction attacks. Our key insight is that the profile of a DNN model's decision boundary can be uniquely characterized by its Universal Adversarial Perturbations (UAPs). UAPs belong to a low-dimensional subspace and piracy models' subspaces are more consistent with victim model's subspace compared with non-piracy model. Based on this, we propose a UAP fingerprinting method for DNN models and train an encoder via contrastive learning that takes fingerprints as inputs, outputs a similarity score. Extensive studies show that our framework can detect model Intellectual Property (IP) breaches with confidence > 99.99 % within only 20 fingerprints of the suspect model. It also has good generalizability across different model architectures and is robust against post-modifications on stolen models.
Added
2026-09-26

Universal Adversarial Perturbations
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi, Pascal Frossard
Why you should read this
Reveals that deep neural networks can be systematically fooled on nearly all natural images using a single, image-agnostic perturbation vector that transfers across different model architectures.
Given a state-of-the-art deep neural network classifier, we show the existence of a universal (image-agnostic) and very small perturbation vector that causes natural images to be misclassified with high probability. We propose a systematic algorithm for computing universal perturbations, and show that state-of-the-art deep neural networks are highly vulnerable to such perturbations, albeit being quasi-imperceptible to the human eye. We further empirically analyze these universal perturbations and show, in particular, that they generalize very well across neural networks. The surprising existence of universal perturbations reveals important geometric correlations among the high-dimensional decision boundary of classifiers. It further outlines potential security breaches with the existence of single directions in the input space that adversaries can possibly exploit to break a classifier on most natural images.
Added
2026-09-14

Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, Matt Fredrikson
Why you should read this
Exposes critical vulnerabilities within contemporary alignment techniques by demonstrating that automated gradient-driven suffix attacks can reliably breach the safety filters of production models.
Because "out-of-the-box" large language models are capable of generating a great deal of objectionable content, recent work has focused on aligning these models in an attempt to prevent undesirable generation. While there has been some success at circumventing these measures -- so-called "jailbreaks" against LLMs -- these attacks have required significant human ingenuity and are brittle in practice. In this paper, we propose a simple and effective attack method that causes aligned language models to generate objectionable behaviors. Specifically, our approach finds a suffix that, when attached to a wide range of queries for an LLM to produce objectionable content, aims to maximize the probability that the model produces an affirmative response (rather than refusing to answer). However, instead of relying on manual engineering, our approach automatically produces these adversarial suffixes by a combination of greedy and gradient-based search techniques, and also improves over past automatic prompt generation methods. Surprisingly, we find that the adversarial prompts generated by our approach are quite transferable, including to black-box, publicly released LLMs. Specifically, we train an adversarial attack suffix on multiple prompts (i.e., queries asking for many different types of objectionable content), as well as multiple models (in our case, Vicuna-7B and 13B). When doing so, the resulting attack suffix is able to induce objectionable content in the public interfaces to ChatGPT, Bard, and Claude, as well as open source LLMs such as LLaMA-2-Chat, Pythia, Falcon, and others. In total, this work significantly advances the state-of-the-art in adversarial attacks against aligned language models, raising important questions about how such systems can be prevented from producing objectionable information. Code is available at github.com/llm-attacks/llm-attacks.
Added
2026-05-08
License
Published with permission
