Built independently by an author, for readers. Read the story and support ChapterPal

keyword

universal adversarial perturbations

Universal adversarial perturbations are input-agnostic modifications or noise patterns that, when applied across a wide variety of different inputs, cause machine learning models to fail or produce erroneous outputs. Unlike standard adversarial attacks that must be custom-crafted for each individual sample, a universal perturbation is a single, fixed modification capable of misleading a model across a large distribution of data points. In computer vision, these perturbations often consist of subtle, human-imperceptible pixel alterations that consistently trigger misclassifications across diverse images, while in natural language processing, they can take the form of transferable text patterns or prompt suffixes that systematically bypass safety alignments. Because they exploit shared geometric vulnerabilities within the decision boundaries of deep neural networks, universal perturbations frequently generalize across distinct network architectures, making them a fundamental focus in the study of artificial intelligence security, robustness, and model auditing.

3 items

Fingerprinting Deep Neural Networks Globally via Universal Adversarial Perturbations

Fingerprinting Deep Neural Networks Globally via Universal Adversarial Perturbations

Zirui Peng, Shaofeng Li, Guoxing Chen, Cheng Zhang, Haojin Zhu, Minhui Xue

OrganizationsCSIRO’s Data61Shanghai Jiao Tong UniversityThe Ohio State UniversityUniversity of Adelaide

Why you should read this

Proposes a practical deep neural network fingerprinting framework that leverages universal adversarial perturbations to capture global decision boundary geometry, enabling model owners to detect intellectual property theft through black-box queries with over 99.99% confidence.

In this paper, we propose a novel and practical mechanism to enable the service provider to verify whether a suspect model is stolen from the victim model via model extraction attacks. Our key insight is that the profile of a DNN model's decision boundary can be uniquely characterized by its Universal Adversarial Perturbations (UAPs). UAPs belong to a low-dimensional subspace and piracy models' subspaces are more consistent with victim model's subspace compared with non-piracy model. Based on this, we propose a UAP fingerprinting method for DNN models and train an encoder via contrastive learning that takes fingerprints as inputs, outputs a similarity score. Extensive studies show that our framework can detect model Intellectual Property (IP) breaches with confidence > 99.99 % within only 20 fingerprints of the suspect model. It also has good generalizability across different model architectures and is robust against post-modifications on stolen models.

Added

2026-09-26

Universal and Transferable Adversarial Attacks on Aligned Language Models

Universal and Transferable Adversarial Attacks on Aligned Language Models

Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, Matt Fredrikson

OrganizationsBosch Center for AICarnegie Mellon UniversityCenter for AI SafetyGoogle

Why you should read this

Exposes critical vulnerabilities within contemporary alignment techniques by demonstrating that automated gradient-driven suffix attacks can reliably breach the safety filters of production models.

Because "out-of-the-box" large language models are capable of generating a great deal of objectionable content, recent work has focused on aligning these models in an attempt to prevent undesirable generation. While there has been some success at circumventing these measures -- so-called "jailbreaks" against LLMs -- these attacks have required significant human ingenuity and are brittle in practice. In this paper, we propose a simple and effective attack method that causes aligned language models to generate objectionable behaviors. Specifically, our approach finds a suffix that, when attached to a wide range of queries for an LLM to produce objectionable content, aims to maximize the probability that the model produces an affirmative response (rather than refusing to answer). However, instead of relying on manual engineering, our approach automatically produces these adversarial suffixes by a combination of greedy and gradient-based search techniques, and also improves over past automatic prompt generation methods. Surprisingly, we find that the adversarial prompts generated by our approach are quite transferable, including to black-box, publicly released LLMs. Specifically, we train an adversarial attack suffix on multiple prompts (i.e., queries asking for many different types of objectionable content), as well as multiple models (in our case, Vicuna-7B and 13B). When doing so, the resulting attack suffix is able to induce objectionable content in the public interfaces to ChatGPT, Bard, and Claude, as well as open source LLMs such as LLaMA-2-Chat, Pythia, Falcon, and others. In total, this work significantly advances the state-of-the-art in adversarial attacks against aligned language models, raising important questions about how such systems can be prevented from producing objectionable information. Code is available at github.com/llm-attacks/llm-attacks.

Added

2026-05-08

License

Published with permission