keyword
neural network classifiers
Neural network classifiers are machine learning models composed of interconnected layers of artificial neurons designed to assign input data into one or more predefined categories. By processing data through successive layers with adjustable weights and mathematical activation functions, these systems automatically learn complex, non-linear features and establish decision boundaries across high-dimensional spaces. The final layer typically outputs class probabilities or categorical predictions, often using functions such as softmax or sigmoid. Widely applied in perceptual domains such as image classification, speech processing, and natural language understanding, these models optimize their parameters during training to minimize classification error and generalize patterns to new, unseen inputs.
2 items

MagNet: A Two-Pronged Defense against Adversarial Examples
Dongyu Meng, Hao Chen
Why you should read this
Proposes MagNet, an attack-agnostic defense framework that combines detector and reformer networks with randomized diversity to identify and neutralize adversarial inputs without modifying the underlying neural network classifier.
Deep learning has shown promising results on hard perceptual problems in recent years. However, deep learning systems are found to be vulnerable to small adversarial perturbations that are nearly imperceptible to human. Such specially crafted perturbations cause deep learning systems to output incorrect decisions, with potentially disastrous consequences. These vulnerabilities hinder the deployment of deep learning systems where safety or security is important. Attempts to secure deep learning systems either target specific attacks or have been shown to be ineffective. In this paper, we propose MagNet, a framework for defending neural network classifiers against adversarial examples. MagNet does not modify the protected classifier or know the process for generating adversarial examples. MagNet includes one or more separate detector networks and a reformer network. Different from previous work, MagNet learns to differentiate between normal and adversarial examples by approximating the manifold of normal examples. Since it does not rely on any process for generating adversarial examples, it has substantial generalization power. Moreover, MagNet reconstructs adversarial examples by moving them towards the manifold, which is effective for helping classify adversarial examples with small perturbation correctly. We discuss the intrinsic difficulty in defending against whitebox attack and propose a mechanism to defend against graybox attack. Inspired by the use of randomness in cryptography, we propose to use diversity to strengthen MagNet. We show empirically that MagNet is effective against most advanced state-of-the-art attacks in blackbox and graybox scenarios while keeping false positive rate on normal examples very low.
Added
2026-09-25

Universal Adversarial Perturbations
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi, Pascal Frossard
Why you should read this
Reveals that deep neural networks can be systematically fooled on nearly all natural images using a single, image-agnostic perturbation vector that transfers across different model architectures.
Given a state-of-the-art deep neural network classifier, we show the existence of a universal (image-agnostic) and very small perturbation vector that causes natural images to be misclassified with high probability. We propose a systematic algorithm for computing universal perturbations, and show that state-of-the-art deep neural networks are highly vulnerable to such perturbations, albeit being quasi-imperceptible to the human eye. We further empirically analyze these universal perturbations and show, in particular, that they generalize very well across neural networks. The surprising existence of universal perturbations reveals important geometric correlations among the high-dimensional decision boundary of classifiers. It further outlines potential security breaches with the existence of single directions in the input space that adversaries can possibly exploit to break a classifier on most natural images.
Added
2026-09-14
