keyword
input-space attacks
Input-space attacks are adversarial techniques in machine learning where an attacker modifies raw input data directly to manipulate a model into making incorrect predictions or activating hidden behaviors such as backdoors. Unlike feature-space or latent-space attacks that manipulate the internal representations learned within intermediate layers of a neural network, input-space attacks operate exclusively on the original data domain, such as changing pixel values in an image, characters in text, or samples in an audio signal, before any model processing occurs. These modifications are typically designed under specific mathematical constraints to remain imperceptible or visually realistic to human observers while still exploiting the model decision boundaries. Operating directly at the data ingestion stage makes input-space attacks highly realistic threat vectors against deployed artificial intelligence systems, as they can be delivered through standard user interfaces without requiring internal access to model architecture or hidden embeddings.
1 item

