Built independently by an author, for readers. Read the story and support ChapterPal

keyword

blackbox attack

A blackbox attack is an adversarial attack on a machine learning model where the attacker has no access to the internal architecture, parameters, training data, or gradients of the target system. In this threat model, the adversary interacts with the system exclusively through its external interface, typically by providing inputs and observing the resulting outputs, such as class predictions or confidence scores. To manipulate the target model into making incorrect predictions, attackers commonly generate adversarial examples using query-based optimization methods or by training a surrogate model to craft adversarial perturbations that transfer effectively to the target system. This approach contrasts with whitebox attacks and closely reflects real-world threat scenarios where proprietary artificial intelligence systems are deployed behind closed interfaces.

1 item