keyword
meaningful perturbation
Meaningful perturbation is an interpretability technique in machine learning used to explain decisions made by complex models by applying controlled, visually or semantically coherent modifications to an input and observing the resulting change in the model output. Rather than inspecting the internal parameters or gradients of a network, this approach systematically alters specific regions or features—such as blurring, occluding, or masking portions of an image—to determine which areas cause the largest drop in prediction confidence. By finding the smallest or most influential modification required to change the decision while avoiding artificial artifacts, meaningful perturbation generates human-interpretable attribution maps that pinpoint the exact parts of an input most responsible for the model classification.
1 item

