keyword
span-aware counterfactuals
Span-aware counterfactuals are modified text samples created by identifying and editing specific substrings or spans responsible for a model prediction, altering those exact segments to flip the target label while keeping the surrounding context intact. In natural language processing and machine learning, this approach focuses perturbations strictly on the specific tokens that causally drive a classification decision, such as policy violations, toxic terms, or key sentiment indicators, and rewrites them into compliant or alternative forms. By localizing edits to relevant text spans rather than altering entire documents uniformly, span-aware counterfactual generation creates targeted training examples that help models isolate genuine causal features from spurious correlations, thereby enhancing robustness against adversarial manipulation and strategic evasion.
1 item

