Disarming Strategic Text: Span-Aware Counterfactuals for Robust Content Moderation

Hardik MeisheriMuhammad Zaid HassanSwati TiwariPuneet ManglaSamarth BharadwajKarthik SankaranarayananAmit Singh

article20250 citations

Presents a span-aware counterfactual data augmentation framework that isolates causal violation spans using multi-LLM consensus and rewrites them into policy-compliant hard negatives to defend content moderation classifiers against adversarial text manipulation.

Listen

Automated content moderation systems struggle to maintain accuracy when users employ strategic evasions, such as euphemisms, misspellings, and deliberate rephrasing. These tactics cause standard machine learning classifiers to rely on superficial patterns—like benign marketing words—rather than the actual violation, leading to missed violations and high false-positive rates. Traditional data augmentation techniques expand vocabulary diversity but fail to alter the core violating phrases, leaving classifiers vulnerable to adversarial manipulation.

The article demonstrates a span-aware counterfactual augmentation framework designed to identify the exact text triggering a policy violation and rewrite it into a compliant alternative. Its primary objective is to evaluate whether training classifiers with these high-quality, label-flipping hard negatives improves robustness and reduces reliance on misleading co-occurrence patterns across multiple content moderation tasks.

The authors implemented a four-phase pipeline using multiple large language models. The system first extracts violating text segments with strict consensus requirements across models, rewrites these segments into policy-compliant language while preserving sentence context, and validates that the label successfully flips to compliant. The pipeline was evaluated on an internal dataset of adult content ads (16,000 training samples), a public multi-label toxic comment dataset, and a diagnostic stress-test dataset designed to measure overfitting to benign trigger words. Downstream classifiers were tested across varying ratios of original to augmented data.

The findings show that this approach substantially boosts moderation performance and robustness. On the internal ad dataset, augmenting training data with 20% span-aware hard negatives increased the Area Under the Precision-Recall Curve by 6.3 percentage points (from 73.3% to 79.0%), whereas naive random text masking degraded performance down to 50.1%. On the diagnostic stress set, the model raised accuracy from 63.4% to 76.5%, proving it successfully learned to ignore benign trigger phrases. The multi-model agreement gate achieved a 61.2% exact span agreement rate, and policy-guided rewrites yielded a 94.3% label-flip success rate. Furthermore, the experiments revealed that the optimal augmentation fraction varies by domain, peaking at 20% for formulaic ad policies and around 5% for broader toxic comment classification.

These results demonstrate that precision and contextual realism are essential when augmenting data for safety-critical moderation systems. Superficial edits such as word deletions fail because they destroy sentence structure, confusing classifiers rather than training them. By systematically generating realistic, policy-compliant alternatives, organizations can reduce compliance risks, lower false alarm rates on harmless content, and improve model resilience without manually annotating vast new datasets.

Organizations deploying automated moderation should integrate targeted, span-aware augmentation into their training workflows, treating the augmentation ratio as a domain-specific tuning parameter to avoid oversaturating models with negative examples. Practitioners should also implement validation mechanisms to ensure generated text genuinely flips classification labels before fine-tuning downstream systems.

Confidence in these findings is supported by consistent gains across multiple benchmarks; however, limitations remain. The pipeline relies on commercial language model application programming interfaces, introducing potential cost and scalability bottlenecks. The framework also experiences occasional errors on complex inputs such as web addresses, emojis, and emerging manipulation tactics not captured by existing policy definitions. Future work should evaluate open-weight models and test resilience against broader distribution shifts.

Cover for Disarming Strategic Text: Span-Aware Counterfactuals for Robust Content Moderation

Abstract

Machine learning systems deployed in the wild must operate reliably despite unreliable inputs, whether arising from distribution shifts, adversarial manipulation, or strategic behavior by users. Content moderation is a prime example: violators deliberately exploit euphemisms, obfuscations, or benign co-occurrence patterns to evade detection, creating unreliable supervision signals for classifiers. We present a span-aware augmentation framework that generates high-quality counterfactual hard negatives to improve robustness under such conditions. Our pipeline combines (i) multi-LLM agreement to extract causal violation spans, (ii) policy-guided rewrites of those spans into compliant alternatives, and (iii) validation via re-inference to ensure only genuine label-flipping counterfactuals are retained. Across real-world ad moderation and toxic comment datasets, this approach consistently reduces spurious correlations and improves robustness to adversarial triggers, with PRAUC gains of up to +6.3 points. We further show that augmentation benefits peak at task-dependent ratios, underscoring the importance of balance in reliable learning. These findings highlight span-aware counterfactual augmentation as a practical path toward reliable ML from strategically manipulated and unreliable text data.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Problem Formulation
  • 4 Method: Span-Aware Counterfactual Augmentation
  • 4.1 Algorithm
  • 5 Experimental Setup
  • 6 Experimental Results
  • 6.1 Error Analysis
  • 6.2 Discussion
  • 7 Limitations and Future Work
  • 7.1 Discussion
  • 8 Conclusion
  • References

Knowls

  1. Knowl 1 — Span-Aware Agreement-Filtered Counterfactual Augmentation Pipeline

    model/method

    The span-aware counterfactual augmentation framework is a multi-stage pipeline designed to generate boundary-challenging, label-flipping hard negatives for text moderation classifiers:

    1. High-Precision Span Identification with Consensus Gating: Each candidate text xx is evaluated by multiple Large Language Models (LLMs) running in parallel. Each LLM predicts whether xx violates a policy and extracts the character-level span SS responsible for the violation. An agreement gate retains only examples that achieve unanimous or majority consensus on both the violation label and exact character-level span boundaries across models.

    2. Policy-Guided Span Replacement: For retained violating examples, the identified span SS is treated as the causal anchor, leaving the remaining non-causal text scaffolding intact. An LLM is prompted with the policy rules π\pi to replace SS with a compliant, contextually fluent phrase, yielding candidate rewrite x′x'.

    3. Re-Inference Validation: Candidate rewrites x′x' are passed back to the LLM ensemble for secondary inference. Only candidates that achieve consensus classification as Compliant\text{Compliant} are accepted into the hard-negative set A\mathcal{A}, filtering out malformed or borderline rewrites.

    4. Augmentation Integration: The validated hard negatives A\mathcal{A} are mixed into classifier fine-tuning mini-batches at a controlled augmentation ratio α\alpha, exposing the classifier to minimally different positive-negative pairs.

  2. Knowl 2 — Causal Spans and Minimal Causal Sets in Content Moderation

    definition

    Let a dataset be denoted as D={(xi,yi)}i=1N\mathcal{D} = \{(x_i, y_i)\}_{i=1}^N, where xix_i represents an input text sequence and yi∈Yy_i \in \mathcal{Y} represents its categorical policy label (with yi≠Complianty_i \neq \text{Compliant} denoting a policy violation).

    • Causal Span: A substring s=xi[a:b]s = x_i[a:b] is defined as causal for violation yiy_i if replacing ss with policy-compliant text is sufficient to change the classification label of xix_i to Compliant\text{Compliant}.

    • Minimal Causal Set: A set of non-overlapping substrings Si={s1,s2,…,sk}S_i = \{s_1, s_2, \dots, s_k\} of xix_i is a minimal causal set if replacing all substrings sj∈Sis_j \in S_i simultaneously flips the label of xix_i to Compliant\text{Compliant}, whereas replacing any proper subset of SiS_i leaves the label as violating.

  3. Knowl 3 — Span-Aware Agreement-Filtered Counterfactual Generation Algorithm

    algorithm

    The counterfactual generation procedure processes a dataset D\mathcal{D} of labeled texts using an ensemble of LLMs M\mathcal{M}, policy guidelines π\pi, and replacement operators R\mathcal{R} to produce a validated set of hard negatives A\mathcal{A}:

    Input: Dataset DD, LLM set MM, policy rules π\pi, replacement methods RR
    Output: Augmented hard negative dataset AA
    A←∅A \leftarrow \emptyset
    for each text x∈Dx \in D do
        Query all m∈Mm \in M for predicted violation label y^(m)\hat{y}^{(m)} and causal span S^(m)\hat{S}^{(m)}
        if label and span agreement across all m∈Mm \in M then
            Mask span S^\hat{S} in xx to obtain xmaskx^{\text{mask}}
            for each method ∈R\in R do
                x′←replace(xmask,S^,π,method)x' \leftarrow \text{replace}(x^{\text{mask}}, \hat{S}, \pi, \text{method})
                y^′←majority({infer(x′,m)}m∈M)\hat{y}' \leftarrow \text{majority}(\{\text{infer}(x', m)\}_{m \in M})
                if y^′=Compliant\hat{y}' = \text{Compliant} then
                    A←A∪{x′}A \leftarrow A \cup \{x'\}
                end if
            end for
        end if
    end for
    return AA

    Each input sample is checked for both label consensus and exact character-level span agreement across LLMs m∈Mm \in \mathcal{M}. If agreement holds, candidate rewrites x′x' are constructed by minimally substituting the causal span. Candidates undergo re-inference validation, and only those whose majority LLM classification flips to Compliant\text{Compliant} are added to A\mathcal{A}.

  4. Knowl 4 — Augmentation-Aware Batch Construction for Classifier Training

    algorithm

    To inject synthesized counterfactual hard negatives without skewing the empirical class balance of the original training distribution, mini-batches of size bb are constructed using a fixed augmentation ratio α∈[0,1]\alpha \in [0, 1]:

    Input: Original dataset DD, augmented counterfactual set AA, augmentation ratio α\alpha, batch size bb
    Output: Training mini-batch BB
    for each mini-batch do
        naug←⌊b×α⌋n_{\text{aug}} \leftarrow \lfloor b \times \alpha \rfloor
        norig←b−naugn_{\text{orig}} \leftarrow b - n_{\text{aug}}
        Sample norign_{\text{orig}} examples from DD preserving original class ratios
        Sample naugn_{\text{aug}} examples from AA uniformly
        B←shuffle(Dorig∪Aaug)B \leftarrow \text{shuffle}(D_{\text{orig}} \cup A_{\text{aug}})
    end for
    return BB

    This construction exposes the downstream classifier to both original texts and their minimally modified compliant counterparts in the same optimization step, encouraging the model to penalize reliance on benign scaffolding tokens.

  5. Knowl 5 — Experimental Setup for Span-Aware Counterfactual Moderation

    experimental setup

    The empirical evaluation assesses robustness, generalization, and classification quality across three datasets and a baseline classifier:

    • Internal Adult Content Dataset: Real-world ad texts governed by adult content policy violations, split into 16,000 training (8,000 compliant, 8,000 violating), 4,000 validation (2,000 compliant, 2,000 violating), and 3,064 gold test samples (2,743 compliant, 321 violating).
    • Kaggle Toxic Comment Dataset: A multi-label benchmark covering six categories (toxic, severe toxic, obscene, threat, insult, identity hate). Subsampled to 8,000 training (7,208 compliant, 792 violating) and 2,000 validation (1,803 compliant, 197 violating) to simulate low-resource settings, evaluated against the full test set of 63,978 samples (57,735 compliant, 6,243 violating).
    • Spy-Camera Diagnostic Stress Set: 400 test instances (250 compliant, 150 violating) containing benign trigger keywords (such as "spy camera") in non-violating contexts to quantify susceptibility to spurious correlations.
    • Models & Hyperparameters: GPT-4o and DeepSeekV3 serve as the LLM ensemble for span identification, rewriting, and validation (averaging 0.3s per extraction, 0.5s per rewrite on GPT-4o API, ~4 hours per 10K samples). Downstream classification is performed with bert-base-uncased (110M parameters) fine-tuned using AdamW with learning rate 2×10−42\times 10^{-4}, batch size 128, maximum 15 epochs with early stopping on validation performance, averaged over 5 random seeds.
    • Primary Metrics: Area Under the Precision-Recall Curve (PRAUC) and Macro-PRAUC across classes.
  6. Knowl 6 — Downstream Classification Gains and Robustness via Span-Aware Augmentation

    empirical result

    Fine-tuning bert-base-uncased on data augmented with span-aware counterfactuals substantially improves Precision-Recall AUC (PRAUC) on ad moderation test sets and diagnostic stress sets, while random span masking degrades performance:

    Method 0% 5% 10% 15%
    Span-aware aug. 73.3% 75.8% 78.2% 79.0%
    Random span mask. 73.3% 65.9% 58.5% 50.1%

    On the internal adult content test set, span-aware augmentation raises PRAUC from a baseline of 73.3% (0% augmentation) to a peak of ~79.3% at α=20%\alpha = 20\% (+6.3 percentage point gain). Beyond α=20%\alpha = 20\%, performance slightly decreases due to oversaturation of compliant negatives diluting sparse violation signals. On the "Spy-Camera" stress set, span-aware augmentation improves PRAUC from 0.634 (baseline) to 0.765 at α=20%\alpha = 20\%, confirming mitigation of spurious trigger-word associations.

  7. Knowl 7 — Span Extraction Agreement and Rewrite Flip Rates

    empirical result

    Evaluation of multi-LLM span identification and replacement mechanisms on violating ad texts demonstrates the necessity of policy-guided rewriting over heuristic replacements:

    • Agreement Gate Quality: Comparing GPT-4o and DeepSeekV3 yields a 90.3% label agreement and a 61.2% exact character-level joint span agreement, acting as a high-precision filter for clean seed data.

    • Label-Flip Success Rates:

    Replacement Method Flip %
    LLM-based (policy-guided) 94.3%
    BERT token replacement 90.7%
    Empty-string removal 90.4%
    Keep [MASK] 88.1%

    Although simple baseline strategies (BERT token filling, empty-string removal, retaining [MASK]) achieve 88–91% flip rates by stripping out violating tokens, policy-guided LLM rewriting achieves the highest flip rate (94.3%) while generating contextually coherent counterfactuals.

  8. Knowl 8 — Task-Dependent Optimal Augmentation Ratios

    empirical result

    The optimal augmentation ratio α\alpha for injecting counterfactual hard negatives depends on the semantic complexity and diversity of the moderation domain:

    • On the Internal Adult Content Dataset, where ad violations follow formulaic phrasing patterns, performance improvements scale with larger augmentation fractions, peaking at α=20%\alpha = 20\% (79.3% PRAUC vs 73.3% baseline).
    • On the Kaggle Toxic Comment Dataset, which spans six diverse multi-label toxicity categories with complex linguistic phrasing, macro-PRAUC peaks at a much smaller augmentation fraction of α≈5%\alpha \approx 5\%.

    In broad and linguistically diverse domains, small injections of hard negatives refine decision boundaries without overwhelming the classifier's representations of varied positive vocabularies, whereas formulaic violations tolerate higher hard-negative ratios before signal dilution occurs.

  9. Knowl 9 — Downstream Degradation from Non-Contextual Span Replacements

    empirical result

    In downstream ablation experiments comparing LLM-guided contextual rewrites against heuristic replacements (BERT token replacement and empty-string removal), both heuristic methods severely degrade downstream PRAUC across all augmentation fractions α\alpha (e.g., dropping from the 73.3% baseline to below 60.0% at α≥5%\alpha \ge 5\%).

    This demonstrates that achieving a high label-flip rate during generation is insufficient on its own; hard negatives must preserve sentence structure, context, and fluency so the classifier learns fine-grained boundary discrimination rather than artifactual, ungrammatical patterns.

  10. Knowl 10 — Limitations of LLM-Guided Span-Aware Counterfactual Moderation

    limitation

    The span-aware counterfactual pipeline has several structural and operational limitations:

    1. Extraction and Rewriting Failures: The system struggles with complex multi-span violations, obfuscated spelling variations, embedded URLs, and emoji-based violations, which can escape exact span consensus.
    2. Distributional Bias to Known Policies: Because rewriting is conditioned on defined policy rules π\pi, emerging zero-day adversarial tactics (e.g., novel scam vectors or misinformation patterns) may remain underrepresented in the generated counterfactuals.
    3. Computational Overhead and Dependency: Reliance on proprietary LLM APIs introduces API latency (approximately 0.3s per span extraction and 0.5s per rewrite) and high operational cost (~4 hours processing time per 10,000 samples), constraining real-time deployment.

Coverage note — None was omitted; all key theoretical definitions, pipeline algorithms, empirical results on both datasets and stress tests, replacement ablations, and stated limitations were fully captured.

References

  1. 1.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding, 2019. URL https://arxiv.org/abs/1810.04805.
  2. 2.Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. Eraser: A benchmark to evaluate rationalized nlp models, 2020. URL https://arxiv.org/abs/1911.03429.
  3. 3.Sergey Edunov, Myle Ott, Michael Auli, and David Grangier. Understanding back-translation at scale. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii, editors, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 489–500, Brussels, Belgium, October-November 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1045. URL https://aclanthology.org/D18-1045/.
  4. 4.Tianyu Gao, Xingcheng Yao, and Danqi Chen. Simcse: Simple contrastive learning of sentence embeddings, 2022. URL https://arxiv.org/abs/2104.08821.
  5. 5.Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou. Evaluating models’ local decision boundaries via contrast sets. In Trevor Cohn, Yulan He, and Yang Liu, editors, Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1307–1323, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.117. URL https://aclanthology.org/2020.findings-emnlp.117/.
  6. 6.Sarthak Jain and Byron C. Wallace. Attention is not explanation, 2019. URL https://arxiv.org/abs/1902.10186.
  7. 7.Harsh Jhamtani and Peter Clark. Learning to explain: Datasets and models for identifying valid reasoning chains in multihop question-answering. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 137–150, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.10. URL https://aclanthology.org/2020.emnlp-main.10/.
  8. 8.Vladimir Karpukhin et al. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, 2020. URL https://aclanthology.org/2020.emnlp-main.550.pdf.
  9. 9.Divyansh Kaushik, Amrith Setlur, Eduard Hovy, and Zachary C. Lipton. Explaining the efficacy of counterfactually augmented data, 2021. URL https://arxiv.org/abs/2010.02114.
  10. 10.Mike Lewis et al. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 7871–7880, 2019.
  11. 11.Xinyu Li et al. Adversarial text generation by search and learning. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 14841–14854, 2023. URL https://aclanthology.org/2023.findings-emnlp.1053/.
  12. 12.Hansa Meghwani, Amit Agarwal, Priyaranjan Pattnayak, Hitesh Laxmichand Patel, and Srikant Panda. Hard negative mining for domain-specific retrieval in enterprise systems, 2025. URL https://arxiv.org/abs/2505.18366.
  13. 13.Joshua Robinson, Ching-Yao Chuang, Suvrit Sra, and Stefanie Jegelka. Contrastive learning with hard negative samples, 2021. URL https://arxiv.org/abs/2010.04592.
  14. 14.Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. In Katrin Erk and Noah A. Smith, editors, Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany, August 2016. Association for Computational Linguistics. doi: 10.18653/v1/P16-1162. URL https://aclanthology.org/P16-1162/.
  15. 15.Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld. Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6707–6723, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.523. URL https://aclanthology.org/2021.acl-long.523/.
  16. 16.Lee Xiong et al. Approximate nearest neighbor negative contrastive estimation for dense text retrieval. In International Conference on Learning Representations (ICLR), 2021.
  17. 17.Yichong Xu et al. Peerda: Data augmentation via modeling peer relation for span identification tasks. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), 2023.
  18. 18.Yiming Yang et al. Enhanced language representation with label knowledge for span extraction. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4822–4834, 2021. URL https://aclanthology.org/2021.emnlp-main.379/.
  19. 19.Yixing Zhou et al. Simans: Simple ambiguous negatives sampling for dense text retrieval. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 443–453, 2022. URL https://aclanthology.org/2022.emnlp-industry.56.pdf.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF