Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment

Di JinZhijing JinJoey Tianyi ZhouPeter Szolovits

article2019AAAI1,419 citations

Proposes TextFooler, a simple and computationally efficient attack framework that systematically deceives pre-trained models like BERT by generating natural, semantics-preserving adversarial examples for classification and textual entailment.

Listen

Modern natural language processing systems are widely deployed in critical applications, yet they remain vulnerable to adversarial manipulation—small, carefully crafted input modifications that deceive algorithms while remaining unnoticed by humans. Generating effective adversarial text has historically been difficult because simple modifications like character typos or word removals degrade fluency and alter underlying meaning. The article evaluates the true robustness of leading natural language models by demonstrating TEXTFOOLER, an automated framework designed to generate natural, meaning-preserving adversarial attacks in realistic operational settings.

The investigation operates under a black-box framework, meaning the attack system requires no access to model parameters, code, or internal gradients; it solely observes output predictions and confidence scores. The method first identifies the most influential words driving a model's decision by measuring prediction score changes when words are removed. It then methodically substitutes these key words with synonyms that preserve grammatical category and sentence-level semantic meaning until the target model produces an incorrect prediction. The article tests this attack across five classification datasets—such as fake news detection and sentiment analysis—and two textual inference datasets, targeting leading architectures including Convolutional Neural Networks, Long Short-Term Memory networks, and Bidirectional Encoder Representations from Transformers (BERT).

The evaluation reveals that even industry-standard models are exceptionally fragile under targeted semantic manipulation. Across almost all tested tasks, the attack reduces model accuracy from benchmark rates above 80%–95% to below 15%, while altering fewer than 20% of the words in a given text. Even the advanced BERT architecture, widely considered more robust, experiences accuracy drops of roughly 5 to 7 times in classification and 9 to 22 times in textual inference. Human evaluations confirm that these modified texts remain grammatically sound and preserve original intent, with human judges maintaining an 85% to 92% classification agreement with the original text labels. Furthermore, the generation process operates efficiently, requiring target model queries that scale linearly with text length.

These findings demonstrate that high benchmark performance does not equate to operational security or genuine linguistic understanding. Organizations relying on standard natural language models face significant vulnerability risks in fraud detection, content moderation, and automated analysis, as adversaries can evade detection through subtle, fluent paraphrasing. Notably, the study finds that adversarial text exhibits transferability across different model architectures and that adversarial samples crafted against higher-performing models like BERT transfer more effectively to others.

To mitigate these risks, organizations should incorporate adversarial text generation directly into their model development workflows. The article demonstrates that retraining models on a mix of standard and adversarially generated examples measurably improves resilience against future attacks. Developers should conduct adversarial stress testing prior to deployment rather than relying solely on standard test accuracy.

Confidence in these findings is supported by rigorous automated benchmarks and multi-rater human evaluations across multiple standard datasets. However, certain limitations remain: real-to-fake news conversion proved significantly more difficult to manipulate than other classification domains, and shorter text inputs present stricter constraints on semantic similarity. Decision-makers should evaluate model robustness within their specific domain constraints when deploying natural language systems.

Cover for Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment

Abstract

Machine learning algorithms are often vulnerable to adversarial examples that have imperceptible alterations from the original counterparts but can fool the state-of-the-art models. It is helpful to evaluate or even improve the robustness of these models by exposing the maliciously crafted adversarial examples. In this paper, we present TextFooler, a simple but strong baseline to generate natural adversarial text. By applying it to two fundamental natural language tasks, text classification and textual entailment, we successfully attacked three target models, including the powerful pre-trained BERT, and the widely used convolutional and recurrent neural networks. We demonstrate the advantages of this framework in three ways: (1) effective---it outperforms state-of-the-art attacks in terms of success rate and perturbation rate, (2) utility-preserving---it preserves semantic content and grammaticality, and remains correctly classified by humans, and (3) efficient---it generates adversarial text with computational complexity linear to the text length. *The code, pre-trained target models, and test examples are available at this https URL.

Table of Contents

  • Introduction
  • Method
  • Problem Formulation
  • Threat Model
  • Step 1: Word Importance Ranking (line 1-6)
  • Step 2: Word Transformer (line 7-30)
  • Experiments
  • Tasks
  • Text Classification
  • Textual Entailment
  • Attacking Target Models
  • Setup of Automatic Evaluation
  • Setup of Human Evaluation
  • Results
  • Automatic Evaluation
  • Human Evaluation
  • Discussion
  • Ablation Study
  • Word Importance Ranking
  • Semantic Similarity Constraint
  • Transferability
  • Adversarial Training
  • Error Analysis
  • Related Work
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — TextFooler Black-Box Adversarial Attack Algorithm

    algorithm

    TextFooler generates adversarial text examples under a black-box setting where the attacker only has access to model prediction scores. The algorithm identifies the most influential words in an input sentence, finds candidate synonyms using counter-fitted embeddings, filters them by part-of-speech (POS) tags and sentence-level semantic similarity, and greedily selects perturbations that flip the model prediction or minimize the confidence score of the true label.

    Input: Sentence X={w1,w2,…,wn}X = \{w_1, w_2, \dots, w_n\}, ground truth label YY, target model FF, sentence similarity function SimSim, sentence similarity threshold ϵ\epsilon, word embeddings EmbEmb over vocabulary VocabVocab
    Output: Adversarial example XadvX_{adv}
    Xadv←XX_{adv} \leftarrow X
    for each word wi∈Xw_i \in X do
        Compute importance score IwiI_{w_i}
    end for
    W←W \leftarrow all words wi∈Xw_i \in X sorted in descending order of IwiI_{w_i}
    Filter out stop words from WW
    for each word wj∈Ww_j \in W do
        Candidates←Candidates \leftarrow top NN synonyms of wjw_j by cosine similarity in EmbEmb with similarity >δ> \delta
        Candidates←POSFilter(Candidates,wj)Candidates \leftarrow POSFilter(Candidates, w_j)
        FinCandidates←{}FinCandidates \leftarrow \{\}
        for each candidate ck∈Candidatesc_k \in Candidates do
            X′←X' \leftarrow replace wjw_j with ckc_k in XadvX_{adv}
            if Sim(X′,Xadv)>ϵSim(X', X_{adv}) > \epsilon then
                Add ckc_k to FinCandidatesFinCandidates
                Yk←F(X′)Y_k \leftarrow F(X')
                Pk←FYk(X′)P_k \leftarrow F_{Y_k}(X')
            end if
        end for
        if there exists ck∈FinCandidatesc_k \in FinCandidates with prediction Yk≠YY_k \neq Y then
            c∗←argmax⁡c∈FinCandidates,Yc≠YSim(X,Xwj→c′)c^* \leftarrow \operatorname{argmax}_{c \in FinCandidates, Y_c \neq Y} Sim(X, X'_{w_j \to c})
            Xadv←X_{adv} \leftarrow replace wjw_j with c∗c^* in XadvX_{adv}
            return XadvX_{adv}
        else if FY(Xadv)>min⁡ck∈FinCandidatesPkF_Y(X_{adv}) > \min_{c_k \in FinCandidates} P_k then
            c∗←argmin⁡ck∈FinCandidatesPkc^* \leftarrow \operatorname{argmin}_{c_k \in FinCandidates} P_k
            Xadv←X_{adv} \leftarrow replace wjw_j with c∗c^* in XadvX_{adv}
        end if
    end for
    return None

    The algorithm operates with query complexity roughly linear in sentence length nn, executing between 2n2n and 8n8n queries per sentence on average.

  2. Knowl 2 — Black-Box Word Importance Scoring Function

    equation

    To select which words to perturb in a black-box setting without gradient access, the importance score IwiI_{w_i} measures the change in prediction confidence for an input sequence X={w1,w2,…,wn}X = \{w_1, w_2, \dots, w_n\} with label Y=F(X)Y = F(X) upon removing word wiw_i. Let X∖{wi}={w1,…,wi−1,wi+1,…,wn}X \setminus \{w_i\} = \{w_1, \dots, w_{i-1}, w_{i+1}, \dots, w_n\} denote the sentence after deleting wiw_i, and let Fy(⋅)F_y(\cdot) denote the prediction probability for class label y∈Yy \in \mathcal{Y}. The score is defined as:

    Iwi={FY(X)−FY(X∖{wi}),if F(X)=F(X∖{wi})=Y(FY(X)−FY(X∖{wi}))+(FYˉ(X∖{wi})−FYˉ(X)),if F(X)=Y,F(X∖{wi})=Yˉ, and Y≠YˉI_{w_i} = \begin{cases} F_Y(X) - F_Y(X \setminus \{w_i\}), & \text{if } F(X) = F(X \setminus \{w_i\}) = Y \\ (F_Y(X) - F_Y(X \setminus \{w_i\})) + (F_{\bar{Y}}(X \setminus \{w_i\}) - F_{\bar{Y}}(X)), & \text{if } F(X) = Y, F(X \setminus \{w_i\}) = \bar{Y}, \text{ and } Y \neq \bar{Y} \end{cases}

    If removing wiw_i does not alter the predicted class, IwiI_{w_i} is the drop in target class confidence FYF_Y. If removing wiw_i changes the prediction to a different class ar{Y}, IwiI_{w_i} combines the decrease in target class probability and the increase in the predicted competing class probability. Words identified as stop words via NLTK and spaCy are filtered out following scoring to maintain grammaticality.

  3. Knowl 3 — Candidate Synonym Generation and Filtering Pipeline

    model/method

    For each selected target word wiw_i, replacement candidates are generated and filtered through a three-stage pipeline:

    1. Synonym Extraction: A candidate pool is initialized with the top N=50N = 50 nearest words in cosine similarity from counter-fitted word embeddings, provided the cosine similarity with wiw_i is strictly greater than δ=0.7\delta = 0.7.
    2. Part-of-Speech (POS) Checking: An off-the-shelf POS tagger (spaCy) is applied to keep only candidate words that share the same POS tag as wiw_i in context, preserving sentence syntax.
    3. Sentence-Level Semantic Similarity Constraint: The original sentence XX and the candidate adversarial sentence XadvX_{adv} (with candidate cc replacing wiw_i) are embedded using the Universal Sentence Encoder (USE). The candidate is retained in the final candidate set only if the cosine similarity between the two sentence embeddings exceeds a predefined semantic threshold ϵ\epsilon.
  4. Knowl 4 — Attack Effectiveness on Text Classification and Textual Entailment

    data/table

    TextFooler was evaluated across five text classification tasks (MR, IMDB, Yelp Polarity, AG's News, Fake News) and two textual entailment tasks (SNLI, MultiNLI matched/mismatched) against WordCNN, WordLSTM, and BERT (base-uncased) on 1,000 randomly sampled test examples.

    Target Model MR IMDB Yelp AG Fake SNLI MNLI (m/mm)
    WordCNN / InferSent
    Original Accuracy (%) 78.0 89.2 93.8 91.5 96.7 84.3 70.9 / 69.6
    After-Attack Accuracy (%) 2.8 0.0 1.1 1.5 15.9 3.5 6.7 / 6.9
    % Perturbed Words 14.3 3.5 8.3 15.2 11.0 18.0 13.8 / 14.6
    Semantic Similarity 0.68 0.89 0.82 0.76 0.82 0.50 0.61 / 0.59
    WordLSTM / ESIM
    Original Accuracy (%) 80.7 89.8 96.0 91.3 94.0 86.5 77.6 / 75.8
    After-Attack Accuracy (%) 3.1 0.3 2.1 3.8 16.4 5.1 7.7 / 7.3
    % Perturbed Words 14.9 5.1 10.6 18.6 10.1 18.1 14.5 / 14.6
    Semantic Similarity 0.67 0.87 0.79 0.63 0.80 0.47 0.59 / 0.59
    BERT
    Original Accuracy (%) 86.0 90.9 97.0 94.2 97.8 89.4 85.1 / 82.1
    After-Attack Accuracy (%) 11.5 13.6 6.6 12.5 19.3 4.0 9.6 / 8.3
    % Perturbed Words 16.7 6.1 13.9 22.0 11.7 18.5 15.2 / 14.6
    Semantic Similarity 0.65 0.86 0.74 0.57 0.76 0.45 0.57 / 0.58

    TextFooler reduces the accuracy of target models from state-of-the-art levels to under 15% in nearly all classification and entailment tasks (except Fake News) while modifying under 20% of original tokens. Models with higher baseline accuracy (BERT) require slightly higher perturbation rates, and linear regression reveals a strong negative correlation between perturbed word percentage and Universal Sentence Encoder semantic similarity (R2=0.94R^2 = 0.94 for classification, R2=0.97R^2 = 0.97 for entailment).

  5. Knowl 5 — Comparison with Existing Adversarial Text Baselines

    data/table

    TextFooler was benchmarked against existing adversarial text attack systems under identical model and dataset setups: LSTM on IMDB, InferSent on SNLI, and LSTM on Yelp.

    Dataset Model Success Rate (%) % Perturbed Words
    IMDB Li et al. (2018) 86.7 6.9
    Alzantot et al. (2018) 97.0 14.7
    TextFooler (Ours) 99.7 5.1
    SNLI Alzantot et al. (2018) 70.0 23.0
    TextFooler (Ours) 95.8 18.0
    Yelp Kuleshov et al. (2018) 74.8 –
    TextFooler (Ours) 97.8 10.6

    TextFooler achieves higher attack success rates while perturbing fewer words across all three evaluated datasets, reaching a 99.7% success rate on IMDB with only 5.1% of words altered.

  6. Knowl 6 — Human Evaluation of Adversarial Text Quality

    empirical result

    To verify that adversarial examples preserve semantic integrity and linguistic fluency, human evaluation was conducted on 100 randomly sampled test instances from MR (targeting WordLSTM) and 100 from SNLI (targeting BERT) with two independent native English speakers:

    1. Grammaticality: Evaluators scored original and adversarial sentences in a shuffled mix on a Likert scale from 1 to 5. Grammaticality was closely preserved: MR original scored 4.22 vs. adversarial 4.01; SNLI original scored 4.50 vs. adversarial 4.27.
    2. Human Prediction Consistency: Human classification agreement between the original and adversarial text was 92% on MR and 85% on SNLI, confirming that the ground-truth classification labels perceived by humans remain consistent.
    3. Semantic Similarity: Judges rated whether adversarial sentences were similar (1.0), ambiguous (0.5), or dissimilar (0.0) to original counterparts, achieving mean scores of 0.91 on MR and 0.86 on SNLI.
  7. Knowl 7 — Ablation of Word Importance Ranking

    data/table

    To evaluate the contribution of the word importance ranking mechanism (Step 1 of Algorithm 1), it was compared against a baseline that randomly selects words to perturb while keeping the perturbation budget and synonym transformation step identical. The target model was BERT-base.

    Metric / Dataset MR AG SNLI
    % Perturbed Words 16.7 22.0 18.5
    Original Accuracy (%) 86.0 94.2 89.4
    After-Attack Accuracy (%) 11.5 12.5 4.0
    After-Attack Accuracy (Random) (%) 68.3 80.8 59.2

    Removing the word importance ranking step causes the after-attack accuracy to rise by over 45 to 68 percentage points across all three datasets, demonstrating that targeted word ranking is necessary for high attack success under low perturbation budgets.

  8. Knowl 8 — Ablation of Sentence-Level Semantic Similarity Constraint

    data/table

    The role of the sentence-level semantic similarity threshold ϵ\epsilon (computed via Universal Sentence Encoder) was assessed by evaluating TextFooler with and without this constraint against BERT-base across four datasets.

    Metric (With / Without) MR IMDB SNLI MNLI(m)
    After-Attack Accuracy (%) 11.5 / 6.2 13.6 / 11.2 4.0 / 3.6 9.6 / 7.9
    % Perturbed Words 16.7 / 14.8 6.1 / 4.0 18.5 / 18.3 15.2 / 14.5
    Query Number 166 / 131 1134 / 884 60 / 57 78 / 70
    Semantic Similarity 0.65 / 0.58 0.86 / 0.82 0.45 / 0.44 0.57 / 0.56

    Omitting the semantic similarity constraint reduces query counts and slightly decreases after-attack accuracy, but causes noticeable degradation in semantic similarity scores and introduces contextually improper substitutions.

  9. Knowl 9 — Transferability of TextFooler Adversarial Examples

    data/table

    Transferability was evaluated by generating adversarial examples on one target model and evaluating them on the remaining two architectures on the IMDB and SNLI test sets.

    IMDB SNLI
    Source \ arget WordCNN WordLSTM BERT InferSent ESIM BERT
    WordCNN / InferSent 0.0 84.9 90.2 0.0 62.7 67.7
    WordLSTM / ESIM 74.9 0.0 87.9 49.4 0.0 59.3
    BERT 84.1 85.1 0.0 58.2 54.6 0.0

    Adversarial examples transfer moderately across model families. Transferability is higher in textual entailment than sentiment classification. Furthermore, adversarial examples generated from the higher-capacity BERT model exhibit greater transferability to other models than vice versa (e.g., reducing ESIM accuracy on SNLI to 54.6% and InferSent to 58.2%).

  10. Knowl 10 — Adversarial Training with TextFooler Examples

    data/table

    To test whether TextFooler examples can improve model robustness, adversarial examples generated from the MR and SNLI training sets were added to the training data, and BERT models were retrained from scratch before undergoing new TextFooler attacks.

    MR SNLI
    Training Regime After-Attack Acc. (%) % Perturbed After-Attack Acc. (%) % Perturbed
    Original Training 11.5 16.7 4.0 18.5
    + Adversarial Training 18.7 21.0 8.3 20.1

    Adversarially trained BERT models show improved robustness against future attacks, exhibiting both higher after-attack accuracy (from 11.5% to 18.7% on MR; from 4.0% to 8.3% on SNLI) and requiring a larger fraction of words to be perturbed to succeed.

Coverage note — None was omitted; all key algorithmic contributions, mathematical formulations, evaluation tables, human studies, ablation analyses, transferability measurements, and adversarial training results have been captured.

References

  1. 1.Alzantot, M.; Sharma, Y.; Elgohary, A.; Ho, B.-J.; Srivastava, M.; and Chang, K.-W. 2018. Generating natural language adversarial examples. arXiv preprint arXiv:1804.07998.
  2. 2.Bowman, S. R.; Angeli, G.; Potts, C.; and Manning, C. D. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics.
  3. 3.Carlini, N., and Wagner, D. 2018. Audio adversarial examples: Targeted attacks on speech-to-text. arXiv preprint arXiv:1801.01944.
  4. 4.Cer, D.; Yang, Y.; Kong, S.-y.; Hua, N.; Limtiaco, N.; John, R. S.; Constant, N.; Guajardo-Cespedes, M.; Yuan, S.; Tar, C.; et al. 2018. Universal sentence encoder. arXiv preprint arXiv:1803.11175.
  5. 5.Chen, Q.; Zhu, X.; Ling, Z.; Wei, S.; Jiang, H.; and Inkpen, D. 2016. Enhanced lstm for natural language inference. arXiv preprint arXiv:1609.06038.
  6. 6.Conneau, A.; Kiela, D.; Schwenk, H.; Barrault, L.; and Bordes, A. 2017. Supervised learning of universal sentence representations from natural language inference data. arXiv preprint arXiv:1705.02364.
  7. 7.Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  8. 8.Feng, S.; Wallace, E.; Grissom, I.; Iyyer, M.; Rodriguez, P.; Boyd-Graber, J.; et al. 2018. Pathologies of neural models make interpretations difficult. arXiv preprint arXiv:1804.07781.
  9. 9.Gagnon-Marchand, J.; Sadeghi, H.; Haidar, M.; Rezagholizadeh, M.; et al. 2018. Salsa-text: self attentive latent space based adversarial text generation. arXiv preprint arXiv:1809.11155.
  10. 10.Gao, J.; Lanchantin, J.; Soffa, M. L.; and Qi, Y. 2018. Blackbox generation of adversarial text sequences to evade deep learning classifiers. arXiv preprint arXiv:1801.04354.
  11. 11.Goodfellow, I. J.; Shlens, J.; and Szegedy, C. 2014. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572.
  12. 12.Goodfellow, I.; Shlens, J.; and Szegedy, C. 2015. Explaining and harnessing adversarial examples. In International Conference on Learning Representations.
  13. 13.Hill, F.; Reichart, R.; and Korhonen, A. 2015. Simlex-999: Evaluating semantic models with (genuine) similarity estimation. Computational Linguistics 41(4):665–695.
  14. 14.Hochreiter, S., and Schmidhuber, J. 1997. Long short-term memory. Neural computation 9(8):1735–1780.
  15. 15.Kim, Y. 2014. Convolutional neural networks for sentence classification. arXiv preprint arXiv:1408.5882.
  16. 16.Kuleshov, V.; Thakoor, S.; Lau, T.; and Ermon, S. 2018. Adversarial examples for natural language classification problems.
  17. 17.Kurakin, A.; Goodfellow, I.; and Bengio, S. 2016a. Adversarial examples in the physical world. arXiv preprint arXiv:1607.02533.
  18. 18.Kurakin, A.; Goodfellow, I. J.; and Bengio, S. 2016b. Adversarial machine learning at scale. CoRR abs/1611.01236.
  19. 19.Li, J.; Ji, S.; Du, T.; Li, B.; and Wang, T. 2018. Textbugger: Generating adversarial text against real-world applications. arXiv preprint arXiv:1812.05271.
  20. 20.Li, J.; Monroe, W.; and Jurafsky, D. 2016. Understanding neural networks through representation erasure. arXiv preprint arXiv:1612.08220.
  21. 21.Liang, B.; Li, H.; Su, M.; Bian, P.; Li, X.; and Shi, W. 2017. Deep text classification can be fooled. arXiv preprint arXiv:1704.08006.
  22. 22.Moosavi-Dezfooli, S.-M.; Fawzi, A.; Fawzi, O.; and Frossard, P. 2017. Universal adversarial perturbations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1765–1773.
  23. 23.Mrkˇsić, N.; Séaghdha, D. O.; Thomson, B.; Gašić, M.; Rojas-Barahona, L.; Su, P.-H.; Vandyke, D.; Wen, T.-H.; and Young, S. 2016. Counter-fitting word vectors to linguistic constraints. arXiv preprint arXiv:1603.00892.
  24. 24.Niven, T., and Kao, H.-Y. 2019. Probing neural network comprehension of natural language arguments. arXiv preprint arXiv:1907.07355.
  25. 25.Pang, B., and Lee, L. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In Proceedings of the 43rd annual meeting on association for computational linguistics, 115–124.
  26. 26.Papernot, N.; McDaniel, P.; Goodfellow, I.; Jha, S.; Celik, Z. B.; and Swami, A. 2017. Practical black-box attacks against machine learning. In Proceedings of the 2017 ACM on Asia Conference on Computer and Communications Security, 506–519. ACM.
  27. 27.Pennington, J.; Socher, R.; and Manning, C. 2014. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), 1532–1543.
  28. 28.Ribeiro, M. T.; Singh, S.; and Guestrin, C. 2018. Semantically equivalent adversarial rules for debugging nlp models. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, 856–865.
  29. 29.Szegedy, C.; Zaremba, W.; Sutskever, I.; Bruna, J.; Erhan, D.; Goodfellow, I.; and Fergus, R. 2013. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199.
  30. 30.Williams, A.; Nangia, N.; and Bowman, S. R. 2017. A broad-coverage challenge corpus for sentence understanding through inference. arXiv preprint arXiv:1704.05426.
  31. 31.Zhang, X.; Zhao, J.; and LeCun, Y. 2015. Character-level convolutional networks for text classification. In Advances in neural information processing systems, 649–657.
  32. 32.Zhao, Z.; Dua, D.; and Singh, S. 2017. Generating natural adversarial examples. arXiv preprint arXiv:1710.11342.

Citation

MLA
Jin, D., et al. “Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment”. arXiv, 2019, http://arxiv.org/abs/1907.11932v6.
APA
Jin, D., Jin, Z., Zhou, J. T., & Szolovits, P. (2019). Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment. arXiv. http://arxiv.org/abs/1907.11932v6
Chicago
Jin, D., Z. Jin, J. T. Zhou, and P. Szolovits. 2019. “Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment”. arXiv. http://arxiv.org/abs/1907.11932v6.
Harvard
Jin, D. et al. (2019) “Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1907.11932v6.
Vancouver
1. Jin D, Jin Z, Zhou JT, Szolovits P (2019) Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment. arXiv

BibTeX

@article{jin2019bert,
  title = {Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment},
  author = {Jin, Di and Jin, Zhijing and Zhou, Joey Tianyi and Szolovits, Peter},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1907.11932v6},
  eprint = {1907.11932}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/