Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment

Vyas RainaAdian LiusieMark J. F. Gales

article2024EMNLP112 citations

Demonstrates that appending short, transferable adversarial phrases to text can trick LLM-as-a-judge evaluators into assigning maximum quality scores regardless of actual content, revealing critical vulnerabilities in zero-shot absolute scoring.

Listen

Organizations increasingly rely on large language models as automated evaluators across critical domains, such as benchmarking new artificial intelligence systems and grading academic examinations. While these automated evaluators correlate well with human judgment without requiring task-specific training, their security and robustness against intentional manipulation have remained largely unexamined. If malicious actors or test candidates can manipulate automated judges, the integrity of academic credentials and industry benchmarks is severely compromised. The article evaluates whether appending short, universal phrases to candidate text can systematically deceive language model evaluators into assigning maximum quality scores regardless of actual content quality.

To investigate this risk under realistic conditions, the researchers developed an attack framework assuming the adversary lacks direct access to the target model's internal weights. Using a relatively small surrogate model, the researchers applied an iterative search method across standard summarization and dialogue evaluation benchmarks to identify short universal attack phrases of one to four words. These phrases were then tested for transferability against several widely used target models, including Llama2-7B, Mistral-7B, and ChatGPT, across two common assessment paradigms: absolute numerical scoring and pairwise comparative assessment.

The analysis yielded several critical findings. First, absolute scoring setups are exceptionally fragile; appending a universal phrase of only four words caused absolute scores to surge near the maximum possible rating, elevating candidate outputs to top rankings regardless of their true quality. Second, these adversarial phrases demonstrated high transferability, meaning attacks optimized on a small, accessible surrogate model successfully fooled distinct and larger proprietary models such as ChatGPT. Third, pairwise comparative assessment—where the system compares two candidate texts side-by-side—proved significantly more robust against universal attacks than absolute scoring, showing only minor score inflation. Finally, the researchers found that evaluating text naturalness using perplexity scoring served as a viable preliminary defense, detecting manipulated inputs with balanced accuracy scores between 70% and 82%.

These results carry substantial operational, academic, and reputational risks for organizations deploying automated language model evaluations. Relying on direct absolute scoring introduces an immediate vulnerability to gaming and fraud, which could distort public model leaderboards or compromise automated grading integrity. The findings demonstrate that automated evaluation systems cannot be assumed secure by default, and organizations must re-evaluate how they deploy automated judging pipelines in high-stakes environments.

To mitigate these vulnerabilities, decision-makers should immediately transition high-stakes evaluation pipelines from direct absolute scoring to pairwise comparative assessment, despite the higher computational cost of processing candidate pairs. Additionally, organizations should integrate input-filtering layers, such as text perplexity detectors, to flag unnatural phrase additions before text reaches the evaluation model. Because current experiments focused on zero-shot evaluation and simple phrase additions, future work should assess whether few-shot prompting, improved prompt design, or adaptive attacks alter these vulnerabilities before fully relying on automated language model evaluation systems.

Cover for Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment

Abstract

Large Language Models (LLMs) are powerful zero-shot assessors used in real-world situations such as assessing written exams and benchmarking systems. Despite these critical applications, no existing work has analyzed the vulnerability of judge-LLMs to adversarial manipulation. This work presents the first study on the adversarial robustness of assessment LLMs, where we demonstrate that short universal adversarial phrases can be concatenated to deceive judge LLMs to predict inflated scores. Since adversaries may not know or have access to the judge-LLMs, we propose a simple surrogate attack where a surrogate model is first attacked, and the learned attack phrase then transferred to unknown judge-LLMs. We propose a practical algorithm to determine the short universal attack phrases and demonstrate that when transferred to unseen models, scores can be drastically inflated such that irrespective of the assessed text, maximum scores are predicted. It is found that judge-LLMs are significantly more susceptible to these adversarial attacks when used for absolute scoring, as opposed to comparative assessment. Our findings raise concerns on the reliability of LLM-as-a-judge methods, and emphasize the importance of addressing vulnerabilities in LLM assessment methods before deployment in high-stakes real-world scenarios.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Zero-shot Assessment with LLMs
  • 3.1 Comparative Assessment
  • 3.2 Absolute Scoring Assessment
  • 4 Adversarial Assessment Attacks
  • 4.1 Attack Threat Model
  • 4.2 Practical Attack Approach
  • 5 Experimental Setup
  • 5.1 Datasets
  • 5.2 LLM Assessment Systems
  • 5.3 Methodology
  • 5.4 Attack Evaluation
  • 6 Results
  • 6.1 Assessment Performance
  • 6.2 Attack on Surrogate Model
  • 6.3 Transferability of the Surrogate Attack
  • 6.4 Attack Detection
  • 7 Conclusions
  • 8 Limitations
  • 9 Risks & Ethics
  • 10 Acknowledgements
  • References
  • A Universal Adversarial Phrases
  • B Analysis of Relative Robustness of Comparative Assessment
  • C Transferability of the Comparative Assessment Attack
  • D Direct Attack on Target Model
  • E Greedy Coordinate Gradient (GCG) Universal Attack
  • F Interpretable Attack Results
  • G LLM Prompts
  • H Attacking Bespoke Assessment Systems
  • I Licensing

Knowls

  1. Knowl 1 — Universal concatenative attacks expose severe weaknesses in LLM scoring

    empirical result

    The paper shows that a short phrase appended to an assessed response can manipulate zero-shot LLM judges into assigning inflated quality scores, even when the underlying response is unchanged. Four-word phrases learned on the FlanT5-xl surrogate frequently push absolute scores close to the maximum and can make the attacked response rank first among otherwise unrelated candidates. The vulnerability is substantially stronger for absolute scoring than for pairwise comparative assessment: comparative attacks produce only mild rank improvements on the surrogate model. Phrases learned on FlanT5-xl also transfer to larger, unseen judges such as Llama2-7B, Mistral-7B, and GPT3.5, although transfer effectiveness depends on the dataset, attribute, target model, and phrase length.

  2. Knowl 2 — Universal and transferable attack objective

    model/method

    Let d(m)d^{(m)} be the context for training example mm, let x1:N(m)x^{(m)}_{1:N} be its NN candidate responses, and let s^\hat{s} be the scores produced by an LLM judge. A concatenative adversarial phrase δ=(δ1,…,δL)\delta=(\delta_1,\ldots,\delta_L) of L≪∣x∣L\ll |x| tokens is appended to a response xx:

    x+δ=(x1,…,x∣x∣,δ1,…,δL).x+\delta=(x_1,\ldots,x_{|x|},\delta_1,\ldots,\delta_L).

    For a candidate ii, only xix_i is attacked; the other candidates remain unchanged. If r^n′(m)(δ)\hat{r}^{\prime(m)}_n(\delta) is the predicted rank of candidate nn in context mm after candidate nn is attacked with δ\delta, where rank 11 is best, the universal phrase minimizes the mean attacked rank:

    rˉ(δ)=1NM∑m=1M∑n=1Nr^n′(m)(δ),δ∗=arg⁡min⁡δrˉ(δ).\bar r(\delta)=\frac{1}{NM}\sum_{m=1}^{M}\sum_{n=1}^{N}\hat{r}^{\prime(m)}_n(\delta), \qquad \delta^*=\arg\min_{\delta}\bar r(\delta).

    Here MM is the number of training contexts and NN is the number of candidates per context. The practical black-box threat model assumes that the adversary learns δ∗\delta^* using a surrogate judge with query or weight access, then applies the same phrase unchanged to an unknown target judge without target-model access.

  3. Knowl 3 — Greedy search for a universal attack phrase

    algorithm

    The paper learns the phrase by appending one vocabulary token at a time and selecting the token that maximizes the attacked candidates' aggregate predicted quality on the training contexts. The method uses a fixed phrase length LL, a vocabulary VV, training data {(d(m),x1:N(m))}m=1M\{(d^{(m)},x^{(m)}_{1:N})\}_{m=1}^{M}, and a surrogate judge FF. At each iteration, two candidate indices are sampled and reused across the vocabulary search. For comparative assessment, the objective sums the probability that the attacked candidate wins in both orderings; for absolute scoring, it sums the attacked candidate's predicted score. After LL iterations, the resulting phrase is returned. The search requires evaluating every trial token on all MM training contexts at each iteration, with two judge evaluations per context for comparative assessment and one for scoring.

    Input: Training contexts and candidate responses; judge FF; vocabulary VV; phrase length LL; assessment mode
    Output: Universal attack phrase δ\delta
    phrase = empty string
    for l = 1 to L
        Sample candidate indices a and b
        best_value = 0
        best_token = none
        for each token v in VV
            trial = phrase + v
            value = 0
            for each training context mm
                if assessment mode is comparative
                    p1=F(xa(m)+trial,xb(m),d(m))p_1 = F(x_a^{(m)} + trial, x_b^{(m)}, d^{(m)})
                    p2=F(xa(m),xb(m)+trial,d(m))p_2 = F(x_a^{(m)}, x_b^{(m)} + trial, d^{(m)})
                    value = value + p1+(1−p2)p_1 + (1 - p_2)
                else if assessment mode is scoring
                    s=F(xa(m)+trial,d(m))s = F(x_a^{(m)} + trial, d^{(m)})
                    value = value + ss
                end if
            end for
            if value > best_value
                best_value = value
                best_token = v
            end if
        end for
        phrase = phrase + best_token
    end for
    return phrase
  4. Knowl 4 — Operational definitions of comparative and absolute LLM assessment

    model/method

    For a context dd and candidate responses x1:Nx_{1:N}, the judge LLM is represented by FF. In comparative assessment, F(xi,xj,d)F(x_i,x_j,d) is the probability that response xix_i is better than response xjx_j when the two responses are presented in that order. To reduce positional bias, the paper evaluates both orderings and defines

    pij=12(F(xi,xj,d)+1−F(xj,xi,d)).p_{ij}=\frac{1}{2}\left(F(x_i,x_j,d)+1-F(x_j,x_i,d)\right).

    The predicted quality of candidate xnx_n is its mean pairwise win probability,

    s^n=1N∑j=1Npnj,\hat{s}_n=\frac{1}{N}\sum_{j=1}^{N}p_{nj},

    which is converted into a predicted ranking. In absolute scoring, the judge directly outputs a score F(xn,d)F(x_n,d), so s^n=F(xn,d)\hat{s}_n=F(x_n,d). When token probabilities are available, the paper instead uses the probability-weighted expected score

    s^n=∑k=1KkPF(k∣xn,d),\hat{s}_n=\sum_{k=1}^{K} kP_F(k\mid x_n,d),

    where KK is the maximum score specified by the prompt, k∈{1,…,K}k\in\{1,\ldots,K\}, and PF(k∣xn,d)P_F(k\mid x_n,d) is normalized over all allowed score values. GPT3.5 was evaluated with direct scores because its API did not provide token probabilities.

  5. Knowl 5 — Experimental design for surrogate learning and transfer

    experimental setup

    The experiments use SummEval and TopicalChat. SummEval contains 100 passages with 16 machine-generated summaries per passage, evaluated for coherency, consistency, fluency, relevance, and their overall average. TopicalChat contains 60 dialogue contexts with 6 generated responses per context, evaluated for coherency, continuity, engagingness, naturalness, and their overall average.

    The judge models are FlanT5-xl (3B parameters), Llama2-7B-chat, Mistral-7B-chat, and GPT3.5. FlanT5-xl is the only encoder-decoder model and is used as the surrogate for learning every universal phrase; the learned phrases are then transferred to Llama2-7B, Mistral-7B, and GPT3.5. Each dataset is split 20%/80% into development and test portions. Phrase learning uses only two candidate texts per context, giving 40 training summaries for SummEval and 24 training responses for TopicalChat. Separate phrases are learned for each dataset, assessment mode, and evaluation attribute. The greedy vocabulary is taken from the NLTK English word corpus, and attack effectiveness is measured by the mean rank of the attacked candidate on the held-out test data.

  6. Knowl 6 — Four-word attacks nearly saturate absolute scores on the surrogate

    empirical result

    On FlanT5-xl, the learned universal phrases have a mild effect in comparative assessment but a large effect in absolute scoring. The following are the reported raw scores before and after appending four-word phrases; comparative and absolute values are on different scales and are not directly comparable.

    Could not parse LaTeX table

    The corresponding rank measurements show that four-word absolute-scoring attacks make the attacked summary or response rank close to first place for nearly all inputs, whereas comparative attacks remain much less effective. Examples of learned four-word absolute phrases include outstandingly superexcellently outstandingly summable for SummEval overall scoring and informative supercomplete impeccable ovated for TopicalChat overall scoring.

  7. Knowl 7 — Transferability varies by dataset, attribute, model, and phrase length

    empirical result

    Absolute-scoring phrases learned only on FlanT5-xl transfer to larger unseen judges. Transfer is especially strong on TopicalChat continuity, where nearly all target systems become highly vulnerable. Transfer is weaker for overall-quality assessment by the more powerful models, suggesting that broad abstract attributes can be more resistant than specific attributes. GPT3.5 is sometimes more vulnerable to shorter phrases; longer phrases can begin to overfit the surrogate model rather than generalize. Transfer on SummEval is mixed across target models and attributes, indicating that dataset complexity affects attack portability. Direct attacks performed with access to a target model are at least as effective as surrogate transfer attacks, with Llama2-7B absolute scoring readily driven toward the highest rank.

  8. Knowl 8 — Perplexity provides an initial attack-detection defense

    model/method

    The paper detects appended adversarial phrases using the negative length-normalized log probability assigned by a base Mistral-7B language model. For a token sequence xx of length ∣x∣|x| and model parameters θ\theta, the reported perplexity score is

    perp⁡(x)=−1∣x∣log⁡Pθ(x).\operatorname{perp}(x)=-\frac{1}{|x|}\log P_{\theta}(x).

    Given a threshold β\beta, an input is classified as adversarial when perp⁡(x)>β\operatorname{perp}(x)>\beta. On a balanced test set containing clean and attacked examples, detection quality is summarized by precision P=TP/(TP+FP)P=\mathrm{TP}/(\mathrm{TP}+\mathrm{FP}), recall R=TP/(TP+FN)R=\mathrm{TP}/(\mathrm{TP}+\mathrm{FN}), and F1=2PR/(P+R)F1=2PR/(P+R), where TP, FP, and FN are true-positive, false-positive, and false-negative counts. The best values obtained while sweeping β\beta were reported as follows; the paper preserves mixed decimal and percentage formatting in the source values.

    Could not parse LaTeX table

    Thus perplexity separates clean and attacked inputs reasonably well in these experiments, with F1 values around 0.7 or higher for SummEval and generally higher values for TopicalChat.

  9. Knowl 9 — Comparative robustness is not explained solely by positional symmetry

    empirical result

    The paper hypothesizes that comparative assessment is harder to attack because the same phrase must satisfy competing objectives across the two prompt orderings: when the attacked response is first, it should increase the probability of the first-choice token, whereas when it is second, it should decrease that probability. An ablation removes this requirement by always placing the attacked response in one fixed position. The resulting asymmetric attacks do not substantially eliminate comparative robustness. For direct attacks on FlanT5-xl, the no-attack mean rank is 8.508.50; with the attacked response always first, the mean ranks for one, two, and three appended words are 6.176.17, 9.809.80, and 6.816.81, respectively, while with the attacked response always second, the mean ranks for one through four words are 9.529.52, 8.168.16, 8.548.54, and 7.067.06. Because lower rank is stronger, these results show only limited and inconsistent improvement relative to the unattacked condition. The paper therefore concludes that another property of comparative assessment, beyond symmetric ordering, likely contributes to its robustness.

  10. Knowl 10 — Scope limitations of the demonstrated attacks and defenses

    limitation

    The study evaluates simple concatenation attacks found by greedy search and a correspondingly simple perplexity detector. It does not establish robustness against more subtle or adaptive attacks designed to evade perplexity detection. The experiments focus on zero-shot assessment; few-shot prompting may be more robust, and the paper does not test whether few-shot examples or attack-resistant prompt designs prevent the demonstrated failures. The authors also note that comparative assessment is a practical defense against the strongest observed attacks, but it requires substantially more inference because pairwise assessment evaluates all ordered candidate pairs rather than one score per candidate.

Coverage note — Secondary appendix material was omitted, including the full list of learned phrases, detailed per-candidate score tables, the Greedy Coordinate Gradient comparison, transfer plots for comparative attacks, and the bespoke UniEval attack study; these analyses refine rather than change the main universal-attack, transferability, and detection findings.

References

  1. 1.Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang. 2018. Generating natural language adversarial examples. pages 2890–2896.
  2. 2.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan. Association for Computational Linguistics.
  3. 3.Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, and Ludwig Schmidt. 2023. Are aligned neural networks adversarially aligned?
  4. 4.Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel. 2020. Extracting training data from large language models. CoRR, abs/2012.07805.
  5. 5.Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. 2023. Jailbreaking black box large language models in twenty queries.
  6. 6.Yi Chen, Rui Wang, Haiyun Jiang, Shuming Shi, and Ruifeng Xu. 2023. Exploring the use of large language models for reference-free text quality evaluation: An empirical study. In Findings of the Association for Computational Linguistics: IJCNLP-AACL 2023 (Findings), pages 361–374, Nusa Dua, Bali. Association for Computational Linguistics.
  7. 7.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  8. 8.Alexander R Fabbri, Wojciech Krysci ´ nski, Bryan Mc- ´ Cann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021. Summeval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics, 9:391–409.
  9. 9.Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023. Gptscore: Evaluate as you desire. arXiv preprint arXiv:2302.04166.
  10. 10.Ji Gao, Jack Lanchantin, Mary Lou Soffa, and Yanjun Qi. 2018. Black-box generation of adversarial text sequences to evade deep learning classifiers. CoRR, abs/1801.04354.
  11. 11.Siddhant Garg and Goutham Ramakrishnan. 2020. BAE: BERT-based adversarial examples for text classification. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6174–6181, Online. Association for Computational Linguistics.
  12. 12.Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015. Explaining and harnessing adversarial examples.
  13. 13.Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tür. 2019. Topical-Chat: Towards Knowledge-Grounded Open-Domain Conversations. In Proc. Interspeech 2019, pages 1891–1895.
  14. 14.Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen. 2024. Catastrophic jailbreak of open-source LLMs via exploiting generation. In The Twelfth International Conference on Learning Representations.
  15. 15.Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. 2023. Baseline defenses for adversarial attacks against aligned language models.
  16. 16.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825.
  17. 17.Haibo Jin, Ruoxi Chen, Andy Zhou, Jinyin Chen, Yang Zhang, and Haohan Wang. 2024. Guard: Role-playing to generate natural-language jailbreakings to test guideline adherence of large language models.
  18. 18.Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2023. Benchmarking cognitive biases in large language models as evaluators.
  19. 19.Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Aaron Jiaxun Li, Soheil Feizi, and Himabindu Lakkaraju. 2024. Certifying llm safety against adversarial prompting.
  20. 20.Raz Lapid, Ron Langberg, and Moshe Sipper. 2023. Open sesame! universal black box jailbreaking of large language models.
  21. 21.Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu. 2020. BERT-ATTACK: Adversarial attack against BERT using BERT. pages 6193–6202.
  22. 22.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  23. 23.Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. 2023a. Autodan: Generating stealthy jailbreak prompts on aligned large language models.
  24. 24.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023b. G-eval: NLG evaluation using gpt-4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522, Singapore. Association for Computational Linguistics.
  25. 25.Yanpei Liu, Xinyun Chen, Chang Liu, and Dawn Song. 2016. Delving into transferable adversarial examples and black-box attacks. CoRR, abs/1611.02770.
  26. 26.Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu. 2023c. Jailbreaking chatgpt via prompt engineering: An empirical study.
  27. 27.Yiqi Liu, Nafise Sadat Moosavi, and Chenghua Lin. 2023d. Llms as narcissistic evaluators: When ego inflates evaluation scores.
  28. 28.Adian Liusie, Potsawee Manakul, and Mark JF Gales. 2023. Zero-shot nlg evaluation through pairware comparisons with llms. arXiv preprint arXiv:2307.07889.
  29. 29.Potsawee Manakul, Adian Liusie, and Mark JF Gales. 2023. Mqag: Multiple-choice question answering and generation for assessing information consistency in summarization. arXiv preprint arXiv:2301.12307.
  30. 30.Shikib Mehri and Maxine Eskenazi. 2020. Unsupervised evaluation of interactive dialog with DialoGPT. In Proceedings of the 21th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 225–235, 1st virtual meeting. Association for Computational Linguistics.
  31. 31.Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi. 2023. Tree of attacks: Jailbreaking black-box llms automatically.
  32. 32.Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee. 2023. Scalable extraction of training data from (production) language models.
  33. 33.Nicolas Papernot, Patrick D. McDaniel, Ian J. Goodfellow, Somesh Jha, Z. Berkay Celik, and Ananthram Swami. 2016. Practical black-box attacks against deep learning systems using adversarial examples. CoRR, abs/1602.02697.
  34. 34.Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, Xuanhui Wang, and Michael Bendersky. 2023. Large language models are effective text rankers with pairwise ranking prompting.
  35. 35.Vyas Raina and Mark Gales. 2023. Sentiment perception adversarial attacks on neural machine translation systems.
  36. 36.Vyas Raina, Mark J.F. Gales, and Kate M. Knill. 2020. Universal Adversarial Attacks on Spoken Language Assessment Systems. In Proc. Interspeech 2020, pages 3855–3859.
  37. 37.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. Comet: A neural framework for mt evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702.
  38. 38.Alexander Robey, Eric Wong, Hamed Hassani, and George J. Pappas. 2023. Smoothllm: Defending large language models against jailbreaking attacks.
  39. 39.Sahar Sadrizadeh, Ljiljana Dolamic, and Pascal Frossard. 2023. A classification-guided approach for adversarial attacks against neural machine translation.
  40. 40.Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2014. Intriguing properties of neural networks.
  41. 41.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  42. 42.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5008–5020.
  43. 43.Jiaan Wang, Yunlong Liang, Fandong Meng, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023a. Is chatgpt a good nlg evaluator? a preliminary study. arXiv preprint arXiv:2303.04048.
  44. 44.Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023b. Large language models are not fair evaluators.
  45. 45.Xiaosen Wang, Hao Jin, and Kun He. 2019. Natural language adversarial attacks and defenses in word level. CoRR, abs/1909.06723.
  46. 46.Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023. Jailbroken: How does llm safety training fail?
  47. 47.Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi. 2024. How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms.
  48. 48.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.
  49. 49.Xinghua Zhang, Bowen Yu, Haiyang Yu, Yangyu Lv, Tingwen Liu, Fei Huang, Hongbo Xu, and Yongbin Li. 2023a. Wider and deeper llm networks are fairer llm evaluators.
  50. 50.Zhexin Zhang, Junxiao Yang, Pei Ke, and Minlie Huang. 2023b. Defending large language models against jailbreaking attacks through goal prioritization.
  51. 51.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2023. A survey of large language models.
  52. 52.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685.
  53. 53.Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022a. Towards a unified multi-dimensional evaluator for text generation. arXiv preprint arXiv:2210.07197.
  54. 54.Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022b. Towards a unified multi-dimensional evaluator for text generation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 2023–2038, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  55. 55.Andy Zhou, Bo Li, and Haohan Wang. 2024. Robust prompt optimization for defending language models against jailbreaking attacks.
  56. 56.Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Yue Zhang, Neil Zhenqiang Gong, and Xing Xie. 2023a. Promptbench: Towards evaluating the robustness of large language models on adversarial prompts.
  57. 57.Lianghui Zhu, Xinggang Wang, and Xinlong Wang. 2023b. Judgelm: Fine-tuned large language models are scalable judges.
  58. 58.Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Furong Huang, and Tong Sun. 2024. AutoDAN: Automatic and interpretable adversarial attacks on large language models.
  59. 59.Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models.

Citation

MLA
Raina, V., et al. “Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 7499–517, https://doi.org/10.18653/v1/2024.emnlp-main.427.
APA
Raina, V., Liusie, A., & Gales, M. (2024). Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 7499–7517. https://doi.org/10.18653/v1/2024.emnlp-main.427
Chicago
Raina, V., A. Liusie, and M. Gales. 2024. “Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 7499–7517. https://doi.org/10.18653/v1/2024.emnlp-main.427.
Harvard
Raina, V., Liusie, A. and Gales, M. (2024) “Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 7499–7517. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.427.
Vancouver
1. Raina V, Liusie A, Gales M (2024) Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 7499–7517

BibTeX

@inproceedings{raina-etal-2024-llm,
    title = "Is {LLM}-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot {LLM} Assessment",
    author = "Raina, Vyas  and
      Liusie, Adian  and
      Gales, Mark",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.427/",
    doi = "10.18653/v1/2024.emnlp-main.427",
    pages = "7499--7517"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/