Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
Vyas RainaAdian LiusieMark J. F. Gales
Demonstrates that appending short, transferable adversarial phrases to text can trick LLM-as-a-judge evaluators into assigning maximum quality scores regardless of actual content, revealing critical vulnerabilities in zero-shot absolute scoring.
Organizations increasingly rely on large language models as automated evaluators across critical domains, such as benchmarking new artificial intelligence systems and grading academic examinations. While these automated evaluators correlate well with human judgment without requiring task-specific training, their security and robustness against intentional manipulation have remained largely unexamined. If malicious actors or test candidates can manipulate automated judges, the integrity of academic credentials and industry benchmarks is severely compromised. The article evaluates whether appending short, universal phrases to candidate text can systematically deceive language model evaluators into assigning maximum quality scores regardless of actual content quality.
To investigate this risk under realistic conditions, the researchers developed an attack framework assuming the adversary lacks direct access to the target model's internal weights. Using a relatively small surrogate model, the researchers applied an iterative search method across standard summarization and dialogue evaluation benchmarks to identify short universal attack phrases of one to four words. These phrases were then tested for transferability against several widely used target models, including Llama2-7B, Mistral-7B, and ChatGPT, across two common assessment paradigms: absolute numerical scoring and pairwise comparative assessment.
The analysis yielded several critical findings. First, absolute scoring setups are exceptionally fragile; appending a universal phrase of only four words caused absolute scores to surge near the maximum possible rating, elevating candidate outputs to top rankings regardless of their true quality. Second, these adversarial phrases demonstrated high transferability, meaning attacks optimized on a small, accessible surrogate model successfully fooled distinct and larger proprietary models such as ChatGPT. Third, pairwise comparative assessment—where the system compares two candidate texts side-by-side—proved significantly more robust against universal attacks than absolute scoring, showing only minor score inflation. Finally, the researchers found that evaluating text naturalness using perplexity scoring served as a viable preliminary defense, detecting manipulated inputs with balanced accuracy scores between 70% and 82%.
These results carry substantial operational, academic, and reputational risks for organizations deploying automated language model evaluations. Relying on direct absolute scoring introduces an immediate vulnerability to gaming and fraud, which could distort public model leaderboards or compromise automated grading integrity. The findings demonstrate that automated evaluation systems cannot be assumed secure by default, and organizations must re-evaluate how they deploy automated judging pipelines in high-stakes environments.
To mitigate these vulnerabilities, decision-makers should immediately transition high-stakes evaluation pipelines from direct absolute scoring to pairwise comparative assessment, despite the higher computational cost of processing candidate pairs. Additionally, organizations should integrate input-filtering layers, such as text perplexity detectors, to flag unnatural phrase additions before text reaches the evaluation model. Because current experiments focused on zero-shot evaluation and simple phrase additions, future work should assess whether few-shot prompting, improved prompt design, or adaptive attacks alter these vulnerabilities before fully relying on automated language model evaluation systems.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). This foundational work introduces universal and transferable adversarial suffix attacks on LLMs, providing the core optimization methodology that the source adapts to the LLM-as-a-judge setting.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This paper establishes the paradigm, benchmark protocols, and systemic biases of using LLMs as automated evaluators, which the source directly audits for adversarial robustness.
- Paper: A Survey on Evaluation of Large Language Models, Yu-Chu Chang et al. (2023). Reading this comprehensive survey provides essential background on the methodologies and reliability challenges in modern LLM evaluation pipelines.
- Paper: Automatically Auditing Large Language Models via Discrete Optimization, Erik Jones et al. (2023). This work establishes discrete coordinate-ascent optimization techniques for auditing LLM failure modes, directly underpinning the surrogate attack algorithms used in the source.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). This study analyzes how competing objectives in safety training lead to alignment vulnerabilities, clarifying why zero-shot LLM judges can be easily fooled by adversarial phrases.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). This survey provides an up-to-date, comprehensive synthesis of the opportunities, architectures, and robustness challenges of the LLM-as-a-judge paradigm.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). This paper presents a formal meta-evaluation framework and taxonomy for mitigating manipulation and biases in LLM evaluator pipelines.
- Paper: JudgeBench: A Benchmark for Evaluating LLM-Based Judges, Sijun Tan et al. (2025). JudgeBench designs an objective evaluation benchmark to systematically stress-test the factual correctness and robustness of LLM-based judges.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). This work introduces LLMBAR to meta-evaluate whether automated LLM judges succumb to superficial adversarial polish rather than true instruction compliance.
- Paper: Split and Merge: Aligning Position Biases in LLM-based Evaluators, Zongjie Li et al. (2024). This paper develops an alignment and prompt-structuring framework (PORTIA) to calibrate and mitigate evaluator inconsistencies and position biases in LLM judges.
- Paper: Improved Techniques for Optimization-Based Jailbreaking on Large Language Models, Xiaojun Jia et al. (2025). This study extends optimization-based adversarial attacks with diverse target templates, advancing beyond the single-target surrogate attack mechanisms examined in the source.
- Paper: Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts, Mikayel Samvelyan et al. (2024). Rainbow Teaming broadens adversarial probing into an open-ended quality-diversity search framework that automatically discovers diverse failure modes.
