What Evidence Do Language Models Find Convincing?
Alexander WanEric WallaceDan Klein
Reveals that retrieval-augmented language models judge evidence primarily by topical relevance rather than human credibility markers like scientific citations or neutral tone, exposing critical vulnerabilities in how AI handles contentious queries.
Retrieval-augmented language models are increasingly deployed to answer complex, subjective, and controversial real-world queries by drawing on internet sources. However, web retrieval frequently surfaces noisy, contradictory, and misleading content. While humans use well-established critical thinking skills, source credibility evaluations, and logical analysis to assess conflicting evidence, little is known about how automated systems decide which sources to trust when generating answers.
The article evaluates what specific text features make evidence convincing to large language models when they are presented with conflicting viewpoints. Specifically, it demonstrates how stylistic properties and query relevance metrics influence model decision-making during open-ended question answering.
To conduct this evaluation, the researchers developed a new benchmark containing 238 controversial questions across 191 categories paired with real-world web evidence retrieved via search engine queries. The setup emulates production systems by providing models with conflicting evidence paragraphs supporting opposite stances and measuring the rate at which model predictions align with each paragraph's viewpoint. The analysis evaluated both leading open-source models, such as LLaMA-2 Chat and Vicuna, and commercial systems including GPT-4 and Claude. The researchers conducted sensitivity analyses on inherent text features like readability, sentiment, perplexity, and semantic similarity, as well as counterfactual experiments where evidence texts were systematically edited to alter style or relevance.
The findings show that language models rely predominantly on relevance metrics rather than stylistic indicators of credibility. First, high semantic embedding similarity and exact keyword overlap between the query and the retrieved text strongly predict which evidence a model favors. Second, simple superficial adjustments to relevance—such as adding a single introductory sentence stating what question the text addresses—substantially increase an evidence source's win rate. Third, stylistic attributes that humans typically associate with credibility, such as neutral tone, scientific references, high informational content, or technical phrasing, had a neutral or even negative effect on model preference. Fourth, models are incapable of reliably articulating source credibility when prompted directly in isolation, despite displaying strong behavioral biases during paired comparisons.
These results demonstrate a substantial divergence between human and machine assessments of evidence quality. By heavily favoring superficial topical relevance over rigorous argumentation or factual credibility, retrieval-augmented models are highly vulnerable to manipulation. Low-quality sources, search engine optimization tactics, and adversarial misinformation campaigns can easily influence model outputs simply by packing relevant keywords or framing text to match user queries.
To mitigate these risks, organizations deploying retrieval-augmented systems must focus on source corpus governance and data filtering rather than relying on the language model to judge evidence validity autonomously. System designers should prioritize curated, pre-verified knowledge repositories and consider incorporating explicit source metadata or prompt constraints. Where topics involve genuine controversy or ambiguity, systems should be designed to present balanced perspectives or seek clarification rather than autonomously resolving conflicts.
These conclusions are bounded by specific experimental conditions, including the evaluation of text-only paragraphs, binary question formats, and two-document comparison setups. While the core findings reliably highlight current algorithmic tendencies across major model families, ongoing research is necessary to evaluate multi-document contexts, non-textual metadata, and evolving model training techniques aimed at better aligning machine judgment with human standards of evidence.
- Paper: Whose Opinions Do Language Models Reflect?, Shibani Santurkar et al. (2023). This study establishes how language models reflect demographic opinions on subjective and contentious queries, providing the foundation for analyzing how models evaluate conflicting evidence on controversial topics.
- Paper: Enabling Large Language Models to Generate Text with Citations, Tianyu Gao et al. (2023). This work introduces benchmark methodology for evaluating citation support and evidence attribution in retrieval-augmented models, which is foundational to examining which evidence features models find convincing.
- Paper: Towards Understanding Sycophancy in Language Models, Mrinank Sharma et al. (2023). It investigates how language models are biased toward confirming user stances rather than objective evidence, explaining the vulnerability of LLM decision-making under contentious contexts.
- Paper: Large Language Models Can Be Easily Distracted by Irrelevant Context, Freda Shi et al. (2023). It provides critical background on how language models are easily swayed or distracted by irrelevant context features, directly motivating counterfactual analyses of evidence document properties.
- Paper: Evidentiality-guided Generation for Knowledge-Intensive NLP Tasks, Akari Asai et al. (2022). This paper explores evidentiality and distinguishing authentic supporting evidence from superficial distractor passages in retrieval-augmented generation.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). It demonstrates how reasoning in language models can be swayed by unfaithful superficial cues rather than grounded facts, contextualizing why LLMs may ignore scientific references.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). This benchmark measures how language models mimic human misconceptions and falsehoods, establishing the baseline challenge of truthfulness and evidence evaluation.
- Paper: Retrieval-Augmented Generation for Large Language Models: A Survey, Yunfan Gao et al. (2023). This survey systematically covers the paradigms and evaluation mechanisms of retrieval-augmented generation models.
- Paper: Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training, Feiteng Fang et al. (2024). It builds directly on findings regarding model vulnerability to noisy or counterfactual context by introducing adaptive adversarial training to enhance RAG noise robustness.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). It develops a unified framework to train LLMs to explicitly rank and filter retrieved contexts, addressing the corpus quality and evidence selection issues exposed in the source.
- Paper: Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA, Minzheng Wang et al. (2024). It extends the evaluation of conflicting and distributed evidence to realistic, long-context multi-document question answering across enterprise and technical domains.
- Paper: Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning, Bowen Jin et al. (2025). It presents a reinforcement learning approach for autonomous search and evidence extraction, moving beyond passive retrieval corpora to active evidence gathering.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). It provides a comprehensive taxonomy and evaluation framework for LLM-as-a-judge systems, extending the inquiry into how model evaluators weight evidence and stylistic features.
- Paper: Split and Merge: Aligning Position Biases in LLM-based Evaluators, Zongjie Li et al. (2024). It tackles the position and stylistic biases of language models acting as evaluators by introducing alignment and calibration strategies for pairwise comparisons.
