What Evidence Do Language Models Find Convincing?

Alexander WanEric WallaceDan Klein

article2024ACL67 citations

Reveals that retrieval-augmented language models judge evidence primarily by topical relevance rather than human credibility markers like scientific citations or neutral tone, exposing critical vulnerabilities in how AI handles contentious queries.

Listen

Retrieval-augmented language models are increasingly deployed to answer complex, subjective, and controversial real-world queries by drawing on internet sources. However, web retrieval frequently surfaces noisy, contradictory, and misleading content. While humans use well-established critical thinking skills, source credibility evaluations, and logical analysis to assess conflicting evidence, little is known about how automated systems decide which sources to trust when generating answers.

The article evaluates what specific text features make evidence convincing to large language models when they are presented with conflicting viewpoints. Specifically, it demonstrates how stylistic properties and query relevance metrics influence model decision-making during open-ended question answering.

To conduct this evaluation, the researchers developed a new benchmark containing 238 controversial questions across 191 categories paired with real-world web evidence retrieved via search engine queries. The setup emulates production systems by providing models with conflicting evidence paragraphs supporting opposite stances and measuring the rate at which model predictions align with each paragraph's viewpoint. The analysis evaluated both leading open-source models, such as LLaMA-2 Chat and Vicuna, and commercial systems including GPT-4 and Claude. The researchers conducted sensitivity analyses on inherent text features like readability, sentiment, perplexity, and semantic similarity, as well as counterfactual experiments where evidence texts were systematically edited to alter style or relevance.

The findings show that language models rely predominantly on relevance metrics rather than stylistic indicators of credibility. First, high semantic embedding similarity and exact keyword overlap between the query and the retrieved text strongly predict which evidence a model favors. Second, simple superficial adjustments to relevance—such as adding a single introductory sentence stating what question the text addresses—substantially increase an evidence source's win rate. Third, stylistic attributes that humans typically associate with credibility, such as neutral tone, scientific references, high informational content, or technical phrasing, had a neutral or even negative effect on model preference. Fourth, models are incapable of reliably articulating source credibility when prompted directly in isolation, despite displaying strong behavioral biases during paired comparisons.

These results demonstrate a substantial divergence between human and machine assessments of evidence quality. By heavily favoring superficial topical relevance over rigorous argumentation or factual credibility, retrieval-augmented models are highly vulnerable to manipulation. Low-quality sources, search engine optimization tactics, and adversarial misinformation campaigns can easily influence model outputs simply by packing relevant keywords or framing text to match user queries.

To mitigate these risks, organizations deploying retrieval-augmented systems must focus on source corpus governance and data filtering rather than relying on the language model to judge evidence validity autonomously. System designers should prioritize curated, pre-verified knowledge repositories and consider incorporating explicit source metadata or prompt constraints. Where topics involve genuine controversy or ambiguity, systems should be designed to present balanced perspectives or seek clarification rather than autonomously resolving conflicts.

These conclusions are bounded by specific experimental conditions, including the evaluation of text-only paragraphs, binary question formats, and two-document comparison setups. While the core findings reliably highlight current algorithmic tendencies across major model families, ongoing research is necessary to evaluate multi-document contexts, non-textual metadata, and evolving model training techniques aimed at better aligning machine judgment with human standards of evidence.

arXiv: 2402.11782
Cover for What Evidence Do Language Models Find Convincing?

Abstract

Retrieval-augmented language models are being increasingly tasked with subjective, contentious, and conflicting queries such as “is aspartame linked to cancer”. To resolve these ambiguous queries, one must search through a large range of websites and consider which, if any, of this evidence do I find convincing? In this work, we study how LLMs answer this question. In particular, we construct CONFLICTINGQA, a dataset that pairs controversial queries with a series of real-world evidence documents that contain different facts (e.g., quantitative results), argument styles (e.g., appeals to authority), and answers (Yes or No). We use this dataset to perform sensitivity and counterfactual analyses to explore which text features most affect LLM predictions. Overall, we find that current models rely heavily on the relevance of a website to the query, while largely ignoring stylistic features that humans find important such as whether a text contains scientific references or is written with a neutral tone. Taken together, these results highlight the importance of RAG corpus quality (e.g., the need to filter misinformation), and possibly even a shift in how LLMs are trained to better align with human judgements.

Table of Contents

  • 1 Introduction
  • 2 Background and Motivations
  • 3 The CONFLICTINGQA Dataset
  • 4 Experimental Results
  • 4.1 Convincingness as Paragraph Win Rate
  • 4.2 Implementation Details
  • 4.3 What Correlates With Convincingness?
  • 4.4 Counterfactual Analysis
  • 5 Discussion & Related Work
  • 6 Conclusion
  • Limitations
  • Acknowledgments
  • References
  • A Additional Details on CONFLICTINGQA
  • B Model-based Data Filtering
  • C Additional Results
  • D Expressing Convincingness in Isolation
  • E Counterfactual Perturbations

Knowls

  1. Knowl 1 — CONFLICTINGQA benchmark for conflicting web evidence

    model/method

    CONFLICTINGQA is a benchmark of contentious binary questions paired with real web evidence that supports opposing answers. The questions are intended to be answerable as Yes/No while still having online sources arguing for both sides. GPT-4 first generated topic categories and then questions conditioned on those categories; the authors manually removed duplicate questions. Question generation covered 191 categories, while the model-filtered LLaMA-2 Chat evaluation subset represented 144 categories. The benchmark is designed to evaluate which evidence paragraph an LLM follows when two retrieved paragraphs disagree.

  2. Knowl 2 — Web retrieval and evidence-paragraph construction

    model/method

    For each question, GPT-3.5 Turbo generated affirmative and negative query statements, and tokens preserved across the rephrasings were enclosed in quotation marks for exact-match Google Search retrieval. The authors retrieved the top k=20k=20 results for each polarity through the Google Search API, extracted raw text with jusText, and excluded visual page features and metadata such as URLs, dates, and headings. Claude Instant and GPT-4 classified each retrieved document as supporting Yes, supporting No, or irrelevant; only documents with an agreeing stance classification were retained, and irrelevant documents were discarded. Each document was divided into 512-token windows with a 256-token stride. TAS-B embeddings were used to compute each window's dot product with the question embedding, and only the highest-scoring window was retained when its score was at least 95. A downstream model-based check further removed paragraphs whose predicted stance disagreed with the assigned stance, producing groups of opposing evidence paragraphs for each question.

  3. Knowl 3 — Scale and source diversity of the evaluation subset

    data/table

    For the LLaMA-2 Chat evaluation subset, CONFLICTINGQA contains 238 questions, 2,208 retrieved paragraphs, an average paragraph length of 365.01 words, and 912 paragraphs that can be compared against at least five opposing paragraphs. Each paragraph participates in an average of 6.54 comparisons. The most common top-level domains are .com with 527 sources, .org with 175, .gov with 59, .edu with 57, and .net with 12; the dataset summary records 39 unique domains. These statistics show that the benchmark compares evidence from many web sources rather than relying on a small collection of curated documents.

  4. Knowl 4 — Paragraph win-rate as model-specific convincingness

    equation

    Let qq be a controversial question, let s∈{yes,no}s\in\{\mathrm{yes},\mathrm{no}\} be an evidence stance, and let Pq,sP_{q,s} be the set of retrieved paragraphs for question qq with stance ss. Let f(p1,p2,q)∈{yes,no}f(p_1,p_2,q)\in\{\mathrm{yes},\mathrm{no}\} be an LLM's binary answer when given paragraphs p1p_1 and p2p_2 together with question qq. For a Yes paragraph pyes∈Pq,yesp_{\mathrm{yes}}\in P_{q,\mathrm{yes}}, its model-specific win-rate is the empirical probability that the model answers Yes when the paragraph is paired with a No paragraph:

    WR(pyes,q)=Ep∼Pq,no[1{f(pyes,p,q)=yes}].\mathrm{WR}(p_{\mathrm{yes}},q)=\mathbb{E}_{p\sim P_{q,\mathrm{no}}}\left[\mathbf{1}\{f(p_{\mathrm{yes}},p,q)=\mathrm{yes}\}\right].

    The analogous definition applies to No paragraphs. Win-rates are computed for both orderings of the two evidence paragraphs because evidence order can bias model predictions. The evaluation retains only win-rate estimates based on at least five distinct opposing paragraphs, so convincingness measures how often a paragraph wins against a set of conflicting alternatives rather than against a single comparison.

  5. Knowl 5 — Cross-model evaluation protocol

    experimental setup

    The study evaluates LLaMA-2 Chat, Vicuna v1.5, WizardLM v1.2, GPT-4, and Anthropic Claude v1 Instant. Each model receives a question and two conflicting paragraphs and is instructed to use only the information in those paragraphs, preventing the evaluation from primarily measuring the model's prior stance on the topic. Open-source models are evaluated through next-token log-probabilities, while closed-source models are prompted to output only Yes or No. For sensitivity analyses, the data are balanced so that each question and answer polarity contains equal numbers of paragraphs with high and low convincingness; this prevents a systematic preference for one answer from affecting the comparison unevenly.

  6. Knowl 6 — Direct credibility ratings do not recover pairwise convincingness

    empirical result

    When an LLM is shown a paragraph in isolation and asked to rate its credibility from one to five, the resulting ratings have little correlation with the paragraph's pairwise win-rate. The isolation prompt asks only for a one-to-five credibility judgment; the study converts the model's label probabilities into a probability-weighted average and supplies examples of low- and high-credibility texts to reduce the model's tendency toward high ratings. In contrast, the win-rate evaluation measures whether the model follows a paragraph when that paragraph competes with an opposing source. Because the two measurements disagree substantially, the study uses pairwise answer behavior as its operational measure of convincingness.

  7. Knowl 7 — Question relevance dominates correlates of convincingness

    empirical result

    The study correlates paragraph win-rate with six automatically computed features: Flesch-Kincaid readability, number of unique lemmas, FLAN-large positive-sentiment probability, GPT-2 Medium perplexity, maximum shared n-gram length between the question and paragraph, and TAS-B question-paragraph embedding similarity. Readability and vocabulary diversity show no meaningful relationship with win-rate across the evaluated models. Sentiment and perplexity have smaller, model-dependent effects. Question-paragraph embedding similarity is strongly associated with win-rate for every evaluated model except GPT-4, while maximum n-gram overlap has a weaker positive association. Thus, measures of how directly a paragraph resembles the question are substantially more predictive than the stylistic measures tested.

  8. Knowl 8 — Counterfactual edits reveal a relevance bias

    empirical result

    The authors use Claude v1 Instant to edit Yes-stance paragraphs and compare their win-rates with a control edit that only appends 'Thanks for reading!'. Relevance-oriented edits include rewriting the paragraph around the question, adding question-related keywords, prefixing it with an explicit statement that the text concerns the question, and adding one sentence that makes its stance obvious. Style-oriented edits add scientific references, unrelated topic information, fake contact information, greater confidence, more technical language, or a more objective tone. Relative to the control, relevance edits substantially increase win-rate, even when they change little of the original content; stylistic edits are generally neutral or reduce win-rate. The maximum shared n-gram lengths between the question and edited paragraph confirm that the intended relevance edits increase lexical overlap: Rewrite Relevance 5.08, Add Single Sentence 3.64, Question Prefix 6.89, Keyword Stuffing 2.47, Add Science Reference 2.38, Add More Info 2.46, Add Contact Info 2.41, Rewrite Confidence 2.17, Rewrite Technical Language 1.97, and Rewrite Objective 2.15. The results indicate that simple relevance cues can make a paragraph more influential without making its evidence more credible by human-oriented standards.

  9. Knowl 9 — Conflicts include underspecification, time dependence, and nonexclusive evidence

    data/table

    The collected conflicts are not limited to two sources making directly incompatible claims. Some questions are underspecified: one source about whether coral snakes occur in Africa discusses Old World coral snakes, whereas an opposing source discusses New World coral snakes in the southern United States. Some conflicts depend on time: one source describes Florida panther numbers and genetic diversity as declining, while another reports a recent population recovery. Other pairs support different answers without addressing the same entity: for artificial-sweetener safety for diabetics, one source discusses cardiovascular risks associated with erythritol while another says that xylitol can be consumed safely by diabetics. These examples show that CONFLICTINGQA captures realistic knowledge conflicts arising from ambiguous scope, changing facts, and heterogeneous entities as well as direct contradiction.

  10. Knowl 10 — Scope and dual-use limitations

    limitation

    CONFLICTINGQA does not cover every kind of controversial question or web text, and the experiments provide only two paragraphs at a time. The benchmark restricts answers to binary Yes/No decisions, whereas real LLM responses may need qualifications, uncertainty, or multiple perspectives. The study focuses primarily on text and omits potentially important credibility signals such as metadata, page design, and visual content. Its findings may also be used adversarially to make misinformation appear convincing or to manipulate retrieval-augmented systems through relevance and search-engine-optimization tactics. The authors therefore identify source restriction, richer input modalities, and improved alignment with human judgments as open directions rather than established solutions.

Coverage note — Detailed prompt templates, auxiliary model-specific appendix plots, and discussion proposals such as abstention or clarification questions were omitted because they support the main benchmark, methods, results, and limitations rather than adding separate load-bearing findings.

References

  1. 1.Sahar Abdelnabi and Mario Fritz. 2023. Fact-Saboteurs: A taxonomy of evidence manipulation attacks against fact-verification systems. In USENIX.
  2. 2.Adept. 2022. ACT-1: Transformer for actions.
  3. 3.Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik R Narasimhan, and Ameet Deshpande. 2023. GEO: Generative engine optimization. arXiv preprint arXiv:2311.09735.
  4. 4.Vasilisa Bashlovkina, Zhaobin Kuang, Riley Matthews, Edward Clifford, Yennie Jun, William W. Cohen, and Simon Baumgartner. 2023. Trusted source alignment in large language models. arXiv preprint arXiv:2311.06697.
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, et al. 2020. Language models are few-shot learners. In NeurIPS.
  6. 6.Daniel Bush and Alex Zaheer. 2019. Bing’s top search results contain an alarming amount of disinformation. Internet Observatory News.
  7. 7.Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading Wikipedia to answer open-domain questions. In ACL.
  8. 8.Hung-Ting Chen, Michael JQ Zhang, and Eunsol Choi. 2022. Rich knowledge sources bring complex knowledge conflicts: Recalibrating models to reflect conflicting evidence. In EMNLP.
  9. 9.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality.
  10. 10.Yibing Du, Antoine Bosselut, and Christopher D Manning. 2022. Synthetic disinformation attacks on automated fact verification systems. In AAAI.
  11. 11.Ronen Eldan and Yuanzhi Li. 2023. Tinystories: How small can language models be and still speak coherent english.
  12. 12.Andrew J. Flanagin and Miriam J. Metzger. 2000. Perceptions of internet information credibility. Journalism & Mass Communication Quarterly.
  13. 13.B. J. Fogg, Cathy Soohoo, David R. Danielson, Leslie Marable, Julianne Stanford, and Ellen R. Tauber. 2003. How do users evaluate the credibility of web sites? A study with over 2,500 participants. In Designing for User Experiences.
  14. 14.Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. PAL: Program-aided language models. In ICML.
  15. 15.Martin Gleize, Eyal Shnarch, Leshem Choshen, Lena Dankin, Guy Moshkowich, Ranit Aharonov, and Noam Slonim. 2019. Are you convinced? Choosing the more convincing evidence with a siamese network. In ACL.
  16. 16.Shai Gretz, Roni Friedman, Edo Cohen-Karlik, Assaf Toledo, Dan Lahav, Ranit Aharonov, and Noam Slonim. 2019. A large-scale dataset for argument quality ranking: Construction and analysis.
  17. 17.Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, and Yuanzhi Li. 2023. Textbooks are all you need.
  18. 18.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In ICML.
  19. 19.Sebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin, and Allan Hanbury. 2021. Efficiently teaching an effective dense retriever with balanced topic aware sampling. In SIGIR.
  20. 20.Robin Jia and Percy Liang. 2017. Adversarial examples for evaluating reading comprehension systems. In EMNLP.
  21. 21.Yonggeol Jo, Minwoo Kim, and Kyungsik Han. 2019. How do humans assess the credibility on web blogs: Qualifying and verifying human factors with machine learning. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems.
  22. 22.Michal Kakol, Michał Jankowski-Lorek, Katarzyna Abramczuk, Adam Wierzbicki, and Michele Catasta. 2013. On the subjectivity and bias of web content credibility evaluations. In WWW.
  23. 23.Michal Kakol, Radoslaw Nielek, and Adam Wierzbicki. 2017. Understanding and predicting web content credibility using the content credibility corpus. Information Processing & Management.
  24. 24.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In EMNLP.
  25. 25.J Peter Kincaid, Robert P Fishburne Jr, Richard L Rogers, and Brad S Chissom. 1975. Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for Navy enlisted personnel. Technical report, Naval Technical Training Command Research Branch.
  26. 26.Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021. Entity-based knowledge conflicts in question answering. In EMNLP.
  27. 27.Miriam J. Metzger, Andrew J. Flanagin, and Ryan Bradley Medders. 2010. Social and heuristic approaches to credibility evaluation online. Journal of Communication.
  28. 28.Grégoire Mialon, Roberto Dessì, Maria Lomeli, Christoforos Nalmpantis, Ram Pasunuru, Roberta Raileanu, Baptiste Rozière, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, et al. 2023. Augmented language models: A survey. In TMLR.
  29. 29.Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020. AmbigQA: Answering ambiguous open-domain questions. In EMNLP.
  30. 30.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. WebGPT: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332.
  31. 31.Liangming Pan, Wenhu Chen, Min-Yen Kan, and William Yang Wang. 2023. Attacking open-domain question answering by injecting misinformation. In AACL.
  32. 32.Alec Radford, Jeffrey Wu, Rewon Child, Dario Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog.
  33. 33.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR.
  34. 34.Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. In-context retrieval-augmented language models. TACL.
  35. 35.Sudha Rao and Hal Daumé III. 2018. Learning to ask good questions: Ranking clarification questions using neural expected value of perfect information. In ACL.
  36. 36.Toran Bruce Richards. 2023. AutoGPT. https://github.com/Significant-Gravitas/AutoGPT.
  37. 37.Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? In ICML.
  38. 38.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:2302.04761.
  39. 39.Dushyant Sharma, Rishabh Shukla, Anil Kumar Giri, and Sumit Kumar. 2019. A brief review on search engine optimization. In Confluence.
  40. 40.Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Rich James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2023. REPLUG: Retrieval-augmented black-box language models. arXiv preprint arXiv:2301.12652.
  41. 41.Haitian Sun, William W Cohen, and Ruslan Salakhutdinov. 2023. Answering ambiguous questions with a database of questions, answers, and revisions. arXiv preprint arXiv:2308.08661.
  42. 42.Assaf Toledo, Shai Gretz, Edo Cohen-Karlik, Roni Friedman, Elad Venezian, Dan Lahav, Michal Jacovi, Ranit Aharonov, and Noam Slonim. 2019. Automatic argument quality assessment – new datasets and methods.
  43. 43.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  44. 44.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. LLaMA 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  45. 45.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. Finetuned language models are zero-shot learners. In ICLR.
  46. 46.Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su. 2023. Adaptive chameleon or stubborn sloth: Unraveling the behavior of large language models in knowledge conflicts. arXiv preprint arXiv:2305.13300.
  47. 47.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. WizardLM: Empowering large language models to follow complex instructions. arXiv preprint arXiv:2304.12244.
  48. 48.Hamed Zamani, Susan T. Dumais, Nick Craswell, Paul N. Bennett, and Gord Lueck. 2020. Generating clarifying questions for information retrieval. The Web Conference.
  49. 49.Michael JQ Zhang and Eunsol Choi. 2021. SituatedQA: Incorporating extra-linguistic contexts into QA. In EMNLP.

Citation

MLA
Wan, A., et al. “What Evidence Do Language Models Find Convincing?”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 7468–84, https://doi.org/10.18653/v1/2024.acl-long.403.
APA
Wan, A., Wallace, E., & Klein, D. (2024). What Evidence Do Language Models Find Convincing?. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7468–7484. https://doi.org/10.18653/v1/2024.acl-long.403
Chicago
Wan, A., E. Wallace, and D. Klein. 2024. “What Evidence Do Language Models Find Convincing?”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7468–84. https://doi.org/10.18653/v1/2024.acl-long.403.
Harvard
Wan, A., Wallace, E. and Klein, D. (2024) “What Evidence Do Language Models Find Convincing?”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7468–7484. Available at: https://doi.org/10.18653/v1/2024.acl-long.403.
Vancouver
1. Wan A, Wallace E, Klein D (2024) What Evidence Do Language Models Find Convincing?. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7468–7484

BibTeX

@inproceedings{wan-etal-2024-evidence,
    title = "What Evidence Do Language Models Find Convincing?",
    author = "Wan, Alexander  and
      Wallace, Eric  and
      Klein, Dan",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.403/",
    doi = "10.18653/v1/2024.acl-long.403",
    pages = "7468--7484"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/