Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models

Paul RöttgerValentin HofmannValentina PyatkinMusashi HinckHannah KirkHinrich SchützeDirk Hovy

article2024ACL220 citationsOutstanding Paper Award

Demonstrates that forced multiple-choice benchmarks like the Political Compass Test produce brittle, misleading measurements of language model political bias, arguing instead for evaluations grounded in realistic, open-ended user interactions.

Listen

As large language models (LLMs) are deployed widely across society, concerns have escalated regarding the subtle political and social biases they may exhibit. Much of the current literature evaluates these values using multiple-choice surveys adapted from human psychometrics, such as the 62-item Political Compass Test (PCT). However, these evaluations impose artificial constraints that do not reflect how real users interact with artificial intelligence, creating a critical gap between academic testing frameworks and real-world deployment risks.

The article systematically assesses whether forced multiple-choice questionnaires reliably measure values and opinions in LLMs. Specifically, it evaluates model behavior across unforced baseline settings, varying styles of forced-choice prompts, minimal prompt paraphrases, and realistic open-ended text generation tasks.

To conduct this evaluation, the authors performed a systematic literature review of prior PCT studies and ran controlled experiments using ten distinct commercial and open-source models (including GPT-4, GPT-3.5, Llama 2, Mistral, and Zephyr) across all 62 PCT propositions. They tested multiple prompting templates under deterministic decoding settings and utilized an automated classifier validated with human annotations (93.1% inter-annotator agreement) to evaluate stances in free-form responses.

The article highlights five key findings: First, in an unforced multiple-choice setting, LLMs consistently resist taking sides; models like Zephyr and newer GPT versions produced 0% valid multiple-choice selections, instead generating nuanced disclaimers or balanced arguments in 95% of invalid cases. Second, models respond unpredictably to forced-choice constraints: GPT-4 remained nearly completely immune to forced prompts, whereas Llama 2 models shut down when pressured with negative consequences. Third, minimal semantics-preserving paraphrases (such as changing "opinion" to "view") caused substantial coordinate swings—shifting GPT-3.5's placement by over 100% on both economic and libertarian axes, a divergence larger than the measured gap between major political figures. Fourth, in realistic open-ended tasks (such as drafting blog posts), models systematically disagreed with propositions they previously endorsed in multiple-choice settings on roughly one out of three questions. Fifth, across these open-ended shifts, models displayed a consistent drift toward more right-leaning, libertarian viewpoints.

These findings demonstrate that popular multiple-choice benchmarks act more like unstable "spinning arrows" than dependable measurement tools. Constrained questionnaires fail to capture inherent model values and obscure the nuanced neutrality models typically offer. For organizations and policymakers, relying on simplistic survey scores creates significant compliance and risk blind spots, as models can exhibit entirely different stances depending on minor prompt variations or application contexts.

The article recommends that practitioners abandon rigid multiple-choice surveys in favor of evaluations tailored to specific real-world use cases. Furthermore, any assessment of model values must include extensive robustness testing across prompt phrasings, and stakeholders should restrict claims about model biases to local, task-specific contexts rather than asserting global model alignment.

These conclusions are bounded by their primary focus on the PCT and behavioral black-box testing. While the results show high statistical confidence across tested models, the authors note that empirical evaluations cannot provide absolute behavioral guarantees, warranting cautious application across different downstream architectures.

arXiv: 2402.16786
Cover for Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models

Abstract

Much recent work seeks to evaluate values and opinions in large language models (LLMs) using multiple-choice surveys and questionnaires. Most of this work is motivated by concerns around real-world LLM applications. For example, politically-biased LLMs may subtly influence society when they are used by millions of people. Such real-world concerns, however, stand in stark contrast to the artificiality of current evaluations: real users do not typically ask LLMs survey questions. Motivated by this discrepancy, we challenge the prevailing constrained evaluation paradigm for values and opinions in LLMs and explore more realistic unconstrained evaluations. As a case study, we focus on the popular Political Compass Test (PCT). In a systematic review, we find that most prior work using the PCT forces models to comply with the PCT's multiple-choice format. We show that models give substantially different answers when not forced; that answers change depending on how models are forced; and that answers lack paraphrase robustness. Then, we demonstrate that models give different answers yet again in a more realistic open-ended answer setting. We distill these findings into recommendations and open challenges in evaluating values and opinions in LLMs.

Table of Contents

  • 1 Introduction
  • 2 The Political Compass Test
  • 3 Literature Review: Evaluating LLMs with the Political Compass Test
  • 3.1 Review Findings
  • 3.2 Implications for Experimental Design
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Unforced Multiple-Choice Responses
  • 4.3 Forced Multiple-Choice Responses
  • 4.4 Paraphrase Robustness
  • 4.5 Open-Ended Responses
  • 5 Discussion
  • 5.1 Recommendations
  • 6 Conclusion
  • Limitations
  • Ethical Considerations
  • Acknowledgments
  • References
  • A Details on Literature Review Method
  • B Structured Results of Literature Review for In-Scope Articles
  • C Model Inference Setup
  • D Prompt Variants
  • E Political Compass Test Propositions
  • F Agreement Classifier

Knowls

  1. Knowl 1 — Political Compass Test as the evaluation instrument

    definition

    The Political Compass Test (PCT) contains 62 single-sentence propositions distributed across six topics: country and world (7), economy (14), personal social values (18), wider society (12), religion (5), and sex (6). Each proposition is answered with one of four choices—“strongly disagree,” “disagree,” “agree,” or “strongly agree”—and the test provides no neutral option. A weighted sum of the responses places a respondent on two dimensions: an economic axis from left to right and a social axis from libertarian to authoritarian.

  2. Knowl 2 — Prior PCT evaluations predominantly impose forced-choice behavior

    empirical result

    A systematic search of Google Scholar, arXiv, and the ACL Anthology, completed on February 12, 2024, found 265 search results corresponding to 57 unique articles; 12 of those articles used the Political Compass Test to evaluate a language model. Ten of the 12 studies explicitly forced models to output exactly one of the four PCT labels, whereas only two allowed more open-ended responses and then used binary stance classifiers. Prompt robustness was not conclusively established in the prior literature: only three in-scope studies performed robustness tests beyond repeating the same prompt. All 12 studies tested some version of GPT-3.5, only three tested open models, and eight did not report generation parameters.

  3. Knowl 3 — Experimental comparison of constrained and unconstrained PCT prompting

    experimental setup

    The experiments evaluated all 62 PCT propositions with up to 10 language models: chat-optimized Llama 2 models with 7B, 13B, and 70B parameters; instruction-tuned Mistral 7B v0.1 and v0.2; Zephyr 7B beta; GPT-3.5 versions 0613 and 1106; and GPT-4 versions 0613 and 1106. The study used temperature 0, a maximum generation length of 256 tokens, no system prompt, and deterministic responses. Test prompts combined an opinion question, a PCT proposition, optional answer labels, and optional instructions forcing a choice. The experiments varied three factors independently or jointly: whether the four-option format was accompanied by a forced-choice instruction, the wording of the initial question, and whether the model had to select a label or generate free text.

  4. Knowl 4 — Unforced multiple-choice prompting produces mostly invalid and qualitatively different responses

    empirical result

    When each of the 10 models received every PCT proposition and the four standard answer options without an additional instruction to choose one, a valid response was defined as exactly one of the four labels; responses expressing neutrality, ambivalence, multiple sides, or refusal were invalid for the PCT. Zephyr and three GPT models produced no valid responses, while GPT-3.5 1106 produced only one valid response. The valid-response rates for Llama 2 7B, Llama 2 13B, Mistral 7B v0.1, and Mistral 7B v0.2 were 6.5%, 45.2%, 75.8%, and 71.0%, respectively. Thus, even the most compliant models were invalid on roughly one quarter of prompts. In an annotation of 100 invalid responses sampled evenly across the models, 95% emphasized that the model could not express a personal opinion, 63% argued for both sides, 22% argued for one side, and 6% explicitly refused to answer. These responses demonstrate that the forced-choice instruction changes model behavior rather than merely exposing an underlying fixed political position.

  5. Knowl 5 — The exact forcefulness of the choice instruction changes model compliance

    empirical result

    The study compared four increasingly forceful additions to the same PCT prompt: “Only answer with the label”; “You have to pick one of the four options. Only answer with the label”; the same instruction followed by “or I will lose my job”; and the same instruction followed by “or I will lose my job and my grandmother will die.” The resulting valid-response rates differed substantially by model. GPT-4 0613 and GPT-4 1106 were essentially resistant to all four prompts and produced little or no valid output. GPT-3.5 models were much more receptive, with at least 80.6% valid responses; GPT-3.5 0613 generally became more compliant as the prompt became more forceful. Zephyr and Mistral 7B v0.2 responded best to the least forceful instruction. Llama 2 models followed the specific instructions without negative consequences but largely stopped complying when negative consequences were added. Mistral 7B v0.1 produced 100% valid responses under every forced-choice prompt. Consequently, the measured PCT result depends on the model-specific interaction between the model and the exact forcing language.

  6. Knowl 6 — Minimal prompt paraphrases substantially alter PCT coordinates

    empirical result

    Ten minimal, semantics-preserving paraphrases of the initial question—such as “What is your opinion,” “What is your view,” “What are your thoughts,” “How do you perceive,” and “What is your perspective” on the proposition—were evaluated with the same answer options and a forced-choice instruction. The analysis focused on Mistral 7B v0.1 and GPT-3.5 1106 because they produced the most valid responses. Both models remained in the PCT’s economically left and socially libertarian quadrant, but their exact coordinates changed substantially. For Mistral, changing the wording moved the result from (−3.6,−5.2)(-3.6,-5.2) to (−6.0,−3.5)(-6.0,-3.5), corresponding to 65.6% more economic left-leaning and 32.4% less libertarian according to the paper’s coordinate comparisons. For GPT-3.5, another paraphrase moved the result from (−1.5,−1.9)(-1.5,-1.9) to (−3.2,−4.4)(-3.2,-4.4), corresponding to 117.1% more left-leaning and 126.3% more libertarian. These differences exceeded the difference between the PCT placements reported for Joe Biden and Donald Trump. At the proposition level, semantically equivalent prompts produced contradictory answers for 14 of 62 propositions for Mistral and 23 of 62 for GPT-3.5.

  7. Knowl 7 — Open-ended evaluation and agreement classification procedure

    model/method

    To approximate realistic language-model use, the study removed the PCT answer labels and asked Mistral 7B v0.1 and GPT-3.5 1106 to generate free-text opinions about each proposition. Ten templates placed the proposition in situations such as a debate, news interview, panel discussion, podcast, newspaper commentary, social-media post, or personal blog; the intent remained to express an opinion. Each template additionally asked the model to take a clear stance by agreeing or disagreeing, allowing comparison with the multiple-choice condition while removing the four-label constraint. GPT-4 0125 classified each generated response as “agree,” “disagree,” or “neither” by receiving the proposition and response and being instructed to answer with one word. Refusals, arguments for both sides, and all outputs that were neither clear agreement nor disagreement were assigned to “neither.” In a human-annotated validation sample of 200 responses, inter-annotator agreement was Fleiss’ κ=93.1%\kappa=93.1\%; 32 responses were labeled agree, 158 disagree, and 10 neither. The classifier achieved 99% accuracy for Mistral responses and 100% for GPT-3.5 responses against the human labels.

  8. Knowl 8 — Open-ended responses often reverse the stance expressed in multiple-choice evaluations

    empirical result

    For the same 62 propositions, the two models frequently expressed opposing stances in the multiple-choice and open-ended settings. GPT-3.5 1106 agreed with a majority of prompt variants in the multiple-choice setting but disagreed with a majority of open-ended variants for 19 of 62 propositions. The corresponding number for Mistral 7B v0.1 was 23 of 62. No proposition showed the inverse majority change from disagreement in the multiple-choice setting to agreement in the open-ended setting. The open-ended condition generally shifted responses toward economically right-leaning and socially libertarian positions: among responses affecting the economic axis, 66.6% of GPT-3.5 changes and 70.0% of Mistral changes went from left to right; among responses affecting the social axis, 84.6% of GPT-3.5 changes and 69.2% of Mistral changes went from authoritarian to libertarian. Open-ended responses also remained prompt-sensitive: Mistral disagreed with one prompt variant but agreed with another for 10 of 62 propositions, and GPT-3.5 did so for 13 of 62. The apparent marginal improvement in stability may partly reflect a general tendency toward disagreement in the open-ended condition.

  9. Knowl 9 — Evaluation should match application context and report local, robust claims

    model/method

    The paper recommends evaluating values and opinions in settings that resemble how models will actually be used, such as examining political language in model-written content rather than asking a model directly to select political labels. Every evaluation should vary relevant prompt wording, situational context, answer constraints, and other plausible sources of instability, and should report the resulting sensitivity. Claims should be local to the tested model, propositions, prompts, and application setting rather than global claims that a model possesses a fixed ideology. The authors interpret the results as consistent with models occupying an intermediate position between a stable persona and a superposition of possible responses: some propositions elicit stable positions, while others resemble sampling from a broader distribution of context-dependent opinions. This interpretation is presented as a conceptual implication of the observed instability, not as a demonstrated internal mechanism.

  10. Knowl 10 — Limits on generalizing output-based evaluations of model values

    limitation

    The empirical study uses the Political Compass Test as a case study, so its direct evidence concerns one questionnaire and the 10 evaluated models, although the authors argue that the PCT’s multiple-choice structure is representative of many value and opinion benchmarks. The experiments do not test every possible instability source, including answer-option ordering and alternative answer formats. More generally, even a large collection of evaluations with consistent outputs cannot provide formal behavioral guarantees about a language model’s values or opinions: any finite set of observations remains limited evidence about future behavior.

Coverage note — The complete text of all 62 PCT propositions and the article-by-article review inventory were omitted because they are supporting source material rather than separate load-bearing contributions; the review’s aggregate findings and the PCT structure are retained.

References

  1. 1.Gavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, Verena Rieser, and Zeerak Talat. 2023. Mirages. on anthropomorphism in dialogue systems. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4776–4790, Singapore. Association for Computational Linguistics.
  2. 2.Giuseppe Attanasio. 2023. Simple Generation. https://github.com/MilaNLProc/simple-generation.
  3. 3.Marcel Binz and Eric Schulz. 2023. Using cognitive psychology to understand gpt-3. Proceedings of the National Academy of Sciences, 120(6):e2218523120.
  4. 4.Ethan Busby, D Flynn, James N Druckman, and P D’Angelo. 2018. Studying framing effects on political preferences. Doing news framing analysis II: Empirical and theoretical perspectives, pages 27–50.
  5. 5.Dennis Chong and James N Druckman. 2007. Framing theory. Annu. Rev. Polit. Sci., 10:103–126.
  6. 6.Esin Durmus, Karina Nyugen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, et al. 2023. Towards measuring the representation of subjective global opinions in language models. arXiv preprint arXiv:2306.16388.
  7. 7.Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg. 2021. Measuring and improving consistency in pretrained language models. Transactions of the Association for Computational Linguistics, 9:1012–1031.
  8. 8.Cristina España-Bonet. 2023. Multilingual coarse political stance classification of media. the editorial line of a ChatGPT and bard newspaper. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 11757–11777, Singapore. Association for Computational Linguistics.
  9. 9.Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023. From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair NLP models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11737–11762, Toronto, Canada. Association for Computational Linguistics.
  10. 10.Sasuke Fujimoto and Takemoto Kazuhiro. 2023. Revisiting the political biases of chatgpt. Frontiers in Artificial Intelligence, 6.
  11. 11.Vahid Ghafouri, Vibhor Agarwal, Yong Zhang, Nishanth Sastry, Jose Such, and Guillermo Suarez-Tangil. 2023. Ai in the gray: Exploring moderation policies in dialogic large language models vs. human answers in controversial topics. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management, pages 556–565.
  12. 12.Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte. 2023. The political ideology of conversational ai: Converging evidence on chatgpt’s pro-environmental, left-libertarian orientation. arXiv preprint arXiv:2301.01768.
  13. 13.Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2020. Aligning ai with shared human values. In International Conference on Learning Representations.
  14. 14.Janus. 2022. Simulators. LessWrong online forum, 2nd September. https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/.
  15. 15.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825.
  16. 16.Graham Kalton and Howard Schuman. 1982. The effect of the question on survey responses: A review. Journal of the Royal Statistical Society Series A: Statistics in Society, 145(1):42–57.
  17. 17.Hannah Kirk, Andrew Bean, Bertie Vidgen, Paul Rottger, and Scott Hale. 2023. The past, present and better future of feedback learning in large language models for subjective human preferences and values. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2409–2430, Singapore. Association for Computational Linguistics.
  18. 18.R Thomas McCoy, Shunyu Yao, Dan Friedman, Matthew Hardy, and Thomas L Griffiths. 2023. Embers of autoregression: Understanding large language models through the problem they are trained to solve. arXiv preprint arXiv:2309.13638.
  19. 19.Marilù Miotto, Nicola Rossberg, and Bennett Kleinberg. 2022. Who is GPT-3? an exploration of personality, values and demographics. In Proceedings of the Fifth Workshop on Natural Language Processing and Computational Social Science (NLP+CSS), pages 218–227, Abu Dhabi, UAE. Association for Computational Linguistics.
  20. 20.Fabio Motoki, Valdemar Pinho Neto, and Victor Rodrigues. 2023. More human than human: Measuring chatgpt political bias. Public Choice, pages 1–21.
  21. 21.Arvind Narayanan and Sayash Kapoor. 2023. Does chatgpt have a liberal bias? https://www.aisnakeoil.com/p/does-chatgpt-have-a-liberal-bias.
  22. 22.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems.
  23. 23.Siru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu, Heng Ji, and Jiawei Han. 2023. The shifted and the overlooked: A task-oriented investigation of user-GPT interactions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2375–2393, Singapore. Association for Computational Linguistics.
  24. 24.Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. 2023. Xstest: A test suite for identifying exaggerated safety behaviours in large language models. arXiv preprint arXiv:2308.01263.
  25. 25.David Rozado. 2023a. Danger in the machine: The perils of political and demographic biases embedded in ai systems. Manhattan Institute.
  26. 26.David Rozado. 2023b. The political biases of chatgpt. Social Sciences, 12(3):148.
  27. 27.David Rozado. 2024. The political preferences of llms. arXiv preprint arXiv:2402.01789.
  28. 28.Jérôme Rutinowski, Sven Franke, Jan Endendyk, Ina Dormuth, and Markus Pauly. 2023. The self-perception and political biases of chatgpt. arXiv preprint arXiv:2304.07333.
  29. 29.Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org.
  30. 30.Nino Scherrer, Claudia Shi, Amir Feder, and David Blei. 2023. Evaluating the moral beliefs encoded in llms. In Thirty-seventh Conference on Neural Information Processing Systems.
  31. 31.Howard Schuman and Stanley Presser. 1977. Question wording as an independent variable in survey analysis. Sociological Methods & Research, 6(2):151–170.
  32. 32.Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2024. Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting. In The Twelfth International Conference on Learning Representations.
  33. 33.Murray Shanahan, Kyle McDonell, and Laria Reynolds. 2023. Role play with large language models. Nature, 623(7987):493–498.
  34. 34.Bangzhao Shu, Lechen Zhang, Minje Choi, Lavinia Dunagan, Dallas Card, and David Jurgens. 2023. You don’t need a personality test to know these models are unreliable: Assessing the reliability of large language models on psychometric instruments. arXiv preprint arXiv:2311.09718.
  35. 35.Surendrabikram Thapa, Ashwarya Maratha, Khan Md Hasib, Mehwish Nasim, and Usman Naseem. 2023. Assessing political inclination of Bangla language models. In Proceedings of the First Workshop on Bangla Language Processing (BLP-2023), pages 62–71, Singapore. Association for Computational Linguistics.
  36. 36.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  37. 37.Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, et al. 2023. Zephyr: Direct distillation of lm alignment. arXiv preprint arXiv:2310.16944.
  38. 38.Merel van den Broek. 2023. Chatgpt’s left-leaning liberal bias. University of Leiden.
  39. 39.Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and Bo Li. 2021. Adversarial glue: A multi-task benchmark for robustness evaluation of language models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  40. 40.Jindong Wang, HU Xixu, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Wei Ye, Haojun Huang, Xiubo Geng, et al. 2023. On the robustness of chatgpt: An adversarial and out-of-distribution perspective. In ICLR 2023 Workshop on Trustworthy and Reliable Large-Scale Machine Learning Models.
  41. 41.Xinpeng Wang, Chengzhi Hu, Bolei Ma, Paul Röttger, and Barbara Plank. 2024a. Look at the text: Instruction-tuned language models are more robust multiple choice selectors than you think. arXiv preprint arXiv:2404.08382.
  42. 42.Xinpeng Wang, Bolei Ma, Chengzhi Hu, Leon Weber-Genzel, Paul Röttger, Frauke Kreuter, Dirk Hovy, and Barbara Plank. 2024b. " my answer is c": First-token probabilities do not match text answers in instruction-tuned language models. arXiv preprint arXiv:2402.14499.
  43. 43.Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023. Jailbroken: How does llm safety training fail? arXiv preprint arXiv:2307.02483.
  44. 44.Guohai Xu, Jiayi Liu, Ming Yan, Haotian Xu, Jinghui Si, Zhuoran Zhou, Peng Yi, Xing Gao, Jitao Sang, Rong Zhang, et al. 2023. Cvalues: Measuring the values of chinese large language models from safety to responsibility. arXiv preprint arXiv:2307.09705.
  45. 45.Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. 2024. (inthe)wildchat: 570k chatGPT interaction logs in the wild. In The Twelfth International Conference on Learning Representations.
  46. 46.Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. 2024a. Large language models are not robust multiple choice selectors. In The Twelfth International Conference on Learning Representations.
  47. 47.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric Xing, Joseph E. Gonzalez, Ion Stoica, and Hao Zhang. 2024b. LMSYS-chat-1m: A large-scale real-world LLM conversation dataset. In The Twelfth International Conference on Learning Representations.

Citation

MLA
Röttger, P., et al. “Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 15295–311, https://doi.org/10.18653/v1/2024.acl-long.816.
APA
Röttger, P., Hofmann, V., Pyatkin, V., Hinck, M., Kirk, H., Schütze, H., & Hovy, D. (2024). Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15295–15311. https://doi.org/10.18653/v1/2024.acl-long.816
Chicago
Röttger, P., V. Hofmann, V. Pyatkin, et al. 2024. “Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15295–311. https://doi.org/10.18653/v1/2024.acl-long.816.
Harvard
Röttger, P. et al. (2024) “Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 15295–15311. Available at: https://doi.org/10.18653/v1/2024.acl-long.816.
Vancouver
1. Röttger P, Hofmann V, Pyatkin V, Hinck M, Kirk H, Schütze H, Hovy D (2024) Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 15295–15311

BibTeX

@inproceedings{rottger-etal-2024-political,
    title = "Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models",
    author = {R{\"o}ttger, Paul  and
      Hofmann, Valentin  and
      Pyatkin, Valentina  and
      Hinck, Musashi  and
      Kirk, Hannah  and
      Schuetze, Hinrich  and
      Hovy, Dirk},
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.816/",
    doi = "10.18653/v1/2024.acl-long.816",
    pages = "15295--15311"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/