Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
Paul RöttgerValentin HofmannValentina PyatkinMusashi HinckHannah KirkHinrich SchützeDirk Hovy
Demonstrates that forced multiple-choice benchmarks like the Political Compass Test produce brittle, misleading measurements of language model political bias, arguing instead for evaluations grounded in realistic, open-ended user interactions.
As large language models (LLMs) are deployed widely across society, concerns have escalated regarding the subtle political and social biases they may exhibit. Much of the current literature evaluates these values using multiple-choice surveys adapted from human psychometrics, such as the 62-item Political Compass Test (PCT). However, these evaluations impose artificial constraints that do not reflect how real users interact with artificial intelligence, creating a critical gap between academic testing frameworks and real-world deployment risks.
The article systematically assesses whether forced multiple-choice questionnaires reliably measure values and opinions in LLMs. Specifically, it evaluates model behavior across unforced baseline settings, varying styles of forced-choice prompts, minimal prompt paraphrases, and realistic open-ended text generation tasks.
To conduct this evaluation, the authors performed a systematic literature review of prior PCT studies and ran controlled experiments using ten distinct commercial and open-source models (including GPT-4, GPT-3.5, Llama 2, Mistral, and Zephyr) across all 62 PCT propositions. They tested multiple prompting templates under deterministic decoding settings and utilized an automated classifier validated with human annotations (93.1% inter-annotator agreement) to evaluate stances in free-form responses.
The article highlights five key findings: First, in an unforced multiple-choice setting, LLMs consistently resist taking sides; models like Zephyr and newer GPT versions produced 0% valid multiple-choice selections, instead generating nuanced disclaimers or balanced arguments in 95% of invalid cases. Second, models respond unpredictably to forced-choice constraints: GPT-4 remained nearly completely immune to forced prompts, whereas Llama 2 models shut down when pressured with negative consequences. Third, minimal semantics-preserving paraphrases (such as changing "opinion" to "view") caused substantial coordinate swings—shifting GPT-3.5's placement by over 100% on both economic and libertarian axes, a divergence larger than the measured gap between major political figures. Fourth, in realistic open-ended tasks (such as drafting blog posts), models systematically disagreed with propositions they previously endorsed in multiple-choice settings on roughly one out of three questions. Fifth, across these open-ended shifts, models displayed a consistent drift toward more right-leaning, libertarian viewpoints.
These findings demonstrate that popular multiple-choice benchmarks act more like unstable "spinning arrows" than dependable measurement tools. Constrained questionnaires fail to capture inherent model values and obscure the nuanced neutrality models typically offer. For organizations and policymakers, relying on simplistic survey scores creates significant compliance and risk blind spots, as models can exhibit entirely different stances depending on minor prompt variations or application contexts.
The article recommends that practitioners abandon rigid multiple-choice surveys in favor of evaluations tailored to specific real-world use cases. Furthermore, any assessment of model values must include extensive robustness testing across prompt phrasings, and stakeholders should restrict claims about model biases to local, task-specific contexts rather than asserting global model alignment.
These conclusions are bounded by their primary focus on the PCT and behavioral black-box testing. While the results show high statistical confidence across tested models, the authors note that empirical evaluations cannot provide absolute behavioral guarantees, warranting cautious application across different downstream architectures.
- Paper: From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models, Shangbin Feng et al. (2023). This paper establishes the use of the Political Compass test to map LLM ideological leanings, providing the exact empirical paradigm that the source systematically critiques and re-evaluates.
- Paper: Whose Opinions Do Language Models Reflect?, Shibani Santurkar et al. (2023). It provides foundational methodology and findings on measuring public opinion and political values in language models via standardized survey benchmarks.
- Paper: Out of One, Many: Using Language Models to Simulate Human Samples, Lisa P. Argyle et al. (2022). It introduces early methods for evaluating political attitudes and public opinion distributions in language models through simulated survey responses.
- Paper: The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values, Hannah Kirk et al. (2023). It surveys the integration and evaluation of subjective human preferences and values in LLMs, framing the broader alignment challenges addressed by the source.
- Paper: Prompting is not a substitute for probability measurements in large language models, Jennifer Hu et al. (2023). It demonstrates how prompting formats systematically diverge from internal model representations, motivating the source's investigation into prompting artifacts in value evaluations.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). It shows that LLM responses and reasoning are easily swayed by subtle input and multiple-choice prompt perturbations, directly underpinning the source's findings on answer instability.
- Paper: A Survey on Evaluation of Large Language Models, Yu-Chu Chang et al. (2023). It maps the landscape of LLM evaluation benchmarks and their structural shortcomings across subjective and ethical tasks.
- Paper: Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text, Sebastian Gehrmann et al. (2023). It analyzes widespread methodological obstacles and validity failures in text evaluation benchmarks that motivate moving to more realistic paradigms.
- Paper: Unintended Impacts of LLM Alignment on Global Representation, Michael J. Ryan et al. (2024). It investigates how post-training alignment procedures alter model value distributions and global cultural representation across open and closed evaluation settings.
- Paper: Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs, Xuhui Zhou et al. (2024). It extends the critique of artificial LLM evaluations by demonstrating how simulated social interactions fail to capture realistic interactive dynamics.
- Paper: Dynamic Evaluation of Large Language Models by Meta Probing Agents, Kaijie Zhu et al. (2024). It proposes dynamic meta-probing agent protocols to evaluate model capabilities and overcome the rigid constraints of static multiple-choice benchmarks.
- Paper: Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?, Nishant Balepur et al. (2024). It provides a deeper probe into how LLMs exploit superficial shortcuts in multiple-choice question formats, complementing the source's critique of forced-choice evaluations.
- Paper: Reasoning Models Generate Societies of Thought, Junsol Kim et al. (2026). It analyzes how advanced reasoning models generate diverse viewpoints and internal dialogue, offering an alternative lens on unconstrained model perspectives.
