SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
Yilun Zhao Kaiyan Zhang Tiansheng Hu Arman Cohan
Presents a community-driven evaluation platform and benchmark containing over 20,000 human expert votes to evaluate how foundation models generate and judge complex scientific literature-grounded responses.
The rapid surge in scientific publishing makes it increasingly difficult for researchers to keep up with developments in their fields. While artificial intelligence models are increasingly used to analyze and synthesize scientific literature, evaluating their performance on open-ended, complex scientific tasks remains a significant challenge. Existing evaluation benchmarks are typically static, narrow in scope, and quickly become obsolete. Meanwhile, automated AI-based evaluators frequently fail to capture subtle domain-specific nuances, and relying purely on manual human expert grading is slow, costly, and difficult to scale.
The article introduces and evaluates SciArena, an open community platform designed to assess how well artificial intelligence models synthesize scientific literature and answer real-world research queries. It also introduces a benchmark to measure how accurately automated AI evaluators can replicate human expert judgments.
To conduct this evaluation, the platform integrates a multi-stage literature retrieval pipeline with 47 state-of-the-art AI models. When a user submits a scientific question, the system retrieves relevant academic passages from a database of over 100 million papers, generates two independent answers with citations using randomly selected models, and collects human preference votes. The platform gathered over 20,000 votes across four major disciplines: natural science, healthcare, engineering, and humanities and social sciences. The authors analyzed model rankings using statistical rating systems, tested for stylistic and length biases, and constructed a 2,000-example benchmark to test automated AI evaluators against human decisions.
The analysis produced several key findings. First, advanced reasoning-oriented systems lead overall performance, with o3, Claude-4.1-Opus, and GPT-5 securing the top three rankings across the evaluated models. Performance varied significantly by subject: o3 led in engineering, while GPT-5 ranked highest in natural sciences and social sciences. Second, leading open-source models demonstrated competitive capability, with GPT-OSS-120B and Deepseek-R1-0528 ranking in the top ten and outperforming several commercial alternatives. Third, human scientific reviewers prioritize citation quality over volume. While general-purpose platforms often suffer from users favoring longer text or higher reference counts, reviewers on SciArena showed a clear preference for correct attribution to relevant literature rather than sheer citation quantity. Finally, automated AI evaluators struggled to judge scientific answers accurately: the best-performing evaluator achieved only 65.1% accuracy compared to human expert preferences, falling well below alignment rates seen on general-purpose benchmarks.
These findings indicate that general-purpose automated evaluation tools are currently insufficient for reliable decision-making in knowledge-intensive domains. Deploying AI for scholarly review requires robust grounding in literature, as models frequently exhibit failure modes such as conflicting with cited evidence, misunderstanding specialized terminology, or providing vague summaries. Furthermore, the strong showing of select open-source models shows that organizations can achieve high-tier literature synthesis performance without relying entirely on costly proprietary services.
Organizations developing or deploying AI for scientific discovery should adopt multi-stage retrieval pipelines and emphasize citation verification rather than relying solely on raw model reasoning. Teams should also refrain from using automated AI judges as the sole quality gate for specialized tasks until more reliable evaluation methods are developed. Decision-makers should combine community-driven expert feedback with continuous benchmark updates to monitor model performance over time.
The results should be interpreted with awareness of certain operational limitations. The evaluation platform currently excludes some earlier model versions and cannot yet interface with certain agent-based research tools that impose usage caps or lack programming interfaces. Nonetheless, the high rate of annotator agreement and strong consistency over time provide high confidence in the overall rankings and findings.
- Paper: Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference, Wei-Lin Chiang et al. (2024). Chatbot Arena establishes the live pairwise human-preference evaluation model that SciArena adapts for scientific answers and model rankings.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). This survey maps the biases and reliability challenges of LLM-as-a-judge methods that SciArena tests directly against expert preferences.
- Paper: Enabling Large Language Models to Generate Text with Citations, Tianyu Gao et al. (2023). ALCE defines measurable citation-quality criteria for retrieval-grounded answers, a key concern in SciArena’s evaluation of scientific responses.
- Paper: A Critical Evaluation of Evaluations for Long-form Question Answering, Fangyuan Xu et al. (2023). This study shows why expert judgments of long-form answers cannot be replaced by generic automatic metrics, framing SciArena’s human-versus-automated evaluation.
- Paper: LitSearch: A Retrieval Benchmark for Scientific Literature Search, Anirudh Ajith et al. (2024). LitSearch’s evaluation of scientific-literature retrieval provides context for SciArena’s pipeline that retrieves research passages before generating answers.
- Paper: JudgeBench: A Benchmark for Evaluating LLM-Based Judges, Sijun Tan et al. (2025). JudgeBench extends the problem SciArena identifies by testing whether automated judges can detect subtle correctness failures rather than merely match human preferences.
- Paper: Benchmarking AI Agents for Addressing Scientific Challenges Across Scales, Tianyu Liu et al. (2026). SciAgentArena carries scientific AI evaluation into multi-step research workflows, extending SciArena’s assessment beyond literature-grounded answers to executable tasks.
- Paper: A Benchmark for Deep Information Synthesis, Debjit Paul et al. (2026). DEEPSYNTH extends the study of complex synthesis evaluation to multi-source research tasks, building on the challenge of assessing open-ended answers without simple ground truth.
