SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks

Yilun Zhao Kaiyan Zhang Tiansheng Hu Arman Cohan

article2025NeurIPS20 citations

Presents a community-driven evaluation platform and benchmark containing over 20,000 human expert votes to evaluate how foundation models generate and judge complex scientific literature-grounded responses.

Listen

The rapid surge in scientific publishing makes it increasingly difficult for researchers to keep up with developments in their fields. While artificial intelligence models are increasingly used to analyze and synthesize scientific literature, evaluating their performance on open-ended, complex scientific tasks remains a significant challenge. Existing evaluation benchmarks are typically static, narrow in scope, and quickly become obsolete. Meanwhile, automated AI-based evaluators frequently fail to capture subtle domain-specific nuances, and relying purely on manual human expert grading is slow, costly, and difficult to scale.

The article introduces and evaluates SciArena, an open community platform designed to assess how well artificial intelligence models synthesize scientific literature and answer real-world research queries. It also introduces a benchmark to measure how accurately automated AI evaluators can replicate human expert judgments.

To conduct this evaluation, the platform integrates a multi-stage literature retrieval pipeline with 47 state-of-the-art AI models. When a user submits a scientific question, the system retrieves relevant academic passages from a database of over 100 million papers, generates two independent answers with citations using randomly selected models, and collects human preference votes. The platform gathered over 20,000 votes across four major disciplines: natural science, healthcare, engineering, and humanities and social sciences. The authors analyzed model rankings using statistical rating systems, tested for stylistic and length biases, and constructed a 2,000-example benchmark to test automated AI evaluators against human decisions.

The analysis produced several key findings. First, advanced reasoning-oriented systems lead overall performance, with o3, Claude-4.1-Opus, and GPT-5 securing the top three rankings across the evaluated models. Performance varied significantly by subject: o3 led in engineering, while GPT-5 ranked highest in natural sciences and social sciences. Second, leading open-source models demonstrated competitive capability, with GPT-OSS-120B and Deepseek-R1-0528 ranking in the top ten and outperforming several commercial alternatives. Third, human scientific reviewers prioritize citation quality over volume. While general-purpose platforms often suffer from users favoring longer text or higher reference counts, reviewers on SciArena showed a clear preference for correct attribution to relevant literature rather than sheer citation quantity. Finally, automated AI evaluators struggled to judge scientific answers accurately: the best-performing evaluator achieved only 65.1% accuracy compared to human expert preferences, falling well below alignment rates seen on general-purpose benchmarks.

These findings indicate that general-purpose automated evaluation tools are currently insufficient for reliable decision-making in knowledge-intensive domains. Deploying AI for scholarly review requires robust grounding in literature, as models frequently exhibit failure modes such as conflicting with cited evidence, misunderstanding specialized terminology, or providing vague summaries. Furthermore, the strong showing of select open-source models shows that organizations can achieve high-tier literature synthesis performance without relying entirely on costly proprietary services.

Organizations developing or deploying AI for scientific discovery should adopt multi-stage retrieval pipelines and emphasize citation verification rather than relying solely on raw model reasoning. Teams should also refrain from using automated AI judges as the sole quality gate for specialized tasks until more reliable evaluation methods are developed. Decision-makers should combine community-driven expert feedback with continuous benchmark updates to monitor model performance over time.

The results should be interpreted with awareness of certain operational limitations. The evaluation platform currently excludes some earlier model versions and cannot yet interface with certain agent-based research tools that impose usage caps or lack programming interfaces. Nonetheless, the high rate of annotator agreement and strong consistency over time provide high confidence in the overall rankings and findings.

Cover for SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks

Abstract

We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific literature understanding and synthesis, SciArena engages the research community directly, following the Chatbot Arena evaluation approach of community voting on model comparisons. By leveraging collective intelligence, SciArena offers a community-driven evaluation of model performance on open-ended scientific tasks that demand literature-grounded, long-form responses. The platform currently supports 47 foundation models and has collected over 20,000 votes from human researchers across diverse scientific domains. Our analysis of the data collected so far confirms its high quality. We discuss the results and insights based on the model ranking leaderboard. To further promote research in building model-based automated evaluation systems for literature tasks, we release SciArena-Eval, a meta-evaluation benchmark based on collected preference data. It measures the accuracy of models in judging answer quality by comparing their pairwise assessments with human votes. Our experiments highlight the benchmark's challenges and emphasize the need for more reliable automated evaluation methods.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 SciArena Platform
  • 3.1 Platform Design
  • 3.2 User Study on SciArena vs. Commercial Platforms for Scientific Literature Tasks
  • 3.3 Leaderboard Ranking with Elo Rating
  • 4 SciArena Data
  • 4.1 SciArena Human Preference Data Collection
  • 4.2 Data Analysis
  • 4.3 Quality Assessment of Collected Human Votes
  • 5 SciArena Leaderboard Analysis
  • 5.1 Main Results
  • 5.2 Preference Analyses in SciArena Evaluation
  • 5.3 Case Analysis
  • 6 SciArena-Eval for Evaluating Model-based Evaluators
  • 6.1 SciArena-Eval Benchmark
  • 6.2 Experiments on SciArena-Eval
  • 7 Discussion
  • References
  • A SciArena Platform
  • A.1 User Study Results
  • A.2 Online Elo Rating System
  • A.3 Prompts for Model Response Generation and Postprocessing
  • A.4 Configurations of Evaluated Foundation Models, As of Jan 11, 2026
  • A.5 Detecting Anomalous Users
  • B SciArena Data
  • B.1 Annotator Biographies of 102 Researchers Involved in Initial Data Collection
  • B.2 SciArena Data Analysis
  • B.3 Prompts for Question Category Classification
  • C SciArena Leaderboard Results
  • C.1 In-depth Analysis
  • C.2 Analysis of Citation Features in SciArena Evaluation
  • C.3 Analysis of o3 Model
  • C.3.1 Examples Illustrating o3’s Detailed Elaboration on Cited Papers
  • C.3.2 Examples Illustrating o3’s More Professional & Precise Terminology
  • C.3.3 Examples Illustrating o3’s Clear Structured Presentation
  • C.3.4 Examples Illustrating o3’s More Comprehensive Coverage
  • C.4 Model Failure Case Analysis
  • C.4.1 Examples Illustrating Errors of Failure to Answer the Question
  • C.4.2 Examples Illustrating Errors of Conflict with Cited Papers
  • C.4.3 Examples Illustrating Errors of Lack of Details
  • C.4.4 Examples Illustrating Errors of Misunderstanding Terminology
  • C.4.5 Examples Illustrating Errors of Incoherent Structure
  • D SciArena-Eval Meta-Evaluation Experiment
  • D.1 Evaluation Prompts

Knowls

  1. Knowl 1 — SciArena’s community preference evaluation workflow

    model/method

    SciArena evaluates scientific literature-grounded answers through blinded pairwise comparisons. A submitted question first passes a harmful-content check using OpenAI’s omni-moderation-latest; two models are then randomly selected from the available model pool and independently produce citation-attributed answers using the question and retrieved paper contexts. Users choose the answer that better meets their information need and may provide a written justification. To reduce presentation-based preference effects, models are prompted to use plain text without Markdown, and a strong language model postprocesses responses to standardize citation formatting and remove Markdown before display. The postprocessor changed from GPT-4o to GPT-4.1 on April 20. As of January 11, 2026, the platform supported 47 models: 25 proprietary and 22 open-source.

  2. Knowl 2 — SciArena’s scientific literature retrieval pipeline

    model/method

    For each user question, SciArena uses a query-decomposition language model to rephrase the question for different search endpoints and extract metadata filters such as publication year, author, or venue. The system uses Semantic Scholar’s continuously updated corpus, with more than 100 million paper abstracts and about 11.7 million full-text papers indexed for snippet search. It retrieves up to 40 full-text snippets and 20 abstracts, then applies a cross-encoder reranker and supplies the 30 highest-ranked results as context to the two answer-generating models. The query-decomposition model was GPT-4o before April 20 and GPT-4.1 afterward.

  3. Knowl 3 — Human preference data and question distribution

    data/table

    SciArena reports 20,832 votes across Natural Science, Healthcare, Humanities and Social Sciences, and Engineering. The vote outcomes were 9,107 preferences for answer A, 9,392 for answer B, 1,746 ties, and 587 judgments that both answers were bad. The initial collection involved 102 researchers; each had authored at least two peer-reviewed papers and completed a one-hour training session. Community votes were included in the leaderboard only after users passed anomaly checks and consented to data collection. GPT-4.1 classification of the collected questions found that 32.94% sought conceptual explanations, 25.36% assessed the state of the art, 23.20% asked about challenges or limitations, 8.76% concerned methodology, 4.78% sought papers, and 4.96% fell into other categories.

  4. Knowl 4 — Bradley–Terry estimation for the SciArena leaderboard

    model/method

    SciArena estimates model strengths from all pairwise votes using a Bradley–Terry logistic model rather than updating ratings match by match. For nn comparisons among MM models, comparison ii has a feature vector Xi∈RMX_i\in\mathbb{R}^M: its entry is +1+1 for the first model, −1-1 for the second, and 00 for all other models. The binary outcome YiY_i equals 1 when the first model wins and 0 otherwise. The strength vector β^∈RM\hat{\beta}\in\mathbb{R}^M minimizes average binary cross-entropy: β^=arg⁡min⁡β∈RM1n∑i=1nCE⁡(σ(Xi⊤β),Yi)\hat{\beta}=\arg\min_{\beta\in\mathbb{R}^M}\frac{1}{n}\sum_{i=1}^{n}\operatorname{CE}(\sigma(X_i^\top\beta),Y_i), where σ\sigma is the sigmoid function and CE⁡\operatorname{CE} is cross-entropy. Because the model has no tie outcome, each tie vote is duplicated and split evenly between a first-model win and a second-model win. SciArena uses 100 bootstrap resamples to estimate confidence intervals for ratings.

  5. Knowl 5 — Leaderboard results across models and scientific disciplines

    empirical result

    Among models with at least 100 votes, the highest overall Bradley–Terry ratings as of January 11, 2026, were o3 (1151.4), Claude-4.1-Opus (1126.4), GPT-5 (1122.5), Gemini-3-Pro-Preview (1086.3), and GPT-5.1 (1079.7). The leading model varied by discipline: GPT-5.1 had the highest Natural Science rating (1161.0), Gemini-2.5-Pro led Healthcare (1162.2), Claude-4.1-Opus led Humanities and Social Sciences (1173.3), and o3 led Engineering (1176.7). Among open-source models, GPT-OSS-120B ranked eighth overall with 1036.5, and DeepSeek-R1-0528 ranked ninth with 1041.8. The results therefore show both strong overall performance by several frontier proprietary models and variation in leaders across scientific fields.

  6. Knowl 6 — Human-vote reliability checks

    empirical result

    SciArena assessed expert-vote reliability using 400 examples: 100 randomly selected questions from each of four disciplines, each independently evaluated by a second expert with a closely aligned research background. Average inter-annotator agreement was 0.82 accuracy and 0.76 weighted Cohen’s κ\kappa; average repeat-annotation self-consistency, measured after at least two weeks, was 0.94 accuracy and 0.91 weighted κ\kappa. By discipline, inter-annotator accuracy / κ\kappa and self-consistency accuracy / κ\kappa were Natural Science 0.82 / 0.76 and 0.94 / 0.91; Healthcare 0.87 / 0.82 and 0.91 / 0.89; Humanities and Social Sciences 0.78 / 0.70 and 0.96 / 0.94; Engineering 0.82 / 0.75 and 0.93 / 0.91. A separate audit sampled 150 community-voted examples; reviewers could not assess 17 because of unfamiliarity with the topic. Of the remaining 133, reviewers judged 98 (73.7%) clearly reasonable, 23 (17.3%) acceptable but ambiguous, and 12 (9.0%) likely incorrect.

  7. Knowl 7 — SciArena-Eval benchmark construction

    model/method

    SciArena-Eval is a meta-evaluation set for measuring whether automated evaluators choose the same answer as human voters on scientific literature questions. It contains 2,000 pairwise comparisons sampled from SciArena votes: 500 examples from each of four scientific disciplines, with 250 examples per discipline in which humans preferred answer A and 250 in which they preferred answer B. Tie votes are excluded because they do not identify a superior answer. An evaluated system receives the question and both citation-attributed responses, selects the better response, and is scored by accuracy against the human preference.

  8. Knowl 8 — Automated evaluators show limited agreement with human preferences

    empirical result

    On SciArena-Eval, the best tested pairwise evaluator, o3, achieved 65.1% accuracy, compared with 50.0% for random choice. Accuracies were: o4-mini 64.8%; GPT-5 with high reasoning 63.2%; GPT-4.1 61.9%; DeepSeek-R1 61.2%; GPT-5 with medium reasoning 60.6%; DeepSeek-V3 60.5%; GPT-4.1-mini 60.5%; Gemini-2.5-Pro-Preview 60.3%; Claude-3.7-Sonnet 59.2%; Qwen3-32B 58.1%; Gemini-2.5-Flash-Preview 57.8%; Llama-4-Scout 57.7%; and Llama-4-Maverick 57.5%. Reasoning-oriented variants outperformed corresponding non-reasoning models in the reported comparisons: o4-mini exceeded GPT-4.1 by 2.9 percentage points, and DeepSeek-R1 exceeded DeepSeek-V3 by 0.7 points. Overall, even the strongest evaluator remained substantially below the higher-than-70% alignment reported for pairwise evaluation on general-purpose benchmarks.

  9. Knowl 9 — Citation attribution and response length affect preferences

    empirical result

    An analysis of 3,000 SciArena voting instances used a Bradley–Terry model with response-style features. The o4-mini model classified citations as supporting, irrelevant, or contradicting; irrelevant and contradicting citations were combined for the preference analysis. The estimated coefficient for citation count was modestly positive (γ=0.039\gamma=0.039), while supporting citations had a positive coefficient (γ=0.155\gamma=0.155) and irrelevant or contradicting citations had a negative coefficient (γ=−0.154\gamma=-0.154). The response-length coefficient was γ=0.141\gamma=0.141. These estimates indicate that citation quantity alone was less influential than whether citations supported the associated claims; the reported length coefficient was also lower than coefficients cited for Chatbot Arena (0.25), Vision Arena (0.27), and Search Arena (0.33).

  10. Knowl 10 — Qualitative analyses identify model strengths and recurring failures

    empirical result

    A human review of 200 voting examples comparing o3 with other high-performing models, focused on Engineering, identified four recurring strengths in o3’s answers: more detailed explanation of cited papers, more precise and professional terminology, clearer organization, and more comprehensive coverage, particularly for questions about challenges and the state of the art. A separate review sampled 100 difficult examples, selected either because users judged both answers poor or because a top-three model lost to another model. Reviewers identified five recurring failure types: not answering the question, making claims that conflict with cited papers, providing insufficient detail, misunderstanding terminology, and presenting information with incoherent structure.

Coverage note — The four-researcher usability study comparing SciArena with commercial search and deep-research platforms, and the detailed per-model configuration inventory, are omitted because they are secondary to the platform, dataset, ranking, and meta-evaluation contributions.

References

  1. 1.Ai2. Olmo 3: Charting a path through the model flow to lead open-source ai. https://allenai.org/blog/olmo3, November 2025.
  2. 2.Anirudh Ajith, Mengzhou Xia, Alexis Chevalier, Tanya Goyal, Danqi Chen, and Tianyu Gao. LitSearch: A retrieval benchmark for scientific literature search. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 15068–15083, Miami, Florida, USA, November 2024. Association for Computational Linguistics.
  3. 3.Bridger Waleed Ammar, Dirk Groeneveld, Chandra Bhagavatula, Iz Beltagy, Miles Crawford, Doug Downey, Jason Dunkelberger, Ahmed Elgohary, Sergey Feldman, Vu A. Ha, Rodney Michael Kinney, Sebastian Kohlmeier, Kyle Lo, Tyler C. Murray, Hsu-Han Ooi, Matthew E. Peters, Joanna L. Power, Sam Skjonsberg, Lucy Lu Wang, Christopher Wilhelm, Zheng Yuan, Madeleine van Zuylen, and Oren Etzioni. Construction of the literature graph in semantic scholar. In North American Chapter of the Association for Computational Linguistics, 2018.
  4. 4.Anthropic. Claude 3.7 sonnet and claude code. https://www.anthropic.com/news/claude-3-7-sonnet, February 2025.
  5. 5.Anthropic. Claude opus 4.1. https://www.anthropic.com/news/claude-opus-4-1/, August 2025.
  6. 6.Anthropic. Introducing claude 4. https://www.anthropic.com/news/claude-4, May 2025.
  7. 7.Akari Asai, Jacqueline He, Rulin Shao, Weijia Shi, Amanpreet Singh, Joseph Chee Chang, Kyle Lo, Luca Soldaini, Sergey Feldman, Mike D’arcy, David Wadden, Matt Latzke, Minyang Tian, Pan Ji, Shengyan Liu, Hao Tong, Bohao Wu, Yanyu Xiong, Luke Zettlemoyer, Graham Neubig, Dan Weld, Doug Downey, Wen tau Yih, Pang Wei Koh, and Hannaneh Hajishirzi. Openscholar: Synthesizing scientific literature with retrieval-augmented lms, 2024.
  8. 8.Akari Asai, Jacqueline He, Rulin Shao, Weijia Shi, Amanpreet Singh, Joseph Chee Chang, Kyle Lo, Luca Soldaini, Sergey Feldman, Mike D’Arcy, David Wadden, Matt Latzke, Minyang Tian, Pan Ji, Shengyan Liu, Hao Tong, Bohao Wu, Yanyu Xiong, Luke S. Zettlemoyer, Graham Neubig, Dan Weld, Doug Downey, Wen tau Yih, Pang Wei Koh, and Hanna Hajishirzi. Openscholar: Synthesizing scientific literature with retrieval-augmented lms. ArXiv, abs/2411.14199, 2024.
  9. 9.Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. Re-evaluating evaluation in text summarization. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9347–9359, Online, November 2020. Association for Computational Linguistics.
  10. 10.Ralph Allan Bradley and Milton E Terry. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345, 1952.
  11. 11.Isabel Cachola, Kyle Lo, Arman Cohan, and Daniel Weld. TLDR: Extreme summarization of scientific documents. In Trevor Cohn, Yulan He, and Yang Liu, editors, Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4766–4777, Online, November 2020. Association for Computational Linguistics.
  12. 12.Hui Chen, Miao Xiong, Yujie Lu, Wei Han, Ailin Deng, Yufei He, Jiaying Wu, Yibo Li, Yue Liu, and Bryan Hooi. Mlr-bench: Evaluating ai agents on open-ended machine learning research, 2025.
  13. 13.Cheng-Han Chiang and Hung yi Lee. Can large language models be an alternative to human evaluations? In Annual Meeting of the Association for Computational Linguistics, 2023.
  14. 14.Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. Chatbot arena: An open platform for evaluating LLMs by human preference. In Forty-first International Conference on Machine Learning, 2024.
  15. 15.Christopher Chou, Lisa Dunlap, Koki Mashita, Krishna Mandal, Trevor Darrell, Ion Stoica, Joseph Gonzalez, and Wei-Lin Chiang. Visionarena: 230k real world user-vlm conversations with preference labels. ArXiv, abs/2412.08687, 2024.
  16. 16.Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel Weld. SPECTER: Document-level representation learning using citation-informed transformers. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2270–2282, Online, July 2020. Association for Computational Linguistics.
  17. 17.Consensus. Consensus – ai for research, 2024. Accessed: 2025-05-02.
  18. 18.Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, and Matt Gardner. A dataset of information-seeking questions and answers anchored in research papers. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou, editors, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4599–4610, Online, June 2021. Association for Computational Linguistics.
  19. 19.Google DeepMind. Gemini – deep research mode, 2024. Accessed: 2025-05-02.
  20. 20.DeepSeek-AI. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, December 2024.
  21. 21.DeepSeek-AI. Deepseek-r1-0528 release. https://api-docs.deepseek.com/news/news250528/, May 2025.
  22. 22.DeepSeek-AI. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.
  23. 23.Yann Dubois, Percy Liang, and Tatsunori Hashimoto. Length-controlled alpacaeval: A simple debiasing of automatic evaluators. In First Conference on Language Modeling, 2024.
  24. 24.Lisa Dunlap, Krishna Mandal, Trevor Darrell, Jacob Steinhardt, and Joseph E Gonzalez. Vibecheck: Discover and quantify qualitative differences in large language models, 2025.
  25. 25.Elicit. Elicit – the ai research assistant, 2024. Accessed: 2025-05-02.
  26. 26.Arpad E Elo. The proposed uscf rating system, its development, theory, and applications. Chess life, 22(8):242–247, 1967.
  27. 27.Ronald Aylmer Fisher. Statistical methods for research workers. Number 5. Oliver and Boyd, 1928.
  28. 28.Google. Gemini 2.5: Our most intelligent ai model. https://blog.google/technology/google-deepmind/gemini-model-thinking-updates-march-2025//, March 2025.
  29. 29.Google. Gemini 2.5 pro preview: even better coding performance. https://developers.googleblog.com/en/gemini-2-5-pro-io-improved-coding-performance/, April 2025.
  30. 30.Google. Gemini 3 pro: Best for complex tasks and bringing creative concepts to life. https://deepmind.google/models/gemini/pro/, November 2025.
  31. 31.Google. Start building with gemini 2.5 flash. https://developers.googleblog.com/en/start-building-with-gemini-25-flash/, April 2025.
  32. 32.Dongfu Jiang, Max Ku, Tianle Li, Yuansheng Ni, Shizhuo Sun, Rongqi Fan, and Wenhu Chen. GenAI arena: An open evaluation platform for generative models. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024.
  33. 33.Tetsu Kasanishi, Masaru Isonuma, Junichiro Mori, and Ichiro Sakata. SciReviewGen: A large-scale dataset for automatic literature review generation. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Findings of the Association for Computational Linguistics: ACL 2023, pages 6695–6715, Toronto, Canada, July 2023. Association for Computational Linguistics.
  34. 34.Rodney Kinney, Chloe Anastasiades, Russell Authur, Iz Beltagy, Jonathan Bragg, Alexandra Buraczynski, Isabel Cachola, Stefan Candra, Yoganand Chandrasekhar, Arman Cohan, Miles Crawford, Doug Downey, Jason Dunkelberger, Oren Etzioni, Rob Evans, Sergey Feldman, Joseph Gorney, David Graham, Fangzhou Hu, Regan Huff, Daniel King, Sebastian Kohlmeier, Bailey Kuehl, Michael Langan, Daniel Lin, Haokun Liu, Kyle Lo, Jaron Lochner, Kelsey MacMillan, Tyler Murray, Chris Newell, Smita Rao, Shaurya Rohatgi, Paul Sayre, Zejiang Shen, Amanpreet Singh, Luca Soldaini, Shivashankar Subramanian, Amber Tanaka, Alex D. Wade, Linda Wagner, Lucy Lu Wang, Chris Wilhelm, Caroline Wu, Jiangjiang Yang, Angele Zamarron, Madeleine Van Zuylen, and Daniel S. Weld. The semantic scholar open data platform, 2025.
  35. 35.Yoonjoo Lee, Yoonjoo Lee, Kyungjae Lee, Sunghyun Park, Dasol Hwang, Jaehyeon Kim, Ho Hin Lee, and Moontae Lee. Qasa: Advanced question answering on scientific articles. In International Conference on Machine Learning, 2023.
  36. 36.Chuhan Li, Ziyao Shangguan, Yilun Zhao, Deyuan Li, Yixin Liu, and Arman Cohan. M3SciQA: A multi-modal multi-document scientific QA benchmark for evaluating foundation models. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024, pages 15419–15446, Miami, Florida, USA, November 2024. Association for Computational Linguistics.
  37. 37.Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline, 2024.
  38. 38.Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Alpacaeval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval, 5 2023.
  39. 39.Gabrielle Kaili-May Liu, Bowen Shi, Avi Caciularu, Idan Szpektor, and Arman Cohan. Mdcure: A scalable pipeline for multi-document instruction-following, 2025.
  40. 40.Haokun Liu, Sicong Huang, Jingyu Hu, Yangqiaoyu Zhou, and Chenhao Tan. Hypobench: Towards systematic and principled benchmarking for hypothesis generation, 2025.
  41. 41.Hongyi Liu, Qingyun Wang, Payam Karisani, and Heng Ji. Named entity recognition under domain shift via metric learning for life sciences. In Kevin Duh, Helena Gomez, and Steven Bethard, editors, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 1–21, Mexico City, Mexico, June 2024. Association for Computational Linguistics.
  42. 42.Yixin Liu, Alex Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, and Dragomir Radev. Revisiting the gold standard: Grounding summarization evaluation with robust human evaluation. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4140–4170, Toronto, Canada, July 2023. Association for Computational Linguistics.
  43. 43.Yixin Liu, Alexander Fabbri, Jiawen Chen, Yilun Zhao, Simeng Han, Shafiq Joty, Pengfei Liu, Dragomir Radev, Chien-Sheng Wu, and Arman Cohan. Benchmarking generation and evaluation capabilities of large language models for instruction controllable summarization. In Kevin Duh, Helena Gomez, and Steven Bethard, editors, Findings of the Association for Computational Linguistics: NAACL 2024, pages 4481–4501, Mexico City, Mexico, June 2024. Association for Computational Linguistics.
  44. 44.Yixin Liu, Kejian Shi, Alexander Fabbri, Yilun Zhao, PeiFeng Wang, Chien-Sheng Wu, Shafiq Joty, and Arman Cohan. ReIFE: Re-evaluating instruction-following evaluation. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors, NAACL. Association for Computational Linguistics, April 2025.
  45. 45.Yao Lu, Yue Dong, and Laurent Charlin. Multi-XScience: A large-scale dataset for extreme multi-document summarization of scientific articles. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8068–8074, Online, November 2020. Association for Computational Linguistics.
  46. 46.Yujie Lu, Dongfu Jiang, Wenhu Chen, William Yang Wang, Yejin Choi, and Bill Yuchen Lin. Wildvision: Evaluating vision-language models in the wild with human preferences. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024.
  47. 47.Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, and Dan Roth. ExpertQA: Expert-curated questions and attributed answers. In Kevin Duh, Helena Gomez, and Steven Bethard, editors, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 3025–3045, Mexico City, Mexico, June 2024. Association for Computational Linguistics.
  48. 48.Meta AI. The llama 4 herd: The beginning of a new era of natively multimodal ai innovation. https://ai.meta.com/blog/llama-4-multimodal-intelligence/, April 2025.
  49. 49.MINIMAX. Minimax-m1, the world’s first open-source, large-scale, hybrid-attention reasoning model. https://www.minimax.io/news/minimaxm1/, June 2025.
  50. 50.Ministral AI. Medium is the new large. https://mistral.ai/news/mistral-medium-3/, May 2025.
  51. 51.Mihran Miroyan, Tsung-Han Wu, Logan King, Tianle Li, Jiayi Pan, Xinyan Hu, Wei-Lin Chiang, Anastasios N. Angelopoulos, Trevor Darrell, Narges Norouzi, and Joseph E. Gonzalez. Search arena: Analyzing search-augmented llms, 2025.
  52. 52.Mistral AI Team. Mistral small 3. https://mistral.ai/news/mistral-small-3/, January 2025.
  53. 53.MoonshotAI. Kimi k2: Open agentic intelligence. https://moonshotai.github.io/Kimi-K2//, July 2025.
  54. 54.mrfakename, Vaibhav Srivastav, Clémentine Fourrier, Lucain Pouget, Yoach Lacombe, main, and Sanchit Gandhi. Text to speech arena. https://huggingface.co/spaces/TTS-AGI/TTS-Arena, 2024.
  55. 55.OpenAI. Chatgpt – deep research mode, 2024. Accessed: 2025-05-02.
  56. 56.OpenAI. Gpt-5.1: A smarter, more conversational chatgpt. https://openai.com/index/gpt-5-1/, November 2025.
  57. 57.OpenAI. Introducing gpt-4.1 in the api. https://openai.com/index/gpt-4-1/, April 2025.
  58. 58.OpenAI. Introducing gpt-5. https://openai.com/index/introducing-gpt-5/, August 2025.
  59. 59.OpenAI. Introducing gpt-oss. https://openai.com/index/introducing-gpt-oss//, August 2025.
  60. 60.OpenAI. Introducing openai o3 and o4-mini. https://openai.com/index/introducing-o3-and-o4-mini/, April 2025.
  61. 61.Perplexity. Perplexity ai – ask anything, 2024. Accessed: 2025-05-02.
  62. 62.Qwen Team. Qwen3: Think deeper, act faster, April 2025.
  63. 63.Qwen Team. Qwq-32b: Embracing the power of reinforcement learning, March 2025.
  64. 64.Scale AI. Advancing frontier model evaluation. Scale AI Blog, 2024.
  65. 65.Scite. Scite – smart citations for research, 2024. Accessed: 2025-05-02.
  66. 66.Aamir Shakir, Darius Koenig, Julius Lipp, and Sean Lee. Boost your search with the crispy mixedbread rerank models, 2024.
  67. 67.Chenglei Si, Diyi Yang, and Tatsunori Hashimoto. Can LLMs generate novel research ideas? a large-scale human study with 100+ NLP researchers. In The Thirteenth International Conference on Learning Representations, 2025.
  68. 68.Amanpreet Singh, Joseph Chee Chang, Chloe Anastasiades, Dany Haddad, Aakanksha Naik, Amber Tanaka, Angele Zamarron, Cecile Nguyen, Jena D. Hwang, Jason Dunkleberger, Matt Latzke, Smita Rao, Jaron Lochner, Rob Evans, Rodney Kinney, Daniel S. Weld, Doug Downey, and Sergey Feldman. Ai2 scholar qa: Organized literature synthesis with attribution, 2025.
  69. 69.Amanpreet Singh, Joseph Chee Chang, Chloe Anastasiades, Dany Haddad, Aakanksha Naik, Amber Tanaka, Angele Zamarron, Cecile Nguyen, Jena D. Hwang, Jason Dunkleberger, Matt Latzke, Smita R Rao, Jaron Lochner, Rob Evans, Rodney Kinney, Daniel S. Weld, Doug Downey, and Sergey Feldman. Ai2 scholar qa: Organized literature synthesis with attribution. 2025.
  70. 70.Amanpreet Singh, Mike D’Arcy, Arman Cohan, Doug Downey, and Sergey Feldman. SciRepEval: A multi-format benchmark for scientific document representations. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5548–5566, Singapore, December 2023. Association for Computational Linguistics.
  71. 71.Shivalika Singh, Yiyang Nan, Alex Wang, Daniel D’souza, Sayash Kapoor, A. Ustun, Sanmi Koyejo, Yuntian Deng, Shayne Longpre, Noah Smith, Beyza Hilal Ermi¸s, Marzieh Fadaee, and Sara Hooker. The leaderboard illusion. 2025.
  72. 72.Shruti Singh, Nandan Sarkar, and Arman Cohan. SciDQA: A deep reading comprehension dataset over scientific papers. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 20908–20923, Miami, Florida, USA, November 2024. Association for Computational Linguistics.
  73. 73.Annalisa Szymanski, Noah Ziems, Heather A. Eicher-Miller, Toby Jia-Jun Li, Meng Jiang, and Ronald A. Metoyer. Limitations of the llm-as-a-judge approach for evaluating llm outputs in expert knowledge tasks, 2024.
  74. 74.Aryan Vichare, Anastasios N. Angelopoulos, Wei-Lin Chiang, Kelly Tang, and Luca Manolache. Webdev arena: A live llm leaderboard for web app development, 2025.
  75. 75.David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. Fact or fiction: Verifying scientific claims. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7534–7550, Online, November 2020. Association for Computational Linguistics.
  76. 76.Bowen Wang, Xinyuan Wang, Jiaqi Deng, Tianbao Xie, Ryan Li, Yanzhe Zhang, Gavin Li, Toh Jing Hua, Ion Stoica, Wei-Lin Chiang, Diyi Yang, Yu Su, Yi Zhang, Zhiguo Wang, Victor Zhong, and Tao Yu. Computer agent arena: Compare and test computer use agents on crowdsourced real-world tasks, 2025.
  77. 77.Chengye Wang, Yifei Shen, Zexi Kuang, Arman Cohan, and Yilun Zhao. Sciver: Evaluating foundation models for multimodal scientific claim verification, 2025.
  78. 78.Qingyun Wang, Doug Downey, Heng Ji, and Tom Hope. SciMON: Scientific inspiration machines optimized for novelty. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 279–299, Bangkok, Thailand, August 2024. Association for Computational Linguistics.
  79. 79.xAI. Grok 3 beta — the age of reasoning agents. https://x.ai/news/grok-3, February 2025.
  80. 80.xAI. Grok 4. https://x.ai/news/grok-4, July 2025.
  81. 81.Zhijian Xu, Yilun Zhao, Manasi Patwardhan, Lovekesh Vig, and Arman Cohan. Can LLMs identify critical limitations within scientific research? a systematic evaluation on AI research papers. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 20652–20706, Vienna, Austria, July 2025. Association for Computational Linguistics.
  82. 82.Nithik Yekollu, Arth Bohra, Ashwin Chirumamilla, Kai Wen, Sai Kolasani Wei-Lin Chiang, Anastasios Angelopoulos, Joseph E. Gonzalez, Ion Stoica, and Shishir G. Patil. Agent arena. 2024.
  83. 83.You.com. You.com – personalized ai search, 2024. Accessed: 2025-05-02.
  84. 84.Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen. Evaluating large language models at evaluating instruction following. In The Twelfth International Conference on Learning Representations, 2024.
  85. 85.Yu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang, Shuiwang Ji, Wei Wang, and Jiawei Han. A comprehensive survey of scientific large language models and their applications in scientific discovery. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 8783–8817, Miami, Florida, USA, November 2024. Association for Computational Linguistics.
  86. 86.Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. Wildchat: 1m chatGPT interaction logs in the wild. In The Twelfth International Conference on Learning Representations, 2024.
  87. 87.Yilun Zhao, Weiyuan Chen, Zhijian Xu, Manasi Patwardhan, Chengye Wang, Yixin Liu, Lovekesh Vig, and Arman Cohan. AbGen: Evaluating large language models in ablation study design and evaluation for scientific research. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12479–12491, Vienna, Austria, July 2025. Association for Computational Linguistics.
  88. 88.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Haotong Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena. ArXiv, abs/2306.05685, 2023.
  89. 89.Yuxiang Zheng, Shichao Sun, Lin Qiu, Dongyu Ru, Cheng Jiayang, Xuefeng Li, Jifan Lin, Binjie Wang, Yun Luo, Renjie Pan, Yang Xu, Qingkai Min, Zizhao Zhang, Yiwen Wang, Wenjie Li, and Pengfei Liu. OpenResearcher: Unleashing AI for accelerated scientific research. In Delia Irazu Hernandez Farias, Tom Hope, and Manling Li, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 209–218, Miami, Florida, USA, November 2024. Association for Computational Linguistics.
  90. 90.Zhipu ai. Glm-4.5: Reasoning, coding, and agentic abililties. https://z.ai/blog/glm-4.5//, July 2025.

Citation

MLA
Zhao, Y., et al. “SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks”. Advances in Neural Information Processing Systems 38, 2025, pp. 117124–69, https://doi.org/10.52202/085713-3532.
APA
Zhao, Y., Zhang, K., Hu, T., Wu, S., Le Bras, R., Liu, Y., Tang, R., Chang, J. C., Dodge, J., Bragg, J., Zhao, C., Hajishirzi, H., Downey, D., & Cohan, A. (2025). SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks. Advances in Neural Information Processing Systems 38, 117124–117169. https://doi.org/10.52202/085713-3532
Chicago
Zhao, Y., K. Zhang, T. Hu, et al. 2025. “SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks”. Advances in Neural Information Processing Systems 38, 117124–69. https://doi.org/10.52202/085713-3532.
Harvard
Zhao, Y. et al. (2025) “SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks”, Advances in Neural Information Processing Systems 38. Neural Information Processing Systems Foundation, Inc. (NeurIPS), pp. 117124–117169. Available at: https://doi.org/10.52202/085713-3532.
Vancouver
1. Zhao Y, Zhang K, Hu T, et al (2025) SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks. In: Advances in Neural Information Processing Systems 38. Neural Information Processing Systems Foundation, Inc. (NeurIPS), pp 117124–117169

BibTeX

@inproceedings{Zhao_2025, series={NeurIPS 2025}, title={SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks}, url={http://dx.doi.org/10.52202/085713-3532}, DOI={10.52202/085713-3532}, booktitle={Advances in Neural Information Processing Systems 38}, publisher={Neural Information Processing Systems Foundation, Inc. (NeurIPS)}, author={Zhao, Yilun and Zhang, Kaiyan and Hu, Tiansheng and Wu, Sihong and Le Bras, Ronan and Liu, Yixin and Tang, Robert and Chang, Joseph Chee and Dodge, Jesse and Bragg, Jonathan and Zhao, Chen and Hajishirzi, Hanna and Downey, Doug and Cohan, Arman}, year={2025}, pages={117124–117169}, collection={NeurIPS 2025} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/