Can We Automate Scientific Reviewing?
Weizhe YuanPengfei LiuGraham Neubig
Evaluates the feasibility of automated peer review by benchmarking state-of-the-art language models on review generation, revealing critical limitations in factuality and bias while outlining practical directions for machine-assisted reviewing.
The exponential growth of scientific publications has placed an unsustainable burden on the peer review system, which relies heavily on limited expert labor. This escalating volume causes delays, inconsistencies, and reviewer fatigue, prompting interest in whether artificial intelligence can assist in or automate the reviewing process.
The article evaluates the feasibility of automating scientific peer review using modern natural language processing models. Specifically, it establishes quantitative criteria for review quality, develops an automated generation system called ReviewAdvisor, and benchmarks its capabilities and limitations against human-written reviews.
To conduct this evaluation, the authors compiled a dataset of 8,877 machine learning papers and 28,119 peer reviews from major conferences. They defined an eight-aspect typology covering areas such as core summaries, clarity, soundness, and originality, annotating 1,000 reviews manually to train a reliable tagging model for the remainder of the corpus. The authors then built a two-stage summarization pipeline using a pre-trained language model, which first extracted salient content from full papers and subsequently generated structured, aspect-aware reviews. System performance was assessed across multiple dimensions, including decisiveness, aspect coverage, informativeness, factual accuracy, and linguistic bias, combining automated metrics with human evaluations.
The analysis revealed several critical findings regarding current automated reviewing capabilities. First, the automated system demonstrated strong core summarization skills, achieving summary accuracy rates between 70% and 93%, which is statistically comparable to human reviewers at around 91%. Second, generated reviews were more comprehensive in breadth, outperforming human reviewers by up to 14% in covering diverse predefined review aspects. However, the system failed significantly in critical evaluation and high-level reasoning: its constructiveness score on negative feedback was only 32% to 44%, compared to human constructiveness of approximately 76%, because the evidence generated to justify negative critiques was frequently non-factual or fabricated. Furthermore, recommendation decisiveness was poor, yielding negative alignment scores between -12% and -38% compared to positive human alignment of roughly 30%. Finally, the system exhibited notable disparities, behaving more harshly toward papers authored by non-native English speakers regarding originality while narrowing gaps in clarity ratings.
These findings demonstrate that automated peer review systems cannot replace human subject matter experts in high-stakes evaluation settings. Deploying generative models autonomously presents severe risks of producing misleading, non-factual justifications and reinforcing linguistic or demographic biases. However, the models show promise as assistive drafting aids. They can quickly generate accurate paper summaries, supply structural templates, and help reviewers and authors identify whether key aspects—such as replicability or baseline comparisons—have been addressed.
Moving forward, organizations and publication venues should restrict automated reviewing technologies to machine-assisted human workflows rather than autonomous decision-making. Future research must prioritize resolving factual hallucinations in critical text generation, integrating external domain knowledge and citation networks, enhancing long-document modeling beyond simple extractive heuristics, and establishing bias-mitigation protocols.
These conclusions are bounded by certain limitations, as the empirical study focused exclusively on machine learning conference papers, utilized heuristic sentence selection to fit model input constraints, and relied on a relatively small sample of 28 papers for in-depth author evaluations of constructiveness. Consequently, while confidence is high that language models can reliably extract and summarize paper contributions, extreme caution is necessary regarding their critical evaluative claims.
- Paper: On Faithfulness and Factuality in Abstractive Summarization, Joshua Maynez et al. (2020). Establishes the foundational empirical taxonomy of hallucinations and factual unfaithfulness in abstractive neural summarization, which directly underpins the analysis of factual fabrications in automated review generation.
- Paper: BLEURT: Learning Robust Metrics for Text Generation, Thibault Sellam et al. (2020). Introduces learned metrics for natural language generation evaluation to move beyond lexical overlap, providing necessary background on evaluating semantic fidelity in model outputs.
- Paper: Automatic Evaluation of Summaries Using N-gram Co-occurrence Statistics, Chin-Yew Lin et al. (2003). Provides the foundational n-gram co-occurrence framework for automatic summarization evaluation that informs standard baseline assessment of generated scientific text.
- Paper: Pre-trained models for natural language processing: A survey, Xipeng Qiu et al. (2020). Surveys the pre-trained language model architectures and sequence-to-sequence adaptation paradigms utilized in two-stage scientific review generation pipelines.
- Paper: Towards Automating Scientific Review with Google's Paper Assistant Tool, Rajesh Jayaram et al. (2026). Advances automated peer review from simple summarization to agentic, verification-focused systems capable of detecting theoretical and mathematical errors in conference submissions.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). Generalizes the evaluation of large language models as automated judges and systematically analyzes the bias and alignment failure modes identified in automated reviewing.
- Paper: FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation, Sewon Min et al. (2023). Develops atomic fact verification methods to directly address and measure the factual inaccuracies and non-factual justifications observed in generated long-form reviews.
- Paper: Enabling Large Language Models to Generate Text with Citations, Tianyu Gao et al. (2023). Addresses the need for grounded evidence in generative text by benchmarking citation support and fact retrieval in long-form language model outputs.
- Paper: Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text, Sebastian Gehrmann et al. (2023). Critically expands on the evaluation pitfalls and human assessment discrepancies encountered when benchmarking complex natural language generation tasks.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). Explores formal frameworks, structured rubrics, and mitigation protocols for utilizing language models as evaluators across multidimensional quality criteria.
- Paper: Idea2Story: An Automated Pipeline for Transforming Research Concepts into Complete Scientific Narratives, Tengyue Xu et al. (2026). Applies conference peer review data and simulated review feedback loops to guide automated scientific paper narrative construction.
- Paper: PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing, Yiwen Song et al. (2026). Incorporates simulated peer feedback into a multi-agent framework designed to draft and iteratively refine scientific research papers.
