Can We Automate Scientific Reviewing?

Weizhe YuanPengfei LiuGraham Neubig

article2022JAIR118 citations

Evaluates the feasibility of automated peer review by benchmarking state-of-the-art language models on review generation, revealing critical limitations in factuality and bias while outlining practical directions for machine-assisted reviewing.

Listen

The exponential growth of scientific publications has placed an unsustainable burden on the peer review system, which relies heavily on limited expert labor. This escalating volume causes delays, inconsistencies, and reviewer fatigue, prompting interest in whether artificial intelligence can assist in or automate the reviewing process.

The article evaluates the feasibility of automating scientific peer review using modern natural language processing models. Specifically, it establishes quantitative criteria for review quality, develops an automated generation system called ReviewAdvisor, and benchmarks its capabilities and limitations against human-written reviews.

To conduct this evaluation, the authors compiled a dataset of 8,877 machine learning papers and 28,119 peer reviews from major conferences. They defined an eight-aspect typology covering areas such as core summaries, clarity, soundness, and originality, annotating 1,000 reviews manually to train a reliable tagging model for the remainder of the corpus. The authors then built a two-stage summarization pipeline using a pre-trained language model, which first extracted salient content from full papers and subsequently generated structured, aspect-aware reviews. System performance was assessed across multiple dimensions, including decisiveness, aspect coverage, informativeness, factual accuracy, and linguistic bias, combining automated metrics with human evaluations.

The analysis revealed several critical findings regarding current automated reviewing capabilities. First, the automated system demonstrated strong core summarization skills, achieving summary accuracy rates between 70% and 93%, which is statistically comparable to human reviewers at around 91%. Second, generated reviews were more comprehensive in breadth, outperforming human reviewers by up to 14% in covering diverse predefined review aspects. However, the system failed significantly in critical evaluation and high-level reasoning: its constructiveness score on negative feedback was only 32% to 44%, compared to human constructiveness of approximately 76%, because the evidence generated to justify negative critiques was frequently non-factual or fabricated. Furthermore, recommendation decisiveness was poor, yielding negative alignment scores between -12% and -38% compared to positive human alignment of roughly 30%. Finally, the system exhibited notable disparities, behaving more harshly toward papers authored by non-native English speakers regarding originality while narrowing gaps in clarity ratings.

These findings demonstrate that automated peer review systems cannot replace human subject matter experts in high-stakes evaluation settings. Deploying generative models autonomously presents severe risks of producing misleading, non-factual justifications and reinforcing linguistic or demographic biases. However, the models show promise as assistive drafting aids. They can quickly generate accurate paper summaries, supply structural templates, and help reviewers and authors identify whether key aspects—such as replicability or baseline comparisons—have been addressed.

Moving forward, organizations and publication venues should restrict automated reviewing technologies to machine-assisted human workflows rather than autonomous decision-making. Future research must prioritize resolving factual hallucinations in critical text generation, integrating external domain knowledge and citation networks, enhancing long-document modeling beyond simple extractive heuristics, and establishing bias-mitigation protocols.

These conclusions are bounded by certain limitations, as the empirical study focused exclusively on machine learning conference papers, utilized heuristic sentence selection to fit model input constraints, and relied on a relatively small sample of 28 papers for in-depth author evaluations of constructiveness. Consequently, while confidence is high that language models can reliably extract and summarize paper contributions, extreme caution is necessary regarding their critical evaluative claims.

Cover for Can We Automate Scientific Reviewing?

Abstract

The rapid development of science and technology has been accompanied by an exponential growth in peer-reviewed scientific publications. At the same time, the review of each paper is a laborious process that must be carried out by subject matter experts. Thus, providing high-quality reviews of this growing number of papers is a significant challenge. In this work, we ask the question “can we automate scientific reviewing?”, discussing the possibility of using natural language processing (NLP) models to generate peer reviews for scientific papers. Because it is non-trivial to define what a “good” review is in the first place, we first discuss possible evaluation metrics that could be used to judge success in this task. We then focus on the machine learning domain and collect a dataset of papers in the domain, annotate them with different aspects of content covered in each review, and train targeted summarization models that take in papers as input and generate reviews as output. Comprehensive experimental results on the test set show that while system-generated reviews are comprehensive, touching upon more aspects of the paper than human-written reviews, the generated texts are less constructive and less factual than human-written reviews for all aspects except the explanation of the core ideas of the papers, which are largely factually correct. Given these results, we pose eight challenges in the pursuit of a good review generation system together with potential solutions, which, hopefully, will inspire more future research in this direction.

We make relevant resource publicly available for use by future research: https://github.com/neulab/ReviewAdvisor. In addition, while our conclusion is that the technology is not yet ready for use in high-stakes review settings we provide a system demo, ReviewAdvisor (http://review.nlpedia.ai/), showing the current capabilities and failings of state-of-the-art NLP models at this task (see demo screenshot in A.2). A review of this paper written by the system proposed in this paper can be found in A.1.

Table of Contents

  • 1. Introduction
  • 2. What Makes a Good Peer Review?
  • 2.1 Peer Review for Scientific Research
  • 2.2 Multi-Perspective Evaluation
  • 2.2.1 D1: Decisiveness
  • 2.2.2 D2: Comprehensiveness
  • 2.2.3 D3: Justification
  • 2.2.4 D4: Accuracy
  • 2.2.5 D5: Kindness
  • 2.2.6 Similarity to Human Reviews
  • 3. Dataset
  • 3.1 Data Collection
  • 3.2 Aspect-enhanced Review Dataset
  • 4. Scientific Review Generation
  • 4.1 Task Formulation
  • 4.2 System Design
  • 4.2.2 Aspect-aware Summarization
  • 5.1 Settings
  • 5.2 What are Systems Good and Bad At?
  • 5.2.1 Weaknesses
  • 5.2.2 Advantages
  • 5.2.3 System Comparisons
  • 5.2.4 Case Study
  • 5.3 Will System Generate Biased Reviews?
  • 5.3.1 Measuring Disparity in Reviews
  • 5.3.2 Nativeness Analysis
  • 5.3.3 Anonymity Analysis
  • 6. Related Work
  • 7. Discussion and Future Directions
  • 7.1 Machine-assisted Reviewing and Authoring
  • 7.2 Challenges and Promising Directions
  • 7.2.1 Model
  • 7.2.2 Datasets
  • 7.2.3 Evaluation
  • 7.3 Conclusion
  • Acknowledgments
  • Appendix A. Details of Experiments
  • A.1 Review of this Paper Written by Our Model
  • A.2 Screenshot of Our Demo System
  • A.3 Details for Evaluation Metrics
  • A.4 Training of Aspect Tagger
  • A.5 Heuristics for Refining Prediction Results
  • A.6 An Example of Automatically Annotated Reviews
  • A.7 Calculation of Aspect Precision and Aspect Recall
  • A.8 Adjusting BART for Long Documents
  • A.9 CE Extraction Details
  • A.9.1 Keywords Filtering
  • A.9.2 Cross Entropy Method
  • A.10 Detailed Analysis and Case Study
  • A.11 Calculation of Aspect Score
  • A.12 Disparity Analysis for All Models
  • Appendix B. Supplemental Material
  • B.1 Dataset Annotation Guideline
  • Motivation
  • Originality
  • Soundness
  • Substance
  • Replicability
  • Meaningful Comparison
  • Clarity
  • References

Knowls

  1. Knowl 1 — Multi-perspective framework for evaluating scientific reviews

    definition

    The paper evaluates a review RR of a paper DD from several complementary perspectives rather than using only text similarity. A good review is treated as decisive, comprehensive, justified, accurate, and kind. Let RmR^m be the paper’s meta-review, let Dec⁡(D)∈{−1,1}\operatorname{Dec}(D)\in\{-1,1\} denote the final accept/reject decision for DD, and let Rec⁡(R)∈{−1,0,1}\operatorname{Rec}(R)\in\{-1,0,1\} denote the review’s reject/neutral/accept recommendation.

    Recommendation Accuracy is

    RAcc⁡(R)=Dec⁡(D)Rec⁡(R),\operatorname{RAcc}(R)=\operatorname{Dec}(D)\operatorname{Rec}(R),

    so a correct decisive recommendation scores 11, an incorrect decisive recommendation scores −1-1, and a neutral recommendation scores 00. Aspect Coverage (ACov) is the fraction of the predefined review aspects mentioned in RR. Aspect Recall (ARec) is the fraction of aspect-and-polarity judgments present in RmR^m that are also covered by RR.

    For justification, Informativeness (Info) is the fraction of negative aspect judgments in RR that have accompanying evidence. If nna(R)n_{na}(R) is the number of negative aspects in RR and nnae(R)n_{nae}(R) is the number of those negative aspects accompanied by evidence, then

    Info⁡(R)=nnae(R)nna(R),\operatorname{Info}(R)=\frac{n_{nae}(R)}{n_{na}(R)},

    with Info set to 11 when RR contains no negative aspects. Aspect-level Constructiveness (ACon) is the fraction of evidence statements that human annotators judge to be valid support; it is set to 11 when no evidence is provided. Summary Accuracy (SAcc) scores the review’s summary as 00, 0.50.5, or 11 for incorrect/absent, partially correct, or correct.

    ROUGE and BERTScore measure similarity to human reference reviews; when several reference reviews exist for one paper, the maximum similarity is used rather than the average. ACov, ARec, ROUGE, and BERTScore can be computed automatically, whereas RAcc, Info, ACon, and SAcc require human judgments. The paper leaves computational measurement of kindness for future work.

  2. Knowl 2 — ASAP-Review dataset of machine-learning peer reviews

    data/table

    The authors construct ASAP-Review, a machine-learning peer-review dataset collected from ICLR papers from 2017–2020 through OpenReview and NeurIPS papers from 2016–2019 through the NeurIPS Proceedings. Each paper is paired with publicly available reference reviews, a meta-review when available, the final accept/reject decision, bibliographic metadata, and structured paper text extracted with AllenAI Science-parse. NeurIPS exposes reviews only for accepted papers, so the dataset contains no publicly available rejected NeurIPS reviews.

    The dataset statistics reported on page 177 are:

    Could not parse LaTeX table

    The resulting resource supplies both training data for review generation and review-level metadata needed for aspect-based evaluation and disparity analysis.

  3. Knowl 3 — Aspect typology and scalable review annotation

    model/method

    ASAP-Review represents review content with eight aspects: Summary (SUM), Motivation/Impact (MOT), Originality (ORI), Soundness/Correctness (SOU), Substance (SUB), Replicability (REP), Meaningful Comparison (CMP), and Clarity (CLA). Every aspect except Summary also receives positive or negative polarity. Thus the annotation scheme contains 15 categories: one unpolarized summary category and seven aspects with two polarities each.

    Six machine-learning or natural-language-processing students manually annotated 1,000 reviews by marking text spans associated with aspects. Each review was labeled by two annotators, and the lowest pairwise Cohen’s kappa was 0.6530.653. To annotate the remaining reviews, the authors trained a BERT-large-cased token classifier on 900 annotated reviews for training and 100 for validation, treating each sentence as an independent sequence-labeling example. The classifier was optimized with Adam at learning rate 5×10−55\times10^{-5} for five epochs, retaining the checkpoint with the lowest validation loss.

    The automatically predicted spans were refined with seven sequential heuristics: merge intervening tokens inside a discontinuous Summary span; retain only the first Summary span; remove punctuation tags that disagree with neighboring tags; fill a one-token gap between identical tags; remove isolated one-token non-outside tags; extend non-Summary spans bounded by outside tags; and truncate or extend Summary by at most five words so that it ends at a period.

    Human evaluation of 300 automatically annotated samples, with three annotators per sample, yielded 92.75% aspect precision and 85.19% aspect recall. The fine-grained precision/recall values reported on page 180 were:

    Could not parse LaTeX table
  4. Knowl 4 — ReviewAdvisor extract-then-generate-and-predict architecture

    model/method

    ReviewAdvisor treats scientific review generation as reviewer-view summarization: unlike author-view summarization, the output should summarize the paper while also expressing aspect-specific critical judgments. Because BART accepts only limited input lengths, the system first extracts salient paper sentences and then generates a review from the extraction.

    The generator is BART initialized from the bart-large-cnn checkpoint. Two generation variants are compared: a vanilla sequence-to-sequence decoder and an aspect-aware multitask decoder. The aspect-aware variant adds a multilayer perceptron over decoder representations and jointly predicts the next review token and its aspect label. Its training objective is

    L=Lseq2seq+αLseqlab,\mathcal{L}=\mathcal{L}_{\mathrm{seq2seq}}+\alpha\mathcal{L}_{\mathrm{seqlab}},

    where Lseq2seq\mathcal{L}_{\mathrm{seq2seq}} is the negative log-likelihood of the reference review tokens, Lseqlab\mathcal{L}_{\mathrm{seqlab}} is the negative log-likelihood of the corresponding aspect labels, and α=0.1\alpha=0.1 is tuned on development-set Aspect Coverage. The aspect labels are auxiliary output supervision rather than additional user-provided input.

    The extraction variants are: the Introduction alone; cross-entropy (CE) extraction from the full paper; and a hybrid consisting of the abstract plus CE extraction. An oracle extractor, constructed separately for analysis, greedily selects sentences maximizing average ROUGE-1, ROUGE-2, and ROUGE-L against a reference review. The aspect-aware system can therefore produce fluent text while associating generated spans with aspects and their polarities; the case study on page 187 shows summary, positive clarity/substance, and negative soundness/substance spans aligned with generated text.

  5. Knowl 5 — Cross-entropy sentence extraction for long scientific papers

    algorithm

    The CE extractor selects a short, diverse set of paper sentences before BART generation. It first retains sentences containing one or more of 48 predefined informative keywords or their inflections, including terms such as “propose,” “introduce,” “compare,” “evaluate,” “outperform,” and “experiment.” It then searches for a subset containing fewer than 30 sentences that maximizes unigram entropy.

    For a candidate concatenation SS of selected sentences, let Len⁡(S)\operatorname{Len}(S) be its number of words, let Count⁡(w)\operatorname{Count}(w) be the count of word ww after lowercasing, punctuation removal, and stop-word removal, and define

    pS(w)=Count⁡(w)Len⁡(S),R(S)=−∑w∈SpS(w)log⁡pS(w).p_S(w)=\frac{\operatorname{Count}(w)}{\operatorname{Len}(S)},\qquad R(S)=-\sum_{w\in S}p_S(w)\log p_S(w).

    For a paper with nn sentences, the cross-entropy search maintains a Bernoulli selection-probability vector pt∈[0,1]np_t\in[0,1]^n. It starts from p0=(1/2,…,1/2)p_0=(1/2,\ldots,1/2), samples N=1000N=1000 binary selection vectors, discards samples selecting more than 30 sentences, scores each sample with R(S)R(S), and retains the top ρ=0.05\rho=0.05 fraction. For sentence jj, the elite-sample estimate is

    p^t,j=∑i=1N1{R(Si)≥γt}1{Xi,j=1}∑i=1N1{R(Si)≥γt},\hat p_{t,j}=\frac{\sum_{i=1}^{N}\mathbf{1}\{R(S_i)\ge\gamma_t\}\mathbf{1}\{X_{i,j}=1\}}{\sum_{i=1}^{N}\mathbf{1}\{R(S_i)\ge\gamma_t\}},

    where XiX_i is sample ii, SiS_i is its selected-sentence concatenation, γt\gamma_t is the elite-score threshold, and 1{⋅}\mathbf{1}\{\cdot\} is the indicator function. The probabilities are smoothed with

    pt=αp^t+(1−α)pt−1,p_t=\alpha\hat p_t+(1-\alpha)p_{t-1},

    using α=0.7\alpha=0.7. Iteration stops when γt\gamma_t remains unchanged for three iterations; the converged probabilities determine the extraction. When more than 90 sentences remain after keyword filtering, the initial probabilities are slightly reduced so that enough valid samples are obtained. The extractor tends to select sentences near both the beginning and end of papers, complementing the introduction-focused baseline with experimental and conclusion content.

  6. Knowl 6 — Experimental configuration for ReviewAdvisor

    experimental setup

    Experiments use 8,742 unique papers and 25,986 paper–review pairs after filtering papers with fewer than 100 full-text words and reviews outside the 100–1,024-word range; the retained reviews represent 92.57% of all reviews. For extraction, papers are truncated to their first 250 sentences when longer than 250 sentences. The split contains 6,993 training papers and 20,757 pairs, 874 validation papers and 2,571 pairs, and 875 test papers and 2,658 pairs.

    All generators initialize from bart-large-cnn. Adam training uses a learning rate warmed linearly from 00 to 4×10−54\times10^{-5} during the first 10% of updates and then linearly decayed to zero over the remaining updates. Models are fine-tuned for five epochs, with the checkpoint having the lowest validation loss selected. Generation uses beam search with beam size 4, minimum and maximum lengths of 100 and 1,024 words according to the BART subword tokenizer, length penalty 2.0, and trigram blocking.

    Automatic evaluation measures ACov, ARec, ROUGE, and BERTScore. Human evaluation measures RAcc, Info, ACon, and SAcc on 28 held-out papers from machine learning, natural language processing, computer vision, and reinforcement learning; none of these papers appears in training. Human-review metrics are averaged across reference reviews for each paper, while ROUGE and BERTScore use the maximum over references.

  7. Knowl 7 — Generated reviews are broad and often factually unreliable

    data/table

    The main test-set comparison, reported on page 185, contrasts human reference reviews with extractive baselines and six extractive-plus-abstractive ReviewAdvisor systems. INTRO, CE, and ABSCE denote Introduction extraction, cross-entropy extraction, and abstract-plus-CE extraction, respectively; ×\times and ✓\checkmark indicate absence or presence of auxiliary aspect prediction. RAcc, ACov, ARec, Info, ACon, and SAcc are reported in percentage points; R-1, R-2, R-L, and BS are ROUGE-1, ROUGE-2, ROUGE-L, and BERTScore.

    Could not parse LaTeX table

    The systems cover more aspects than human reviews: maximum ACov is 63.96 versus 49.85, and maximum Info is 100.00 versus 97.97. However, their recommendations and negative evidence are unreliable: the best generated RAcc is −11.54-11.54 versus human RAcc 30.32, a gap of 41.86 points, and the best generated ACon is 43.78 versus human ACon 75.67, a gap of 31.89 points. Summary Accuracy is comparatively strong: four of six systems exceed 80%, and the top three are not statistically distinguishable from human reviews. Extractive-plus-abstractive systems consistently outperform pure extractive systems on ROUGE and BERTScore, while full-paper CE extraction generally improves aspect coverage over using only the Introduction.

  8. Knowl 8 — Characteristic strengths and failure modes of generated reviews

    empirical result

    Human inspection and metric analysis show that ReviewAdvisor is strongest at summarizing a paper’s central contribution and weakest at high-level assessment. The generated reviews often correctly describe the core idea, can cover more review aspects than human references, and can supply evidence-like sentences from the paper. The aspect-aware model also produces fluent spans with recognizable aspect and polarity associations; its attention often focuses on contribution indicators such as “we propose” for summaries, prior work and contribution descriptions for originality, experimental settings and reported numbers for substance, and citation-like “et al.” contexts for meaningful comparison.

    The main failures are non-factual assessment and weak discrimination between strong and weak papers. Generated negative comments frequently lack valid support even when they contain an ostensible reason. The systems also imitate frequent training-set phrasing: the sentence “The paper is well-written and easy to follow” occurs in more than 90% of generated reviews, although the exact sentence appears in more than 10% of training papers. Generated reviews ask substantially fewer questions than human reviews, averaging 0.32 questions per review versus 2.04 in reference reviews. These findings motivate using generated reviews as drafts, evidence-organizing aids, or review templates rather than as autonomous high-stakes evaluations.

  9. Knowl 9 — Fine-grained disparity analysis reveals group-dependent review behavior

    empirical result

    The paper introduces a disparity analysis based on aspect polarity. For a review group GiG_i, the aspect score is the percentage of positive occurrences of an aspect; if the aspect is absent, its score is set to 0.50.5 to represent neutrality. A signed, normalized disparity δ(R,G)\delta(R,G) compares the scores of two groups G=[G0,G1]G=[G_0,G_1], and the difference between generated-review and reference-review disparities is

    Δ(Rg,Rr,G)=δ(Rg,G)−δ(Rr,G),\Delta(R^g,R^r,G)=\delta(R^g,G)-\delta(R^r,G),

    where RgR^g denotes system-generated reviews and RrR^r denotes human reference reviews. Positive values indicate that generated reviews favor G0G_0 more than human reviews do.

    The test set contains 651 native-author papers and 224 non-native-author papers, with accept rates of 66.51% and 50.00%, respectively; it also contains 613 anonymous and 217 non-anonymous papers, with accept rates of 57.59% and 78.34%. Nativeness is assigned using author nationality and, when needed, institution location. Anonymity is assigned according to whether a paper was released as a preprint before half a month after the submission deadline.

    The disparity differences reported on page 190 are:

    Could not parse LaTeX table

    Human reviewers give native-author papers higher scores in most aspects, with the clearest native advantage in Clarity. The automatic system narrows the human Clarity disparity but increases the native advantage in Originality, with a nativeness disparity difference of +18.71+18.71 for Originality and −13.32-13.32 for Clarity. Both human and generated reviews favor non-anonymous papers across all aspects. The total absolute disparity difference is smaller for anonymity (28.00) than for nativeness (43.39), although the authors caution that observed disparities do not by themselves establish bias because group-level paper quality is not fully controlled.

  10. Knowl 10 — Limits of fully automated scientific reviewing

    limitation

    The authors conclude that automatic review generation is not ready to replace expert reviewers in high-stakes settings. The system lacks reliable factuality, external knowledge of the surrounding literature, and robust understanding of a paper’s scientific merit; it also inherits or creates disparities across author groups. Its current usefulness is therefore framed as machine-assisted reviewing or authoring: presenting multiple generated drafts, linking aspect comments to source evidence, helping readers grasp a paper’s core idea, and suggesting possible weaknesses that humans must verify.

    The paper identifies eight directions needed for a stronger system: long-document modeling beyond the two-stage extraction workaround; sequence-to-sequence pretraining on scientific text; explicit use of scientific document structure; retrieval, citation graphs, or knowledge graphs for external knowledge; more open and fine-grained review data; more accurate parsing of tables, figures, and other paper structure; improved fairness and bias evaluation; and factuality, confidence, and calibration assessment. These limitations constrain the interpretation of the reported results to machine-learning peer reviews generated from the paper text itself.

Coverage note — The appendix’s exhaustive heuristic examples, per-model disparity plots, demo screenshot, and auxiliary implementation trials were omitted because they elaborate the included methods and findings rather than constitute separate load-bearing contributions.

References

  1. 1.Angelidis, S., & Lapata, M. (2018). Summarizing opinions: Aspect extraction meets sentiment prediction and they are both weakly supervised. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 3675–3686, Brussels, Belgium. Association for Computational Linguistics.
  2. 2.Anjum, O., Gong, H., Bhat, S., Hwu, W.-M., & Xiong, J. (2019). PaRe: A paper-reviewer matching approach using a common topic space. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 518–528, Hong Kong, China. Association for Computational Linguistics.
  3. 3.August, T., Kim, L., Reinecke, K., & Smith, N. A. (2020). Writing strategies for science communication: Data and computational analysis. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 5327–5344, Online. Association for Computational Linguistics.
  4. 4.Bartoli, A., De Lorenzo, A., Medvet, E., & Tarlao, F. (2016). Your paper has been accepted, rejected, or whatever: Automatic generation of scientific paper reviews. In International Conference on Availability, Reliability, and Security, pp. 19–28. Springer.
  5. 5.Beltagy, I., Lo, K., & Cohan, A. (2019). Scibert: A pretrained language model for scientific text. arXiv preprint arXiv:1903.10676.
  6. 6.Bolukbasi, T., Chang, K.-W., Zou, J., Saligrama, V., & Kalai, A. (2016). Man is to computer programmer as woman is to homemaker? debiasing word embeddings..
  7. 7.Bornmann, L., & Mutz, R. (2015). Growth rates of modern science: A bibliometric analysis based on the number of publications and cited references. Journal of the Association for Information Science and Technology, 66 (11), 2215–2222.
  8. 8.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020). Language models are few-shot learners. arXiv preprint arXiv:2005.14165.
  9. 9.Cachola, I., Lo, K., Cohan, A., & Weld, D. (2020a). TLDR: Extreme summarization of scientific documents. In Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 4766–4777, Online. Association for Computational Linguistics.
  10. 10.Cachola, I., Lo, K., Cohan, A., & Weld, D. S. (2020b). Tldr: Extreme summarization of scientific documents. ArXiv, abs/2004.15011.
  11. 11.Chakraborty, S., Goyal, P., & Mukherjee, A. (2020). Aspect-based sentiment analysis of scientific reviews. arXiv preprint arXiv:2006.03257.
  12. 12.Chen, Y.-C., & Bansal, M. (2018). Fast abstractive summarization with reinforce-selected sentence rewriting. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vol. 1, pp. 675–686.
  13. 13.Cohan, A., Dernoncourt, F., Kim, D. S., Bui, T., Kim, S., Chang, W., & Goharian, N. (2018a). A discourse-aware attention model for abstractive summarization of long documents. In NAACL-HLT.
  14. 14.Cohan, A., Dernoncourt, F., Kim, D. S., Bui, T., Kim, S., Chang, W., & Goharian, N. (2018b). A discourse-aware attention model for abstractive summarization of long documents. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pp. 615–621, New Orleans, Louisiana. Association for Computational Linguistics.
  15. 15.Cohan, A., & Goharian, N. (2017). Scientific article summarization using citation-context and article’s discourse structure. arXiv preprint arXiv:1704.06619.
  16. 16.De Bellis, N. (2009). Bibliometrics and citation analysis: from the science citation index to cybermetrics. scarecrow press.
  17. 17.Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186.
  18. 18.Dou, Z.-Y., Liu, P., Hayashi, H., Jiang, Z., & Neubig, G. (2020). Gsum: A general framework for guided neural abstractive summarization. arXiv preprint arXiv:2010.08014.
  19. 19.Efron, B. (1992). Bootstrap methods: another look at the jackknife. In Breakthroughs in statistics, pp. 569–593. Springer.
  20. 20.Erera, S., Shmueli-Scheuer, M., Feigenblat, G., Nakash, O., Boni, O., Roitman, H., Cohen, D., Weiner, B., Mass, Y., Rivlin, O., Lev, G., Jerbi, A., Herzig, J., Hou, Y., Jochim, C., Gleize, M., Bonin, F., & Konopnicki, D. (2019). A summarization system for scientific documents. In EMNLP/IJCNLP.
  21. 21.Feigenblat, G., Roitman, H., Boni, O., & Konopnicki, D. (2017). Unsupervised query-focused multi-document summarization using the cross entropy method. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’17, p. 961–964, New York, NY, USA. Association for Computing Machinery.
  22. 22.Frermann, L., & Klementiev, A. (2019). Inducing document structure for aspect-based summarization. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 6263–6273, Florence, Italy. Association for Computational Linguistics.
  23. 23.Gao, Y., Eger, S., Kuznetsov, I., Gurevych, I., & Miyao, Y. (2019). Does my rebuttal matter? insights from a major NLP conference. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 1274–1290, Minneapolis, Minnesota.
  24. 24.Gehrmann, S., Deng, Y., & Rush, A. (2018). Bottom-up abstractive summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 4098–4109.
  25. 25.Guu, K., Lee, K., Tung, Z., Pasupat, P., & Chang, M.-W. (2020). Realm: Retrieval-augmented language model pre-training. arXiv preprint arXiv:2002.08909.
  26. 26.Hayashi, H., Budania, P., Wang, P., Ackerson, C., Neervannan, R., & Neubig, G. (2020). Wikiasp: A dataset for multi-domain aspect-based summarization. Transactions of the Association for Computational Linguistics (TACL).
  27. 27.He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778.
  28. 28.Hermann, K. M., Kocisky, T., Grefenstette, E., Espeholt, L., Kay, W., Suleyman, M., & Blunsom, P. (2015). Teaching machines to read and comprehend. In Advances in Neural Information Processing Systems, pp. 1684–1692.
  29. 29.Hou, Y., Jochim, C., Gleize, M., Bonin, F., & Ganguly, D. (2019). Identification of tasks, datasets, evaluation metrics, and numeric scores for scientific leaderboards construction. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 5203–5213, Florence, Italy. Association for Computational Linguistics.
  30. 30.Huang, J.-B. (2018). Deep paper gestalt. arXiv preprint arXiv:1812.08775.
  31. 31.Jain, S., van Zuylen, M., Hajishirzi, H., & Beltagy, I. (2020). SciREX: A challenge dataset for document-level information extraction. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 7506–7516, Online. Association for Computational Linguistics.
  32. 32.Jecmen, S., Zhang, H., Liu, R., Shah, N. B., Conitzer, V., & Fang, F. (2020). Mitigating manipulation in peer review via randomized reviewer assignments. arXiv preprint arXiv:2006.16437.
  33. 33.Jefferson, T., Alderson, P., Wager, E., & Davidoff, F. (2002a). Effects of editorial peer review: a systematic review. Jama, 287 (21), 2784–2786.
  34. 34.Jefferson, T., Wager, E., & Davidoff, F. (2002b). Measuring the quality of editorial peer review. Jama, 287 (21), 2786–2790.
  35. 35.Jha, R., Abu-Jbara, A., & Radev, D. (2013). A system for summarizing scientific topics starting from keywords. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 572–577, Sofia, Bulgaria. Association for Computational Linguistics.
  36. 36.Jha, R., Coke, R., & Radev, D. R. (2015a). Surveyor: A system for generating coherent survey articles for scientific topics. In Bonet, B., & Koenig, S. (Eds.), Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, January 25-30, 2015, Austin, Texas, USA, pp. 2167–2173. AAAI Press.
  37. 37.Jha, R., Finegan-Dollak, C., King, B., Coke, R., & Radev, D. (2015b). Content models for survey generation: A factoid-based evaluation. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 441–450, Beijing, China. Association for Computational Linguistics.
  38. 38.Jin, J., Geng, Q., Zhao, Q., & Zhang, L. (2017). Integrating the trend of research interest for reviewer assignment. In Proceedings of the 26th International Conference on World Wide Web Companion, pp. 1233–1241.
  39. 39.Kang, D., Ammar, W., Dalvi, B., van Zuylen, M., Kohlmeier, S., Hovy, E., & Schwartz, R. (2018). A dataset of peer reviews (peerread): Collection, insights and nlp applications. In Meeting of the North American Chapter of the Association for Computational Linguistics (NAACL), New Orleans, USA.
  40. 40.Kingma, D., & Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  41. 41.Koncel-Kedziorski, R., Bekal, D., Luan, Y., Lapata, M., & Hajishirzi, H. (2019). Text Generation from Knowledge Graphs with Graph Transformers. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 2284–2293, Minneapolis, Minnesota. Association for Computational Linguistics.
  42. 42.Langford, J., & Guzdial, M. (2015). The arbitrariness of reviews, and advice for school administrators..
  43. 43.Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., & Zettlemoyer, L. (2019). Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. ArXiv, abs/1910.13461.
  44. 44.Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., K¨uttler, H., Lewis, M., Yih, W.-t., Rockt¨aschel, T., et al. (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks. arXiv preprint arXiv:2005.11401.
  45. 45.Lin, C.-Y., & Hovy, E. (2003). Automatic evaluation of summaries using n-gram co-occurrence statistics. In Proceedings of the 2003 Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics, pp. 150–157.
  46. 46.Lo, K., Wang, L. L., Neumann, M., Kinney, R., & Weld, D. (2020). S2ORC: The semantic scholar open research corpus. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 4969–4983, Online. Association for Computational Linguistics.
  47. 47.Luan, Y., He, L., Ostendorf, M., & Hajishirzi, H. (2018). Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction. arXiv preprint arXiv:1808.09602.
  48. 48.Luu, K., Koncel-Kedziorski, R., Lo, K., Cachola, I., & Smith, N. A. (2020). Citation text generation. ArXiv, abs/2002.00317.
  49. 49.Manzoor, E., & Shah, N. B. (2020). Uncovering latent biases in text: Method and application to peer review..
  50. 50.Mohammad, S., Dorr, B., Egan, M., Hassan, A., Muthukrishan, P., Qazvinian, V., Radev, D., & Zajic, D. (2009). Using citations to generate surveys of scientific paradigms. In Proceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics, pp. 584–592, Boulder, Colorado. Association for Computational Linguistics.
  51. 51.Nallapati, R., Zhai, F., & Zhou, B. (2017). Summarunner: A recurrent neural network based sequence model for extractive summarization of documents. ArXiv, abs/1611.04230.
  52. 52.Narayan, S., Cohen, S. B., & Lapata, M. (2018). Don’t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium.
  53. 53.Nguyen, J., S´anchez-Hernández, G., Agell, N., Rovira, X., & Angulo, C. (2018). A decision support tool using order weighted averaging for conference review assignment. Pattern Recognition Letters, 105, 114–120.
  54. 54.Paulus, R., Xiong, C., & Socher, R. (2017). A deep reinforced model for abstractive summarization. arXiv preprint arXiv:1705.04304.
  55. 55.Qiao, F., Xu, L., & Han, X. (2018). Modularized and attention-based recurrent convolutional neural network for automatic academic paper aspect scoring. In International Conference on Web Information Systems and Applications, pp. 68–76. Springer.
  56. 56.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI blog, 1 (8), 9.
  57. 57.Rae, J. W., Potapenko, A., Jayakumar, S. M., & Lillicrap, T. (2020). Compressive transformers for long-range sequence modelling. ArXiv, abs/1911.05507.
  58. 58.Rogers, A., & Augenstein, I. (2020). What can we do to improve peer review in NLP?. In Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 1256–1262, Online. Association for Computational Linguistics.
  59. 59.Rubinstein, R. Y., & Kroese, D. P. (2013). The cross-entropy method: a unified approach to combinatorial optimization, Monte-Carlo simulation and machine learning. Springer Science & Business Media.
  60. 60.Smith, R. (2006). Peer review: A flawed process at the heart of science and journals. Journal of the Royal Society of Medicine, 99, 178 – 182.
  61. 61.Stanovsky, G., Smith, N. A., & Zettlemoyer, L. (2019). Evaluating gender bias in machine translation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 1679–1684, Florence, Italy. Association for Computational Linguistics.
  62. 62.Stelmakh, I., Shah, N., & Singh, A. (2019). On testing for biases in peer review. In Advances in Neural Information Processing Systems, pp. 5286–5296.
  63. 63.Subramanian, S., Li, R., Pilault, J., & Pal, C. (2019). On extractive and abstractive neural document summarization with transformer language models. arXiv preprint arXiv:1909.03186.
  64. 64.Tabah, A. N. (1999). Literature dynamics: Studies on growth, diffusion, and epidemics. Annual review of information science and technology (ARIST), 34, 249–86.
  65. 65.Tomkins, A., Zhang, M., & Heavlin, W. D. (2017). Reviewer bias in single-versus double-blind peer review. Proceedings of the National Academy of Sciences, 114 (48), 12708–12713.
  66. 66.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. In Advances in neural information processing systems, pp. 5998–6008.
  67. 67.Von Bearnensquash, C. (2010). Paper gestalt. Secret Proceedings of Computer Vision and Pattern Recognition (CVPR).
  68. 68.Wadden, D., Lin, S., Lo, K., Wang, L. L., van Zuylen, M., Cohan, A., & Hajishirzi, H. (2020). Fact or fiction: Verifying scientific claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 7534–7550, Online. Association for Computational Linguistics.
  69. 69.Wang, K., & Wan, X. (2018). Sentiment analysis of peer review texts for scholarly papers. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, pp. 175–184.
  70. 70.Wang, Q., Zeng, Q., Huang, L., Knight, K., Ji, H., & Rajani, N. F. (2020a). Reviewrobot: Explainable paper review generation based on knowledge synthesis. In Proceedings of INLG.
  71. 71.Wang, Q., Zeng, Q., Huang, L., Knight, K., Ji, H., & Rajani, N. F. (2020b). ReviewRobot: Explainable paper review generation based on knowledge synthesis. In Proceedings of the 13th International Conference on Natural Language Generation, pp. 384–397, Dublin, Ireland. Association for Computational Linguistics.
  72. 72.Xiao, W., & Carenini, G. (2019). Extractive summarization of long documents by combining global and local context. ArXiv, abs/1909.08089.
  73. 73.Xing, X., Fan, X., & Wan, X. (2020). Automatic generation of citation texts in scholarly papers: A pilot study. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 6181–6190, Online. Association for Computational Linguistics.
  74. 74.Xiong, W., & Litman, D. (2011). Automatically predicting peer-review helpfulness. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pp. 502–507, Portland, Oregon, USA. Association for Computational Linguistics.
  75. 75.Yasunaga, M., Kasai, J., Zhang, R., Fabbri, A. R., Li, I., Friedman, D., & Radev, D. R. (2019a). Scisummnet: A large annotated corpus and content-impact models for scientific paper summarization with citation networks. In AAAI.
  76. 76.Yasunaga, M., Kasai, J., Zhang, R., Fabbri, A. R., Li, I., Friedman, D., & Radev, D. R. (2019b). Scisummnet: A large annotated corpus and content-impact models for scientific paper summarization with citation networks. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 - February 1, 2019, pp. 7386–7393. AAAI Press.
  77. 77.Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2019). Bertscore: Evaluating text generation with bert. arXiv, arXiv–1904.
  78. 78.Zhao, J., Wang, T., Yatskar, M., Ordonez, V., & Chang, K.-W. (2018). Gender bias in coreference resolution: Evaluation and debiasing methods. arXiv preprint arXiv:1804.06876.

Citation

MLA
Yuan, W., et al. “Can We Automate Scientific Reviewing?”. arXiv, 2021, http://arxiv.org/abs/2102.00176v1.
APA
Yuan, W., Liu, P., & Neubig, G. (2021). Can We Automate Scientific Reviewing?. arXiv. http://arxiv.org/abs/2102.00176v1
Chicago
Yuan, W., P. Liu, and G. Neubig. 2021. “Can We Automate Scientific Reviewing?”. arXiv. http://arxiv.org/abs/2102.00176v1.
Harvard
Yuan, W., Liu, P. and Neubig, G. (2021) “Can We Automate Scientific Reviewing?”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2102.00176v1.
Vancouver
1. Yuan W, Liu P, Neubig G (2021) Can We Automate Scientific Reviewing?. arXiv

BibTeX

@article{yuan2021can,
  title = {Can We Automate Scientific Reviewing?},
  author = {Yuan, Weizhe and Liu, Pengfei and Neubig, Graham},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2102.00176v1},
  eprint = {2102.00176}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/