Generative Language Models for Paragraph-Level Question Generation
Asahi UshioFernando Alva-ManchegoJosé Camacho-Collados
Introduces QG-Bench, a unified multilingual and multi-domain benchmark that standardizes paragraph-level question generation evaluation across eight languages and multiple domains using sequence-to-sequence language models.
Automatic question generation—creating relevant, natural questions from text given a specific target answer—is critical for training question answering systems, augmenting training data, and building automated educational tools. However, research in this area has suffered from fragmented datasets, non-standardized evaluation protocols, and arbitrary model selections, making it difficult to systematically measure progress and choose optimal architectures.
The article addresses this gap by establishing QG-Bench, a standardized benchmark designed to systematically train, evaluate, and compare sequence-to-sequence language models for paragraph-level question generation across diverse domains and languages.
To construct this benchmark, the authors unified standard datasets into a consistent paragraph-level format spanning English benchmark data, ten specialized domains across two question styles (objective and subjective), and seven non-English languages. Using this data, the study fine-tuned several popular sequence-to-sequence language models (T5, BART, mT5, and mBART) via an optimized two-phase hyperparameter search. The analysis combined standard automated metrics (BLEU4, ROUGE-L, METEOR, BERTScore, and MoverScore) with a large-scale manual human evaluation of 15,000 judgments and an assessment of domain and cross-lingual transferability.
The findings show that properly tuned standard architectures outperform previously reported state-of-the-art specialized models, with T5-Large achieving the strongest performance across metrics (BLEU4 of 27.21 on standard English data) and high human ratings (2.80 out of 3.0 in answerability). Sub-optimal tuning had previously degraded model scores by around 2 points in BLEU4 and ROUGE-L. Furthermore, providing full paragraph context rather than single sentences significantly improved question answerability. Crucially, human evaluation revealed that traditional metrics like BLEU4 and ROUGE-L correlate poorly with human judgments, whereas METEOR and MoverScore track answerability much more reliably, and BERTScore excels at tracking grammar and understandability. Finally, cross-domain and multilingual zero-shot transfer yielded poor results; mT5-Small zero-shot BLEU4 dropped to near zero across non-English languages, though initializing models on large English datasets before fine-tuning on domain-specific data achieved the best domain adaptation.
These results indicate that organizations deploying question generation systems should not rely on off-the-shelf zero-shot transfer or default hyperparameter settings. For specialized domains or non-English applications, models require domain-specific fine-tuning and careful optimization. Additionally, evaluation frameworks relying strictly on BLEU or ROUGE risk misjudging system quality, potentially leading to flawed production deployments.
Practitioners should adopt paragraph-level input representations with highlighted answers and implement transfer learning by pre-training on large general datasets before in-domain fine-tuning. Evaluation suites should replace BLEU with METEOR, MoverScore, and BERTScore, or use downstream task performance as a quality proxy. Future development should focus on expanding methods to joint question-answer generation and scaling non-English training resources.
Confidence in these findings is high for standard paragraph contexts in English and moderate-to-high resource languages. However, readers should note that results are constrained to single-hop extractive questions under 500 tokens and may not generalize to long-form documents, complex multi-hop reasoning, or true low-resource languages.
- Paper: BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension, Mike Lewis et al. (2020). Introduces the sequence-to-sequence BART denoising autoencoder architecture, which serves as one of the primary foundational models benchmarked and fine-tuned for paragraph-level question generation.
- Paper: BLEURT: Learning Robust Metrics for Text Generation, Thibault Sellam et al. (2020). Presents learned transformer-based metrics to address the known shortcomings of surface n-gram metrics like BLEU and ROUGE in evaluating generated text quality.
- Paper: How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation, Chia-Wei Liu et al. (2016). Provides foundational empirical evidence on why standard overlap metrics correlate poorly with human judgment in generative sequence tasks.
- Paper: Natural Questions: A Benchmark for Question Answering Research, Tom Kwiatkowski et al. (2019). Establishes standard question answering datasets and paragraph-level context extraction frameworks that inform the target data formulations used in question generation.
- Paper: Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text, Sebastian Gehrmann et al. (2023). Surveys systemic obstacles across natural language generation evaluation protocols and extends the critique of surface overlap metrics like BLEU and ROUGE highlighted in QG-Bench.
- Paper: TRUE: Re-evaluating Factual Consistency Evaluation, Or Honovich et al. (2022). Builds on generative evaluation methodology by standardizing factual consistency benchmarks using question generation and answering frameworks.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). Advances reference-free natural language generation evaluation by using large language models to overcome the correlation bottlenecks of traditional automated metrics.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). Extends multilingual benchmark evaluation across generative models to further assess cross-lingual transfer disparities and translation-based generation pipelines.
- Paper: ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems, Jon Saad-Falcon et al. (2024). Applies synthetic question generation techniques to automate evaluation pipelines and measure faithfulness in retrieval-augmented generation systems.
