Generative Language Models for Paragraph-Level Question Generation

Asahi UshioFernando Alva-ManchegoJosé Camacho-Collados

article2022EMNLP58 citations

Introduces QG-Bench, a unified multilingual and multi-domain benchmark that standardizes paragraph-level question generation evaluation across eight languages and multiple domains using sequence-to-sequence language models.

Listen

Automatic question generation—creating relevant, natural questions from text given a specific target answer—is critical for training question answering systems, augmenting training data, and building automated educational tools. However, research in this area has suffered from fragmented datasets, non-standardized evaluation protocols, and arbitrary model selections, making it difficult to systematically measure progress and choose optimal architectures.

The article addresses this gap by establishing QG-Bench, a standardized benchmark designed to systematically train, evaluate, and compare sequence-to-sequence language models for paragraph-level question generation across diverse domains and languages.

To construct this benchmark, the authors unified standard datasets into a consistent paragraph-level format spanning English benchmark data, ten specialized domains across two question styles (objective and subjective), and seven non-English languages. Using this data, the study fine-tuned several popular sequence-to-sequence language models (T5, BART, mT5, and mBART) via an optimized two-phase hyperparameter search. The analysis combined standard automated metrics (BLEU4, ROUGE-L, METEOR, BERTScore, and MoverScore) with a large-scale manual human evaluation of 15,000 judgments and an assessment of domain and cross-lingual transferability.

The findings show that properly tuned standard architectures outperform previously reported state-of-the-art specialized models, with T5-Large achieving the strongest performance across metrics (BLEU4 of 27.21 on standard English data) and high human ratings (2.80 out of 3.0 in answerability). Sub-optimal tuning had previously degraded model scores by around 2 points in BLEU4 and ROUGE-L. Furthermore, providing full paragraph context rather than single sentences significantly improved question answerability. Crucially, human evaluation revealed that traditional metrics like BLEU4 and ROUGE-L correlate poorly with human judgments, whereas METEOR and MoverScore track answerability much more reliably, and BERTScore excels at tracking grammar and understandability. Finally, cross-domain and multilingual zero-shot transfer yielded poor results; mT5-Small zero-shot BLEU4 dropped to near zero across non-English languages, though initializing models on large English datasets before fine-tuning on domain-specific data achieved the best domain adaptation.

These results indicate that organizations deploying question generation systems should not rely on off-the-shelf zero-shot transfer or default hyperparameter settings. For specialized domains or non-English applications, models require domain-specific fine-tuning and careful optimization. Additionally, evaluation frameworks relying strictly on BLEU or ROUGE risk misjudging system quality, potentially leading to flawed production deployments.

Practitioners should adopt paragraph-level input representations with highlighted answers and implement transfer learning by pre-training on large general datasets before in-domain fine-tuning. Evaluation suites should replace BLEU with METEOR, MoverScore, and BERTScore, or use downstream task performance as a quality proxy. Future development should focus on expanding methods to joint question-answer generation and scaling non-English training resources.

Confidence in these findings is high for standard paragraph contexts in English and moderate-to-high resource languages. However, readers should note that results are constrained to single-hop extractive questions under 500 tokens and may not generalize to long-form documents, complex multi-hop reasoning, or true low-resource languages.

arXiv: 2210.03992
Cover for Generative Language Models for Paragraph-Level Question Generation

Abstract

Powerful generative models have led to recent progress in question generation (QG). However, it is difficult to measure advances in QG research since there are no standardized resources that allow a uniform comparison among approaches. In this paper, we introduce QG-Bench, a multilingual and multidomain benchmark for QG that unifies existing question answering datasets by converting them to a standard QG setting. It includes general-purpose datasets such as SQuAD (Rajpurkar et al., 2016) for English, datasets from ten domains and two styles, as well as datasets in eight different languages. Using QG-Bench as a reference, we perform an extensive analysis of the capabilities of language models for the task. First, we propose robust QG baselines based on fine-tuning generative language models. Then, we complement automatic evaluation based on standard metrics with an extensive manual evaluation, which in turn sheds light on the difficulty of evaluating QG models. Finally, we analyse both the domain adaptability of these models as well as the effectiveness of multilingual models in languages other than English. QG-Bench is released along with the fine-tuned models presented in the paper,¹ which are also available as a demo.²

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 QG-Bench: A Unified Question Generation Benchmark
  • 3.1 Data Collection and Unification
  • 3.2 Data Statistics
  • 4 LMs for Question Generation
  • 4.1 Task Formulation
  • 4.2 Language Model Fine-tuning
  • 4.3 Experimental Setup
  • 5 Automatic Evaluation
  • 5.1 Evaluation Metrics
  • 5.2 Results
  • 6 Analysis
  • 6.1 Model Input
  • 6.2 Manual Evaluation
  • 6.3 Domain Adaptation
  • 7 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Parameter Optimization
  • A.1 Best Parameters
  • A.2 Fine-tuning without Optimization
  • B Manual Evaluation
  • B.1 Sample Outputs
  • B.2 Spearman's Correlation
  • B.3 William test
  • B.4 Guidelines
  • C Unsupervised QA-based Evaluation
  • D Additional Analysis
  • D.1 Zero-shot Multilingual Transfer
  • D.2 Zero-shot Domain Transfer

Knowls

  1. Knowl 1 — QG-Bench: A Multilingual and Multi-Domain Question Generation Benchmark

    definition

    QG-Bench is a standardized question generation (QG) benchmark that unifies question answering datasets into a consistent four-feature format: (paragraph,sentence,question,answer)(\text{paragraph}, \text{sentence}, \text{question}, \text{answer}), where the answer is conditioned to be an exact sub-string of a sentence within the paragraph.

    The benchmark comprises three sub-categories:

    1. General-Purpose English QG: SQuAD v1.1 based on Wikipedia (75,722 train / 10,570 validation / 11,877 test instances).
    2. Domain-Specific English QG: SQuADShifts (spanning Amazon, Wikipedia, News, and Reddit domains; splits created by reserving half for test and splitting the remainder 1:2 for validation:train) and SubjQA (containing subjective customer review questions across Book, Electronics, Grocery, Movie, Restaurant, and Trip domains).
    3. Multilingual QG: Monolingual SQuAD-style datasets across seven non-English languages: JAQuAD (Japanese), GerQuAD (German), SberQuAD (Russian), KorQuAD (Korean), FQuAD (French), Spanish SQuAD (Spanish), and Italian SQuAD (Italian), with non-overlapping paragraph splits.
  2. Knowl 2 — Highlight Token Formulation for Sequence-to-Sequence Question Generation

    model/method

    Question generation (QG) is framed as conditional sequence generation where a model generates a natural question q^\hat{q} given an input sequence xx by maximizing conditional log-likelihood:

    q^=arg⁡max⁡qP(q∣x)\hat{q} = \arg\max_{q} P(q \mid x)

    To condition the model on a specific target answer a=[a1,…,a∣a∣]a = [a_1, \dots, a_{|a|}] located inside a context passage c=[c1,…,c∣c∣]c = [c_1, \dots, c_{|c|}], the input sequence xx is constructed by inserting special highlight delimiter tokens ⟨hl⟩\langle\text{hl}\rangle surrounding the answer span:

    x=[c1,…,⟨hl⟩,a1,…,a∣a∣,⟨hl⟩,…,c∣c∣]x = [c_1, \dots, \langle\text{hl}\rangle, a_1, \dots, a_{|a|}, \langle\text{hl}\rangle, \dots, c_{|c|}]

    This conditioning admits three structural variants:

    • Paragraph-level (Answer-aware): cc is the entire paragraph containing the target answer aa.
    • Sentence-level (Answer-aware): cc is restricted to the individual sentence containing aa.
    • Answer-free: cc is the entire paragraph, but ⟨hl⟩\langle\text{hl}\rangle delimits the sentence containing the answer rather than the answer itself.
  3. Knowl 3 — Two-Phase Parameter Optimization Algorithm for QG Fine-Tuning

    algorithm

    A two-phase grid search optimization procedure finds robust hyperparameter configurations for fine-tuning sequence-to-sequence language models on question generation tasks.

    Input: Training set DtrainD_{train}, validation set DvalD_{val}, pre-trained model MM
    Input: Hyperparameter search space H=LR×LS×BS\mathcal{H} = \mathcal{LR} \times \mathcal{LS} \times \mathcal{BS}
           where LR={0.0001,0.00005,0.00001}\mathcal{LR} = \{0.0001, 0.00005, 0.00001\}, LS={0.0,0.15}\mathcal{LS} = \{0.0, 0.15\}, BS={64,128,256,512}\mathcal{BS} = \{64, 128, 256, 512\}
    Output: Optimal fine-tuned model M∗M^*
    Phase 1: Coarse Screening
    for each configuration hi∈Hh_i \in \mathcal{H} (24 configurations total) do
        Initialize Mi←MM_i \leftarrow M
        Train MiM_i on DtrainD_{train} with hyperparameter setting hih_i for 2 epochs
        Evaluate BLEU4 score SiS_i of MiM_i on DvalD_{val}
    end for
    Rank configurations by SiS_i and select the top 5 candidates Htop5⊂H\mathcal{H}_{top5} \subset \mathcal{H}
    Phase 2: Full Fine-Tuning
    for each configuration hj∈Htop5h_j \in \mathcal{H}_{top5} do
        Continue training the checkpoint MjM_j on DtrainD_{train} with setting hjh_j until validation performance plateaus
        Record final validation BLEU4 score Sj′S'_j
    end for
    M∗←arg⁡max⁡MjSj′M^* \leftarrow \arg\max_{M_j} S'_j
    return M∗M^*

    Fixed parameters across runs: random seed =1= 1, beam size =4= 4, maximum input token length =512= 512, maximum output token length =34= 34 for fine-tuning and 6464 for evaluation. For T5 models, the prefix string generate question: is prepended to each input text.

  4. Knowl 4 — Comparative Performance of Generative Language Models on SQuAD Question Generation

    data/table

    When fine-tuned using the two-phase parameter optimization on paragraph-level SQuAD v1.1, standard sequence-to-sequence models (T5 and BART) achieve state-of-the-art results compared to previously proposed QG-tailored architectures.

    Model Parameters BLEU4 ROUGE-L METEOR BERTScore MoverScore
    NQG (Du et al.) 30M 12.28 39.75 16.62 - -
    UniLM (Dong et al.) 340M 22.78 51.57 25.49 - -
    UniLMv2 (Bao et al.) 110M 24.70 52.13 26.33 - -
    ProphetNet (Qi et al.) 340M 23.91 52.26 26.60 - -
    ERNIE-GEN (Xiao et al.) 340M 25.40 52.84 26.92 - -
    BARTBASE\text{BART}_{\text{BASE}} 140M 24.68 52.66 26.05 90.87 64.47
    BARTLARGE\text{BART}_{\text{LARGE}} 400M 26.17 53.85 27.07 91.00 64.99
    T5SMALL\text{T5}_{\text{SMALL}} 60M 24.40 51.43 25.84 90.45 63.89
    T5BASE\text{T5}_{\text{BASE}} 220M 26.13 53.33 26.97 90.84 64.74
    T5LARGE\text{T5}_{\text{LARGE}} 770M 27.21 54.13 27.70 91.00 65.29

    T5LARGE\text{T5}_{\text{LARGE}} achieves the highest score across all five automatic metrics. T5SMALL\text{T5}_{\text{SMALL}} (60M parameters) attains performance comparable to UniLMv2\text{UniLMv2} (110M parameters) while using nearly half the parameters.

  5. Knowl 5 — Correlation Analysis Between Automatic Metrics and Human Judgments in Question Generation

    empirical result

    Spearman rank correlation coefficients between automatic metrics (BLEU4, ROUGE-L, METEOR, BERTScore with RoBERTa-large, and MoverScore with DistilBERT-base) and human ratings across 3,000 generated questions (500 SQuAD examples across 6 QG models evaluated by 5 native English judges) demonstrate metric-specific alignments:

    • Answerability correlation: METEOR (r=0.41r = 0.41) and MoverScore (r=0.39r = 0.39) exhibit the highest correlation with human judgment of whether the question can be answered by the target answer, followed by ROUGE-L (r=0.36r = 0.36), BLEU4 (r=0.34r = 0.34), and BERTScore (r=0.34r = 0.34).
    • Grammaticality correlation: BERTScore (r=0.23r = 0.23) and MoverScore (r=0.21r = 0.21) correlate most strongly, followed by METEOR (r=0.20r = 0.20), ROUGE-L (r=0.19r = 0.19), and BLEU4 (r=0.18r = 0.18).
    • Understandability correlation: BERTScore (r=0.30r = 0.30) and MoverScore (r=0.28r = 0.28) achieve the highest correlation, followed by METEOR (r=0.27r = 0.27), ROUGE-L (r=0.25r = 0.25), and BLEU4 (r=0.23r = 0.23).

    All correlation differences are statistically significant according to Williams tests (p<0.005p < 0.005). Traditional n-gram overlap metrics (BLEU4 and ROUGE-L) show the weakest correlation across all three human evaluation criteria.

  6. Knowl 6 — Human Evaluation of Question Generation Across Grammaticality, Understandability, and Answerability

    data/table

    Human evaluation of 500 randomly sampled SQuAD test instances evaluated on a 1-to-3 scale by 5 judges per sample (15,000 judgments total) reveals that answerability is the primary criterion separating model performance, while grammaticality and understandability remain high across pre-trained language models.

    Model Ans. Gra. Und. BLEU4 ROUGE-L METEOR BERTScore MoverScore
    NQG (LSTM) 1.21 2.35 2.63 3.33 14.30 33.53 88.27 58.25
    BARTLARGE\text{BART}_{\text{LARGE}} 2.70 2.89 2.93 16.15 29.93 51.35 90.95 65.44
    T5SMALL\text{T5}_{\text{SMALL}} 2.51 2.83 2.90 13.43 27.38 48.86 90.41 64.27
    T5LARGE\text{T5}_{\text{LARGE}} 2.80 2.93 2.95 17.56 30.42 52.00 90.94 66.09
    T5LARGE\text{T5}_{\text{LARGE}} (sent-level) 2.47 2.91 2.95 14.88 27.49 48.97 90.76 64.53
    T5LARGE\text{T5}_{\text{LARGE}} (answer-free) 2.46 2.91 2.95 13.62 26.82 47.37 90.20 64.00

    Inter-annotator agreement (Fleiss's kappa) is 0.610.61 for Answerability (substantial agreement), 0.360.36 for Understandability (fair agreement), and 0.300.30 for Grammaticality (fair agreement).

  7. Knowl 7 — Effect of Context Granularity and Answer Conditioning on Question Generation Quality

    empirical result

    Paragraph-level answer-aware QG models consistently outperform sentence-level and answer-free variants across model architectures and metric evaluations on SQuAD:

    1. Paragraph-level vs. Sentence-level: Full paragraph inputs outperform sentence-only inputs (e.g., T5LARGE\text{T5}_{\text{LARGE}} achieves BLEU4 27.21 paragraph-level vs. 25.36 sentence-level; ROUGE-L 54.13 vs. 52.53; METEOR 27.70 vs. 26.28). This indicates that models effectively leverage global discourse context rather than relying solely on local sentence cues.
    2. Answer-aware vs. Answer-free: Conditioning on the explicit answer span via highlight tokens is critical for targeted question generation (e.g., T5LARGE\text{T5}_{\text{LARGE}} answer-free drops to BLEU4 24.27, ROUGE-L 51.30, METEOR 25.67). In human evaluation, dropping the answer conditioning decreases answerability from 2.80 to 2.46 for T5LARGE\text{T5}_{\text{LARGE}}, matching the drop observed when reducing context from paragraph to sentence level (2.47).
  8. Knowl 8 — Domain Adaptation Strategies for Question Generation

    empirical result

    Evaluating three adaptation protocols across ten domain splits (SQuADShifts: Amazon, Wikipedia, News, Reddit; SubjQA: Book, Electronics, Grocery, Movie, Restaurant, Trip) demonstrates the following:

    1. Initialization Effect: Initializing the QG model with SQuAD pre-fine-tuning and subsequently fine-tuning on in-domain training data consistently outperforms both (a) fine-tuning purely on in-domain data without SQuAD initialization and (b) direct zero-shot transfer from a SQuAD fine-tuned model.
    2. Style Shift vs. Domain Shift: For SQuADShifts (which shares SQuAD's extractive, factual question style), zero-shot transfer from SQuAD retains relatively high METEOR scores. For SubjQA (which features subjective questions from user reviews), zero-shot SQuAD transfer fails severely, showing that question-style distribution shifts present a much harder obstacle for QG transfer than topical domain shifts.
  9. Knowl 9 — Multilingual Question Generation Performance and Cross-Lingual Transfer Limitations

    data/table

    Multilingual sequence-to-sequence models (extmT5extSMALL ext{mT5}_{ ext{SMALL}}, extmT5extBASE ext{mT5}_{ ext{BASE}}, and extmBART ext{mBART}) fine-tuned on native monolingual datasets achieve lower scores than English models, driven largely by training set size limitations. In addition, zero-shot cross-lingual transfer from English SQuAD fails catastrophically across target languages.

    Setting / Target Language Model BLEU4 ROUGE-L METEOR BERTScore MoverScore
    Monolingual Fine-Tuned
    English (SQuAD) mT5BASE\text{mT5}_{\text{BASE}} 23.03 50.67 25.18 90.23 63.60
    Japanese (JAQuAD) mT5BASE\text{mT5}_{\text{BASE}} 32.54 52.67 30.58 81.77 59.68
    Russian (SberQuAD) mBART\text{mBART} 18.80 34.18 29.30 87.18 65.88
    Korean (KorQuAD) mT5BASE\text{mT5}_{\text{BASE}} 12.18 28.57 29.62 84.52 83.36
    Spanish (Spanish SQuAD) mT5BASE\text{mT5}_{\text{BASE}} 10.15 25.45 23.43 84.47 59.62
    Italian (Italian SQuAD) mT5BASE\text{mT5}_{\text{BASE}} 7.70 22.51 18.00 81.16 57.11
    French (FQuAD) mT5SMALL\text{mT5}_{\text{SMALL}} 8.55 28.56 17.51 80.71 56.50
    German (GerQuAD) mT5BASE\text{mT5}_{\text{BASE}} 0.87 11.10 13.65 80.39 55.73
    Zero-Shot Transfer from English mT5SMALL\text{mT5}_{\text{SMALL}}
    Russian 0.00 0.99 1.78 70.89 49.10
    Japanese 0.00 6.08 0.51 66.08 46.53
    Italian 0.54 5.01 5.89 72.60 50.23
    Korean 0.00 0.06 0.73 66.34 45.86
    Spanish 0.59 5.21 6.02 74.94 50.62
    German 0.00 1.56 4.81 73.53 50.37
    French 1.71 15.84 8.24 72.91 50.96
  10. Knowl 10 — Unsupervised Question Answering-Based Evaluation of QG Models

    model/method

    Unsupervised QA-based evaluation measures the intrinsic quality of a QG model by training an extractive QA model entirely on synthetic question-answer pairs produced by the QG model, evaluating downstream QA accuracy on human-annotated data.

    1. Data Generation: The target QG model generates questions for all 1,000,000 paragraph-answer pairs collected from Wikipedia.
    2. QA Training: A BERTBASE-CASED\text{BERT}_{\text{BASE-CASED}} model is fine-tuned strictly on the resulting 1M synthetic QA instances.
    3. Evaluation: Downstream question answering performance is measured via Exact Match (EM) and F1 scores on the official SQuAD validation set.

    Evaluation results reflect QG model generation capability:

    • Synthetic data from T5LARGE\text{T5}_{\text{LARGE}} yields 70.8670.86 F1 / 59.0459.04 EM.
    • Synthetic data from BARTLARGE\text{BART}_{\text{LARGE}} yields 70.4070.40 F1 / 58.6058.60 EM.
    • Synthetic data from T5BASE\text{T5}_{\text{BASE}} yields 70.3370.33 F1 / 58.1458.14 EM.
    • Synthetic data from BARTBASE\text{BART}_{\text{BASE}} yields 70.1070.10 F1 / 58.4658.46 EM.
    • Synthetic data from T5SMALL\text{T5}_{\text{SMALL}} yields 68.9068.90 F1 / 56.9656.96 EM.

Coverage note — None was omitted; the knowls cover the unified benchmark design, LM fine-tuning formulation, optimization algorithm, main SQuAD results, human evaluation and metric correlations, input ablations, domain adaptation, multilingual baselines, and unsupervised QA-based evaluation.

References

  1. 1.Fernando Alva-Manchego, Carolina Scarton, and Lucia Specia. 2021. The (un)suitability of automatic evaluation metrics for text simplification. Computational Linguistics, 47(4):861–889.
  2. 2.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4623–4637, Online. Association for Computational Linguistics.
  3. 3.Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, Nan Yang, Xiaodong Liu, Yu Wang, Jianfeng Gao, Songhao Piao, Ming Zhou, et al. 2020. Unilmv2: Pseudo-masked language models for unified language model pre-training. In International Conference on Machine Learning, pages 642–652. PMLR.
  4. 4.Max Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel, Pontus Stenetorp, and Douwe Kiela. 2021. Improving question answering model robustness with synthetic adversarial data generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8830–8848, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  5. 5.Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020. Reevaluating evaluation in text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9347–9359, Online. Association for Computational Linguistics.
  6. 6.Johannes Bjerva, Nikita Bhutani, Behzad Golshan, Wang-Chiew Tan, and Isabelle Augenstein. 2020. SubjQA: A Dataset for Subjectivity and Review Comprehension. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5480–5494, Online. Association for Computational Linguistics.
  7. 7.Shuyang Cao and Lu Wang. 2021. Controllable open-ended question generation with a new question type ontology. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6424–6439, Online. Association for Computational Linguistics.
  8. 8.Carrino Casimiro Pio, Costa-jussa Marta R., and Fonollosa Jose A. R. 2019. Automatic Spanish Translation of the SQuAD Dataset for Multilingual Question Answering. arXiv e-prints, page arXiv:1912.05200v1.
  9. 9.Ying-Hong Chan and Yao-Chung Fan. 2019. A recurrent BERT-based model for question generation. In Proceedings of the 2nd Workshop on Machine Reading for Question Answering, pages 154–162, Hong Kong, China. Association for Computational Linguistics.
  10. 10.Yiran Chen, Zhenqiao Song, Xianze Wu, Danqing Wang, Jingjing Xu, Jiaze Chen, Hao Zhou, and Lei Li. 2021. Mtg: A benchmarking suite for multilingual text generation. arXiv preprint arXiv:2108.07140.
  11. 11.Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020. TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Linguistics, 8:454–470.
  12. 12.Danilo Croce, Alexandra Zelenanska, and Roberto Basili. 2018. Neural learning for question answering in italian. In AIIA 2018 – Advances in Artificial Intelligence*, pages 389–402, Cham. Springer International Publishing.
  13. 13.Michael Denkowski and Alon Lavie. 2014. Meteor universal: Language specific translation evaluation for any target language. In Proceedings of the Ninth Workshop on Statistical Machine Translation, pages 376–380, Baltimore, Maryland, USA. Association for Computational Linguistics.
  14. 14.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  15. 15.Martin d'Hoffschmidt, Wacim Belblidia, Quentin Heinrich, Tom Brendlé, and Maxime Vidal. 2020. FQuAD: French question answering dataset. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1193–1208, Online. Association for Computational Linguistics.
  16. 16.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019. Unified language model pre-training for natural language understanding and generation. Advances in Neural Information Processing Systems, 32.
  17. 17.Xinya Du and Claire Cardie. 2018. Harvesting paragraph-level question-answer pairs from Wikipedia. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1907–1917, Melbourne, Australia. Association for Computational Linguistics.
  18. 18.Xinya Du, Junru Shao, and Claire Cardie. 2017. Learning to ask: Neural question generation for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1342–1352, Vancouver, Canada. Association for Computational Linguistics.
  19. 19.Pavel Efimov, Andrey Chertok, Leonid Boytsov, and Pavel Braslavski. 2020. Sberquad–russian reading comprehension dataset: Description and analysis. In International Conference of the Cross-Language Evaluation Forum for European Languages, pages 3–15. Springer.
  20. 20.Michael Heilman and Noah A. Smith. 2010. Good question! statistical ranking for question generation. In Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, pages 609–617, Los Angeles, California. Association for Computational Linguistics.
  21. 21.Gautier Izacard and Edouard Grave. 2020a. Distilling knowledge from reader to retriever for question answering.
  22. 22.Gautier Izacard and Edouard Grave. 2020b. Leveraging passage retrieval with generative models for open domain question answering.
  23. 23.Robin Jia, Mike Lewis, and Luke Zettlemoyer. 2021. Question answering infused pre-training of general-purpose contextualized representations. arXiv preprint arXiv:2106.08190.
  24. 24.Ranjay Krishna, Michael Bernstein, and Li Fei-Fei. 2019. Information maximizing visual question generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2008–2018.
  25. 25.Igor Labutov, Sumit Basu, and Lucy Vanderwende. 2015. Deep questions without deep understanding. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 889–898, Beijing, China. Association for Computational Linguistics.
  26. 26.J Richard Landis and Gary G Koch. 1977. The measurement of observer agreement for categorical data. biometrics, pages 159–174.
  27. 27.Dong Bok Lee, Seanie Lee, Woo Tae Jeong, Donghwan Kim, and Sung Ju Hwang. 2020. Generating diverse and consistent QA pairs from contexts with information-maximizing hierarchical conditional VAEs. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 208–224, Online. Association for Computational Linguistics.
  28. 28.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020a. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  29. 29.Patrick Lewis, Ludovic Denoyer, and Sebastian Riedel. 2019. Unsupervised question answering by cloze translation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4896–4910, Florence, Italy. Association for Computational Linguistics.
  30. 30.Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020b. MLQA: Evaluating cross-lingual extractive question answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7315–7330, Online. Association for Computational Linguistics.
  31. 31.Patrick Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Küttler, Aleksandra Piktus, Pontus Stenetorp, and Sebastian Riedel. 2021. PAQ: 65 million probably-asked questions and what you can do with them. Transactions of the Association for Computational Linguistics, 9:1098–1115.
  32. 32.Seungyoung Lim, Myungji Kim, and Jooyoul Lee. 2019. Korquad1. 0: Korean qa dataset for machine reading comprehension. arXiv preprint arXiv:1909.07005.
  33. 33.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  34. 34.David Lindberg, Fred Popowich, John Nesbit, and Phil Winne. 2013. Generating natural language questions to support learning on-line. In Proceedings of the 14th European Workshop on Natural Language Generation, pages 105–114, Sofia, Bulgaria. Association for Computational Linguistics.
  35. 35.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742.
  36. 36.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  37. 37.Luis Enrico Lopez, Diane Kathryn Cruz, Jan Christian Blaise Cruz, and Charibeth Cheng. 2020. Transformer-based end-to-end question generation. arXiv preprint arXiv:2005.01107, 4.
  38. 38.John Miller, Karl Krauth, Benjamin Recht, and Ludwig Schmidt. 2020. The effect of natural distribution shift on question answering models. In International Conference on Machine Learning, pages 6905–6916. PMLR.
  39. 39.Ruslan Mitkov and Le An Ha. 2003. Computer-aided generation of multiple-choice tests. In Proceedings of the HLT-NAACL 03 Workshop on Building Educational Applications Using Natural Language Processing, pages 17–22.
  40. 40.Timo Möller, Julian Risch, and Malte Pietsch. 2021. Germanquad and germandpr: Improving non-english question answering and passage retrieval.
  41. 41.Preksha Nema and Mitesh M. Khapra. 2018. Towards a better metric for evaluating question generation systems. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3950–3959, Brussels, Belgium. Association for Computational Linguistics.
  42. 42.Liangming Pan, Yuxi Xie, Yansong Feng, Tat-Seng Chua, and Min-Yen Kan. 2020. Semantic graphs for generating deep questions. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1463–1475, Online. Association for Computational Linguistics.
  43. 43.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  44. 44.Bhargavi Paranjape, Matthew Lamm, and Ian Tenney. 2021. Retrieval-guided counterfactual generation for qa. arXiv preprint arXiv:2110.07596.
  45. 45.Ethan Perez, Patrick Lewis, Wen-tau Yih, Kyunghyun Cho, and Douwe Kiela. 2020. Unsupervised question decomposition for question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8864–8880, Online. Association for Computational Linguistics.
  46. 46.Raul Puri, Ryan Spring, Mohammad Shoeybi, Mostofa Patwary, and Bryan Catanzaro. 2020. Training question answering models from synthetic data. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5811–5826, Online. Association for Computational Linguistics.
  47. 47.Valentina Pyatkin, Paul Roit, Julian Michael, Yoav Goldberg, Reut Tsarfaty, and Ido Dagan. 2021. Asking it all: Generating contextualized questions for any semantic role. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1429–1441, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  48. 48.Weizhen Qi, Yu Yan, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, and Ming Zhou. 2020. ProphetNet: Predicting future n-gram for sequence-to-SequencePre-training. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 2401–2410, Online. Association for Computational Linguistics.
  49. 49.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  50. 50.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  51. 51.Ehud Reiter. 2018. A structured review of the validity of BLEU. Computational Linguistics, 44(3):393–401.
  52. 52.Vasile Rus, Brendan Wyse, Paul Piwek, Mihai Lintean, Svetlana Stoyanchev, and Christian Moldovan. 2010. The first question generation shared task evaluation challenge. In Proceedings of the 6th International Natural Language Generation Conference. Association for Computational Linguistics.
  53. 53.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108.
  54. 54.Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021. Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in NLP. Transactions of the Association for Computational Linguistics, 9:1408–1424.
  55. 55.Siamak Shakeri, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Feng Nan, Zhiguo Wang, Ramesh Nallapati, and Bing Xiang. 2020. End-to-end synthetic data generation for domain adaptation of question answering systems. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5445–5460, Online. Association for Computational Linguistics.
  56. 56.ByungHoon So, Kyuhong Byun, Kyungwon Kang, and Seongjin Cho. 2022. Jaquad: Japanese question answering dataset for machine reading comprehension. arXiv preprint arXiv:2202.01764.
  57. 57.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neural networks. Advances in neural information processing systems, 27.
  58. 58.Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman. 2017. NewsQA: A machine comprehension dataset. In Proceedings of the 2nd Workshop on Representation Learning for NLP, pages 191–200, Vancouver, Canada. Association for Computational Linguistics.
  59. 59.George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, et al. 2015. An overview of the bioasq large-scale biomedical semantic indexing and question answering competition. BMC bioinformatics, 16(1):1–28.
  60. 60.Shuohang Wang and Jing Jiang. 2016. Machine comprehension using match-lstm and answer pointer. arXiv preprint arXiv:1608.07905.
  61. 61.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  62. 62.Dongling Xiao, Han Zhang, Yukun Li, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2021. Ernie-gen: an enhanced multi-flow pre-training and fine-tuning framework for natural language generation. In Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence, pages 3997–4003.
  63. 63.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  64. 64.Zhilin Yang, Bhuwan Dhingra, Ye Yuan, Junjie Hu, William W Cohen, and Ruslan Salakhutdinov. 2017. Words or characters? fine-grained gating for reading comprehension. In ICLR (Poster).
  65. 65.Shiyue Zhang and Mohit Bansal. 2019. Addressing semantic drift in question generation for semi-supervised question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2495–2509, Hong Kong, China. Association for Computational Linguistics.
  66. 66.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.
  67. 67.Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019. MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 563–578, Hong Kong, China. Association for Computational Linguistics.
  68. 68.Qingyu Zhou, Nan Yang, Furu Wei, Chuanqi Tan, Hangbo Bao, and Ming Zhou. 2017. Neural question generation from text: A preliminary study. In National CCF Conference on Natural Language Processing and Chinese Computing, pages 662–671. Springer.

Citation

MLA
Ushio, A., et al. “Generative Language Models for Paragraph-Level Question Generation”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 670–88, https://doi.org/10.18653/v1/2022.emnlp-main.42.
APA
Ushio, A., Alva-Manchego, F., & Camacho-Collados, J. (2022). Generative Language Models for Paragraph-Level Question Generation. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 670–688. https://doi.org/10.18653/v1/2022.emnlp-main.42
Chicago
Ushio, A., F. Alva-Manchego, and J. Camacho-Collados. 2022. “Generative Language Models for Paragraph-Level Question Generation”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 670–88. https://doi.org/10.18653/v1/2022.emnlp-main.42.
Harvard
Ushio, A., Alva-Manchego, F. and Camacho-Collados, J. (2022) “Generative Language Models for Paragraph-Level Question Generation”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 670–688. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.42.
Vancouver
1. Ushio A, Alva-Manchego F, Camacho-Collados J (2022) Generative Language Models for Paragraph-Level Question Generation. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 670–688

BibTeX

@inproceedings{ushio-etal-2022-generative,
    title = "Generative Language Models for Paragraph-Level Question Generation",
    author = "Ushio, Asahi  and
      Alva-Manchego, Fernando  and
      Camacho-Collados, Jose",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.42/",
    doi = "10.18653/v1/2022.emnlp-main.42",
    pages = "670--688"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/