RNG-KBQA: Generation Augmented Iterative Ranking for Knowledge Base Question Answering
Xi YeSemih YavuzKazuma HashimotoYingbo ZhouCaiming Xiong
Modern knowledge bases contain massive stores of structured information, but querying them typically requires specialized database languages. Natural language question answering systems aim to make these databases accessible to non-technical users. However, existing methods struggle to answer complex queries that involve new combinations of concepts or entirely unseen database elements. Pure generation models fail to invent unseen schema items reliably, while ranking-based systems struggle to cover the massive combinatorial space of possible database queries.
The article demonstrates and evaluates a hybrid "Rank-and-Generate" framework, named RnG-KBQA, designed to overcome these coverage and generalization bottlenecks. The primary objective is to evaluate whether coupling an iterative contrastive ranking model with a sequence-to-sequence generation model can accurately answer questions across both familiar and novel domains.
The researchers designed a two-stage approach using pre-trained language models. First, candidate database queries are enumerated up to two steps away from identified entities and evaluated using a contrastive bi-encoder ranker trained via iterative negative bootstrapping. Second, a sequence-to-sequence generator takes the user's question and the top-ranked candidate queries to synthesize the final executable query, effectively repairing missing constraints or operations. An execution-guided decoding fallback ensures that returned queries are syntactically valid and executable. The system was benchmarked on two primary datasets: GRAILQA, which explicitly tests generalization across standard, compositional, and zero-shot settings, and WEBQSP, a standard benchmark for multi-hop question answering.
The evaluation produced four key findings. First, the proposed framework established a new state of the art on GRAILQA, achieving an exact match score of 68.8% and an F1 score of 74.4%, outperforming the previous leading baseline by 10.7 exact match points and 9.1 F1 points. Second, the system demonstrated exceptional zero-shot generalization on unseen schema items, exceeding the prior best baseline by 16.7 F1 points (69.2% versus 52.5%). Third, on the WEBQSP benchmark, the framework attained a top-performing 75.6% F1 score, surpassing earlier systems that relied on perfect oracle entity linking. Fourth, ablation analyses showed that omitting either the ranker or the generator caused substantial performance drops of up to 27.5 and 5.3 F1 points, respectively, confirming that both stages are necessary for robust performance.
These findings indicate that combining candidate ranking with generative refinement resolves the trade-off between search space coverage and generalization. Rather than requiring exhaustive rule enumeration or relying on unconstrained generation, systems can leverage rankers to retrieve relevant schema context and use generators to compose precise logic. For enterprise operations, this approach reduces the cost and risk of deploying natural language interfaces across evolving databases without requiring extensive retraining for new data schemas.
Organizations developing automated querying or knowledge base interfaces should adopt two-stage rank-and-generate architectures over single-stage generative or ranking pipelines. Engineering teams should also incorporate execution-guided validation during inference to guarantee valid database outputs. Future technical work should focus on extending candidate generation beyond two-hop paths and exploring constrained decoding mechanisms to further minimize false constraints in zero-shot queries.
The findings are supported by strong benchmark performance across diverse query types. However, users should note key limitations: the generation stage occasionally introduces incorrect constraints in ambiguous queries, and performance gains are less pronounced when queries involve zero-shot relations that cannot be captured in initial candidate enumeration. Overall, confidence in the methodology is high for both standard and compositional querying tasks.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). This foundational paper establishes the retrieval-augmented generation paradigm that RnG-KBQA builds upon and adapts for structured knowledge base question answering.
- Paper: Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering, Gautier Izacard et al. (2021). This work introduces the Fusion-in-Decoder architecture for conditioning sequence-to-sequence generators on multiple retrieved passages, providing the conceptual foundation for RnG-KBQA's generation stage.
- Paper: REALM: Retrieval-Augmented Language Model Pre-Training, Kelvin Guu et al. (2020). It provides foundational principles for integrating explicit neural retrieval with language generation, directly informing rank-and-generate frameworks.
- Paper: A Review of Relational Machine Learning for Knowledge Graphs, Maximilian Nickel et al. (2015). This paper offers essential background on relational machine learning and path-ranking methods over large knowledge graphs that underpin KBQA candidate generation.
- Paper: TIARA: Multi-grained Retrieval for Robust Question Answering over Large Knowledge Base, Yiheng Shu et al. (2022). TIARA builds directly upon rank-and-generate KBQA methods by introducing multi-grained retrieval and constrained decoding on the same GrailQA and WebQSP benchmarks.
- Paper: FlexKBQA: A Flexible LLM-Powered Framework for Few-Shot Knowledge Base Question Answering, Zhenyu Li et al. (2024). FlexKBQA extends knowledge base question answering to few-shot regimes by using language models for synthetic query generation and execution-guided self-training.
- Paper: Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs, Liyi Chen et al. (2024). Plan-on-Graph advances complex multi-hop KBQA by introducing adaptive path exploration and dynamic self-correction on benchmarks like GrailQA and WebQSP.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). RankRAG generalizes the interplay of candidate ranking and answer generation into a unified instruction-tuned language model architecture.
- Paper: Knowledge Base Question Answering by Case-based Reasoning over Subgraphs, Rajarshi Das et al. (2022). CBR-SUBG explores an alternative non-parametric approach to KBQA by performing case-based reasoning over retrieved subgraphs without requiring full logical form generation.
