keyword
answer prediction
Answer prediction is the computational process in artificial intelligence and natural language processing where an automated model infers and outputs the correct response to a given question using contextual information such as text passages, structured knowledge graphs, or images. Depending on the architecture and task formulation, the process can be framed as extracting a specific span of text from reference documents, selecting from a candidate set of predefined answers, or generating free-form natural language. To produce accurate predictions, models typically combine representation learning, multi-hop reasoning, and attention mechanisms to align the semantics of the query with relevant multimodal or textual evidence.
2 items

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, Akiko Aizawa
Why you should read this
Presents 2WikiMultiHopQA, a benchmark that combines structured Wikidata knowledge with text to guarantee multi-hop questions require true multi-step inference while providing explicit reasoning paths to evaluate model explanations.
A multi-hop question answering (QA) dataset aims to test reasoning and inference skills by requiring a model to read multiple paragraphs to answer a given question. However, current datasets do not provide a complete explanation for the reasoning process from the question to the answer. Further, previous studies revealed that many examples in existing multi-hop datasets do not require multi-hop reasoning to answer a question. In this study, we present a new multi-hop QA dataset, called 2WikiMultiHopQA, which uses structured and unstructured data. In our dataset, we introduce the evidence information containing a reasoning path for multi-hop questions. The evidence information has two benefits: (i) providing a comprehensive explanation for predictions and (ii) evaluating the reasoning skills of a model. We carefully design a pipeline and a set of templates when generating a question-answer pair that guarantees the multi-hop steps and the quality of the questions. We also exploit the structured format in Wikidata and use logical rules to create questions that are natural but still require multi-hop reasoning. Through experiments, we demonstrate that our dataset is challenging for multi-hop models and it ensures that multi-hop reasoning is required.
Added
2026-09-25


Hierarchical Question-Image Co-Attention for Visual Question Answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, Devi Parikh
Why you should read this
Introduces a hierarchical co-attention model for visual question answering that jointly computes attention over relevant image regions and multi-level language structures across word, phrase, and question representations.
A number of recent works have proposed attention models for Visual Question Answering (VQA) that generate spatial maps highlighting image regions relevant to answering the question. In this paper, we argue that in addition to modeling "where to look" or visual attention, it is equally important to model "what words to listen to" or question attention. We present a novel co-attention model for VQA that jointly reasons about image and question attention. In addition, our model reasons about the question (and consequently the image via the co-attention mechanism) in a hierarchical fashion via a novel 1-dimensional convolution neural networks (CNN). Our model improves the state-of-the-art on the VQA dataset from 60.3% to 60.5%, and from 61.6% to 63.3% on the COCO-QA dataset. By using ResNet, the performance is further improved to 62.1% for VQA and 65.4% for COCO-QA.
Added
2026-09-24
