keyword
question hierarchy
A question hierarchy is a structured, multi-tiered framework used in natural language processing and question-answering systems to categorize or represent queries across multiple levels of granularity. In question classification, it refers to a layered semantic taxonomy that organizes questions and their expected answer types from broad, coarse classes down into specific, fine-grained categories to guide targeted answer extraction. In computational and multimodal language models, it can also describe a multi-level linguistic architecture that analyzes query text across hierarchical scales—such as words, phrases, and complete sentences—enabling systems to capture both local detail and overall context for more accurate reasoning and retrieval.
2 items

Learning Question Classifiers
Xin Li, Dan Roth
Why you should read this
Presents a hierarchical machine learning approach that classifies open-domain questions into a two-layered taxonomy of 50 fine-grained answer types to guide candidate retrieval and verification in question answering systems.
In order to respond correctly to a free form factual question given a large collection of texts, one needs to understand the question to a level that allows determining some of the constraints the question imposes on a possible answer. These constraints may include a semantic classification of the sought after answer and may even suggest using different strategies when looking for and verifying a candidate answer. This paper presents a machine learning approach to question classification. We learn a hierarchical classifier that is guided by a layered semantic hierarchy of answer types, and eventually classifies questions into fine-grained classes. We show accurate results on a large collection of free-form questions used in TREC 10.
Added
2026-09-24

Hierarchical Question-Image Co-Attention for Visual Question Answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, Devi Parikh
Why you should read this
Introduces a hierarchical co-attention model for visual question answering that jointly computes attention over relevant image regions and multi-level language structures across word, phrase, and question representations.
A number of recent works have proposed attention models for Visual Question Answering (VQA) that generate spatial maps highlighting image regions relevant to answering the question. In this paper, we argue that in addition to modeling "where to look" or visual attention, it is equally important to model "what words to listen to" or question attention. We present a novel co-attention model for VQA that jointly reasons about image and question attention. In addition, our model reasons about the question (and consequently the image via the co-attention mechanism) in a hierarchical fashion via a novel 1-dimensional convolution neural networks (CNN). Our model improves the state-of-the-art on the VQA dataset from 60.3% to 60.5%, and from 61.6% to 63.3% on the COCO-QA dataset. By using ResNet, the performance is further improved to 62.1% for VQA and 65.4% for COCO-QA.
Added
2026-09-24
