MMCoQA: Conversational Question Answering over Text, Tables, and Images
Yongqi LiWenjie LiLiqiang Nie
Introduces multimodal conversational question answering across text, tables, and images alongside the new MMConvQA benchmark dataset and an end-to-end baseline framework to support multi-turn reasoning over diverse knowledge sources.
Modern conversational assistants must handle complex, multi-turn dialogues, yet existing systems predominantly rely on single knowledge sources such as text passages or structured knowledge graphs. This creates a critical limitation in practical environments where answering user inquiries requires navigating diverse media, including visual content, tables, and unstructured documents. As users naturally shift topics and ask follow-up questions across multiple formats, conversational interfaces must dynamically determine the appropriate medium, verify consistent facts across sources, and combine complementary information from different formats to generate accurate responses.
The article addresses this operational gap by introducing the multimodal conversational question answering task and establishing a dedicated benchmark dataset along with a baseline architecture. The study aims to evaluate how effectively conversational artificial intelligence systems can retrieve relevant evidence from mixed-format repositories and accurately extract answers across multi-turn interactions.
To establish a standard benchmark, the authors developed a dataset comprising 1,179 conversations and 5,753 question-answer pairs linked to an open-domain repository of over 218,000 text passages, 10,000 structured tables, and 57,000 images. The dataset was constructed by decomposing complex multi-step questions into coherent dialogue sequences and annotating each turn with corresponding answers, supporting evidence, and self-contained question rewrites. The authors also developed an end-to-end baseline framework that processes contextual questions, conducts dense retrieval across modalities using independent neural encoders, detects the target modality, and extracts final answers through specialized text, table, and image extractors.
Key findings demonstrate that existing text-only conversational models and single-turn multimodal architectures fail to handle this combined task effectively. Across the full search repository, the baseline framework achieved a retrieval recall of 41.53% among the top-2,000 candidates, but answer extraction accuracy remained low, with an exact match rate of only 1.36%. However, when ground-truth evidence was manually supplied to bypass the retrieval step, exact match accuracy rose significantly to 22.03%, highlighting that multimodal evidence retrieval is currently the primary performance bottleneck. Furthermore, performance on image-based questions was substantially lower than on text or table inquiries, and embedding visualizations revealed that visual representations remained isolated rather than properly aligned with semantic concepts in text and tables.
These results indicate that deploying real-world conversational systems across heterogeneous enterprise knowledge bases carries substantial risks of retrieval failure and inaccurate answers under current architectures. Standard dense retrieval techniques developed for text documents do not transfer seamlessly to multimodal data. For decision-makers and technical leaders, this means conversational assistants cannot reliably perform cross-media reasoning without targeted algorithmic improvements.
Organizations developing multimodal conversational agents should prioritize research into unified multimodal embedding spaces to improve cross-media retrieval and enhance visual extraction modules. Incorporating automated question-rewriting and pretraining on conversational data provide measurable performance gains and should be integrated into system pipelines. Future work should focus on scaling dataset diversity and addressing remaining limitations, such as the relatively modest conversational sample size and the challenges of multi-hop cross-modal reasoning.
- Paper: CoQA: A Conversational Question Answering Challenge, Siva Reddy et al. (2018). It introduces the foundational conversational question answering challenge (CoQA) that established multi-turn dialogue over knowledge passages upon which MMCoQA adds multimodal evidence.
- Paper: VQA: Visual Question Answering, Stanislaw Antol et al. (2015). It defines the core visual question answering task and baseline paradigms that ground MMCoQA's visual retrieval and reasoning components.
- Paper: OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge, Kenneth Marino et al. (2019). It formalizes visual question answering requiring external knowledge retrieval, directly motivating MMCoQA's multi-source evidence extraction setup.
- Paper: Hierarchical Question-Image Co-Attention for Visual Question Answering, Jiasen Lu et al. (2016). It presents hierarchical co-attention mechanisms across questions and visual inputs, supplying core architectural concepts for multimodal question-evidence fusion.
- Paper: Multimodal Machine Learning: A Survey and Taxonomy, Tadas Baltrušaitis et al. (2017). It provides a taxonomy of representation, fusion, and alignment across heterogeneous modalities that frames MMCoQA's multimodal evidence retrieval design.
- Paper: PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers, Weizhe Lin et al. (2024). It directly advances multi-modal evidence retrieval across text and image formats, scaling up the retrieval techniques essential for multimodal question answering systems like MMCoQA.
- Paper: Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond, Jinze Bai et al. (2023). It develops an end-to-end conversational vision-language model (Qwen-VL-Chat) capable of dialogue, text reading, and visual localization across complex multimodal QA interactions.
- Paper: Improved Baselines with Visual Instruction Tuning, Haotian Liu et al. (2024). It enhances visual instruction tuning baselines to improve conversational and fine-grained visual question answering in multimodal models.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). It builds a unified large multimodal framework that generalizes conversational visual question answering across single-image, multi-image, and video contexts.
- Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). It extends conversational multimodal question answering by benchmarking long-term memory, temporal reasoning, and multi-session dialogue consistency in LLM agents.
- Paper: M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought, Qiguang Chen et al. (2024). It extends multimodal question answering to complex multi-step and multi-domain chain-of-thought reasoning across textual and visual sources.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). It provides a comprehensive survey of modern multimodal large language models, synthesizing advances in the multimodal dialogue and reasoning tasks initiated by works like MMCoQA.
