MMCoQA: Conversational Question Answering over Text, Tables, and Images

Yongqi LiWenjie LiLiqiang Nie

article2022ACL55 citations

Introduces multimodal conversational question answering across text, tables, and images alongside the new MMConvQA benchmark dataset and an end-to-end baseline framework to support multi-turn reasoning over diverse knowledge sources.

Listen

Modern conversational assistants must handle complex, multi-turn dialogues, yet existing systems predominantly rely on single knowledge sources such as text passages or structured knowledge graphs. This creates a critical limitation in practical environments where answering user inquiries requires navigating diverse media, including visual content, tables, and unstructured documents. As users naturally shift topics and ask follow-up questions across multiple formats, conversational interfaces must dynamically determine the appropriate medium, verify consistent facts across sources, and combine complementary information from different formats to generate accurate responses.

The article addresses this operational gap by introducing the multimodal conversational question answering task and establishing a dedicated benchmark dataset along with a baseline architecture. The study aims to evaluate how effectively conversational artificial intelligence systems can retrieve relevant evidence from mixed-format repositories and accurately extract answers across multi-turn interactions.

To establish a standard benchmark, the authors developed a dataset comprising 1,179 conversations and 5,753 question-answer pairs linked to an open-domain repository of over 218,000 text passages, 10,000 structured tables, and 57,000 images. The dataset was constructed by decomposing complex multi-step questions into coherent dialogue sequences and annotating each turn with corresponding answers, supporting evidence, and self-contained question rewrites. The authors also developed an end-to-end baseline framework that processes contextual questions, conducts dense retrieval across modalities using independent neural encoders, detects the target modality, and extracts final answers through specialized text, table, and image extractors.

Key findings demonstrate that existing text-only conversational models and single-turn multimodal architectures fail to handle this combined task effectively. Across the full search repository, the baseline framework achieved a retrieval recall of 41.53% among the top-2,000 candidates, but answer extraction accuracy remained low, with an exact match rate of only 1.36%. However, when ground-truth evidence was manually supplied to bypass the retrieval step, exact match accuracy rose significantly to 22.03%, highlighting that multimodal evidence retrieval is currently the primary performance bottleneck. Furthermore, performance on image-based questions was substantially lower than on text or table inquiries, and embedding visualizations revealed that visual representations remained isolated rather than properly aligned with semantic concepts in text and tables.

These results indicate that deploying real-world conversational systems across heterogeneous enterprise knowledge bases carries substantial risks of retrieval failure and inaccurate answers under current architectures. Standard dense retrieval techniques developed for text documents do not transfer seamlessly to multimodal data. For decision-makers and technical leaders, this means conversational assistants cannot reliably perform cross-media reasoning without targeted algorithmic improvements.

Organizations developing multimodal conversational agents should prioritize research into unified multimodal embedding spaces to improve cross-media retrieval and enhance visual extraction modules. Incorporating automated question-rewriting and pretraining on conversational data provide measurable performance gains and should be integrated into system pipelines. Future work should focus on scaling dataset diversity and addressing remaining limitations, such as the relatively modest conversational sample size and the challenges of multi-hop cross-modal reasoning.

Li et al (2022).pdf
Cover for MMCoQA: Conversational Question Answering over Text, Tables, and Images

Abstract

The rapid development of conversational assistants accelerates the study on conversational question answering (QA). However, the existing conversational QA systems usually answer users' questions with a single knowledge source, e.g., paragraphs or a knowledge graph, but overlook the important visual cues, let alone multiple knowledge sources of different modalities. In this paper, we hence define a novel research task, i.e., multimodal conversational question answering (MMCoQA), aiming to answer users' questions with multi-modal knowledge sources via multi-turn conversations. This new task brings a series of research challenges, including but not limited to priority, consistency, and complementarity of multimodal knowledge. To facilitate the data-driven approaches in this area, we construct the first multimodal conversational QA dataset, named MMConvQA. Questions are fully annotated with not only natural language answers but also the corresponding evidence and valuable decontextualized self-contained questions. Meanwhile, we introduce an end-to-end baseline model, which divides this complex research task into question understanding, multi-modal evidence retrieval, and answer extraction. Moreover, we report a set of benchmarking results, and the results indicate that there is ample room for improvement.

Table of Contents

  • 2 Related Work
  • 3 Dataset Construction
  • 4 Dataset Analysis
  • 5 MAE Model
  • 5.1 Problem Formulation
  • 5.2 Question and Multimodal Knowledge Encoder
  • 5.3 Evidence Retrieval
  • 5.4 Adaptive Answer Extraction
  • 6 Experiments
  • 6.1 Evaluation Protocols
  • 6.2 Baseline Models
  • 6.3 Results Analysis
  • 7 Conclusion
  • Acknowledgments
  • References
  • A Annotation Guidelines
  • A.1 Decompose complex questions
  • A.2 Refine conversational questions
  • B Training Details
  • C Implementation Details

Knowls

  1. Knowl 1 — Multimodal conversational question answering task

    definition

    MMCoQA is the task of answering a sequence of user questions by retrieving evidence from a shared collection containing passages, tables, and images. For a question at turn kk, the system must use the preceding questions and answers, determine which modality or combination of modalities is useful, retrieve the relevant evidence, and produce a natural-language answer.

    The task introduces three interacting challenges. Priority requires selecting the most suitable modality when the question does not specify one. Consistency requires using agreement between modalities, such as an image and a passage describing the same entity, to verify an answer. Complementarity requires combining evidence from different modalities when no single item is sufficient. The required modality can also change from one turn to the next.

  2. Knowl 2 — MMConvQA construction from complex multimodal questions

    model/method

    MMConvQA is constructed by converting complex questions from the MMQA question pool into multi-turn conversations. The process first generates artificial candidate conversations: a seed question is selected, entities in its question and answer are identified, a subsequent entity is sampled, and another unused question containing that entity is selected as a follow-up. Questions are removed from the pool after selection, and turns are added until no matching question remains or a maximum conversation length is reached.

    Human annotators then transform the artificial conversations. For a complex question, the annotator uses its logical type and provided intermediate answers to decompose it into independent, self-contained questions. The supported composition patterns include Compose, which nests a question producing an entity inside another question; Intersect, which intersects the lists returned by two questions; and Compare, which compares two entities that can be linked to table cells. The annotator subsequently rewrites, reorders, or deletes questions to produce a coherent dialogue, and may discard an entire low-quality conversation. The final dataset retains both the natural conversational questions and their corresponding self-contained versions.

    Four computer-science students with NLP research experience performed annotation using training, annotation, checking, modification, and re-checking. Each conversation was checked by another annotator and returned for modification when it did not meet the quality requirements.

  3. Knowl 3 — MMConvQA dataset scale and multimodal composition

    data/table

    MMConvQA contains conversational questions, their natural-language answers, self-contained versions, and one supporting passage, table, or image for each question. Its scale and modality distribution are:

    Could not parse LaTeX table

    The dataset therefore requires retrieval over a substantially larger multimodal collection than the number of annotated questions. Text questions are the largest group, but table and image questions together account for 54.4% of all questions, making modality selection a central part of the task. Answers are usually passage spans, table cells, or image titles, although some require visual interpretation beyond directly copying the evidence; 9.9% of questions require a list of answers, and the average answer length is 2.11 words.

  4. Knowl 4 — MAE multimodal encoding and dense evidence retrieval

    model/method

    The baseline Multimodal Conversational QA system with Adaptive Extractors (MAE) represents a conversation-aware question and every item in a multimodal knowledge collection in a shared retrieval space. Let kk be the current turn, let qkq_k be its question, and let Hk={q1,a1,…,qk−1,ak−1}H_k=\{q_1,a_1,\ldots,q_{k-1},a_{k-1}\} be the preceding questions and answers. Let C=Cp∪Ct∪CiC=C_p\cup C_t\cup C_i contain passages CpC_p, tables CtC_t, and images CiC_i. A sliding-window context filter produces a reformulated question qk′q'_k from qkq_k and HkH_k.

    The question representation and item representations are computed as

    vq=WqFq(qk′),vpj=WpFp(pj),vtj=WtFt(tj′),vij=WiFi(ij).v_q=W_qF_q(q'_k),\qquad v_p^j=W_pF_p(p_j),\qquad v_t^j=W_tF_t(t'_j),\qquad v_i^j=W_iF_i(i_j).

    Here, pj∈Cpp_j\in C_p is a passage, tj′t'_j is table jj linearized row by row, and ij∈Cii_j\in C_i is an image. FqF_q, FpF_p, and FtF_t are BERT-based encoders; FiF_i is an ImageNet-pretrained ResNet; and WqW_q, WpW_p, WtW_t, and WiW_i are learned projection matrices. The vectors vqv_q, vpjv_p^j, vtjv_t^j, and vijv_i^j all lie in the same dd-dimensional space.

    The knowledge encoders do not depend on the question. They are run offline over all items, after which the encoded collection is frozen. For a question vector vqv_q and an item vector vxv_x, MAE uses the inner product sa(vq,vx)=vqTvxs_a(v_q,v_x)=v_q^{\mathsf T}v_x as the retrieval score and retains the top NrN_r items as the evidence list. In the reported implementation, d=128d=128 and Nr=10N_r=10.

  5. Knowl 5 — Adaptive modality-specific answer extraction

    model/method

    MAE predicts answers from the retrieved evidence using separate extractors for passages, tables, and images. A modality classifier receives the reformulated question and predicts a three-way probability vector

    sb=softmax⁡ ⁣(WcFc(qk′))∈R3,s_b=\operatorname{softmax}\!\left(W_cF_c(q'_k)\right)\in\mathbb{R}^{3},

    where FcF_c is a question encoder, WcW_c is a classifier matrix, and the three coordinates correspond to text, table, and image modalities.

    The text extractor applies span prediction to each retrieved passage. The table extractor concatenates the reformulated question with a row-linearized table and predicts start and end tokens in the resulting sequence. The image extractor obtains a ResNet feature for the image, concatenates the question with every answer string observed during training, and combines the resulting BERT token representations with the visual feature before predicting answer start and end positions.

    For a candidate answer, the extraction score is the mean of its predicted start-token probability psp_s and end-token probability pep_e:

    sc=ps+pe2.s_c=\frac{p_s+p_e}{2}.

    MAE ranks candidate answers using the sum of the retrieval score, modality score, and extraction score, sa+sb+scs_a+s_b+s_c, selecting the highest-scoring candidate.

  6. Knowl 6 — MAE training and evaluation configuration

    experimental setup

    MAE first pretrains the three knowledge encoders for 12 epochs with batch size 4 and Adam at learning rate 0.00010.0001. After all passages, tables, and images are encoded offline, the knowledge-encoder parameters are frozen. The question encoder and answer extractors are then trained with batch size 1 and Adam at learning rate 0.00010.0001. The checkpoint with the best development-set performance is evaluated on the test set using a server with an NVIDIA 3090 GPU.

    Evidence retrieval is evaluated using Recall and NDCG over the top 2,000 retrieved items. Answer extraction is evaluated using word-level macro-average F1 and Exact Match (EM). The experiments include ordinary evidence retrieval, in which the system must find the evidence, and an evidence-given condition, in which the supporting evidence is manually inserted into the retrieved list if it was not retrieved. Diagnostic variants remove conversation context, replace the reformulated question with its gold self-contained question, append previous gold answers, rewrite the question with a separate question-rewriting procedure, or pretrain retrieval on ORQuAC data.

  7. Knowl 7 — Benchmark performance on MMConvQA

    data/table

    The benchmark compares MAE with an open-retrieval conversational QA system, a multimodal QA system, and MAE diagnostic variants. ER means that supporting evidence must be retrieved; EG means that supporting evidence is manually included. Recall and NDCG use the top 2,000 retrieved items. The reported values are:

    Could not parse LaTeX table

    The comparison demonstrates that MAE has higher test retrieval Recall than ORConvQA, but answer F1 and EM remain low when retrieval is required. Supplying the correct evidence raises MAE test F1 from 2.19 to 28.33 and test EM from 1.36 to 22.03, showing that answer extraction is substantially more capable when retrieval errors are removed.

  8. Knowl 8 — Conversation context and retrieval are the principal MAE bottlenecks

    empirical result

    The MAE diagnostics show that conversational context is necessary: on the test set under evidence retrieval, removing context reduces Recall from 41.53 to 33.28, F1 from 2.19 to 1.74, and EM from 1.36 to 0.82. Providing the gold self-contained question instead of the model's reformulation increases the same metrics to Recall 63.39, F1 6.29, and EM 3.73, indicating that question understanding and reformulation remain difficult.

    The largest gap appears between retrieved-evidence and evidence-given conditions. With supporting evidence manually inserted, MAE reaches test F1 28.33 and EM 22.03, compared with 2.19 and 1.36 when evidence must be retrieved. The paper therefore identifies multimodal evidence retrieval as the dominant current bottleneck. The retrieval variants also show a benefit from additional supervision: ORQuAC pretraining raises test Recall from 41.53 to 42.71 and test F1 from 2.19 to 3.59, while question rewriting raises them to 46.32 and 4.91, respectively.

  9. Knowl 9 — MMConvQA exhibits dynamic modality and conversational phenomena

    empirical result

    The dataset analysis shows that modality requirements change frequently across a dialogue. Among the 1,179 conversations, 57.7% use two different knowledge modalities and 24.4% use all three modalities. In the first transition, approximately 70% of table questions become text or image questions at the next turn, illustrating that a single fixed modality is inadequate.

    The linguistic analysis of 100 sampled development-set follow-up questions found that approximately 33% can be answered without conversational history, 57% contain explicit coreference markers such as “he,” “she,” or “it,” and 10% refer implicitly to a prior entity or event without an explicit marker. About 24% change the conversation topic. Conversational questions average 14.4 words, while their self-contained gold versions average 15.5 words; the shorter conversational form reflects coreference and ellipsis. Answers average 2.11 words, and 9.9% of questions require answer lists.

  10. Knowl 10 — Limitations of the baseline multimodal representation

    limitation

    MAE performs worst on image questions in the evidence-given modality analysis. The paper attributes this partly to the coarse image extractor, which predicts answers by combining a global ResNet feature with text representations rather than modeling fine-grained visual-language interactions.

    The visualization of multimodal embeddings further shows that image representations form isolated clusters, while some text and table representations overlap only partially. Thus, items with related meanings are not organized in a well-aligned common space, limiting inner-product dense retrieval across modalities. The paper concludes that document-oriented dense retrieval must be substantially adapted to represent semantic relations among passages, tables, and images.

Coverage note — No substantial contributed component was omitted; auxiliary loss equations and related-work comparisons were left out because they are implementation details or background rather than standalone contributions.

References

  1. 1.Raviteja Anantha, Svitlana Vakulenko, Zhucheng Tu, Shayne Longpre, Stephen Pulman, and Srinivas Chappidi. 2021a. Open-domain question answering goes conversational via question rewriting. In Proceedings of the International Conference of the North American Chapter of the Association for Computational Linguistics, pages 520–534.
  2. 2.Raviteja Anantha, Svitlana Vakulenko, Zhucheng Tu, Shayne Longpre, Stephen Pulman, and Srinivas Chappidi. 2021b. Open-domain question answering goes conversational via question rewriting. In Proceedings of the International Conference of the North American Chapter of the Association for Computational Linguistics, pages 520–534.
  3. 3.Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wentau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018. Quac: Question answering in context. In Proceedings of the International Conference on Empirical Methods in Natural Language Processing, pages 2174–2184.
  4. 4.Philipp Christmann, Rishiraj Saha Roy, Abdalghani Abujabal, Jyotsna Singh, and Gerhard Weikum. 2019. Look before you hop: Conversational question answering over knowledge graphs using judicious context expansion. In Proceedings of the International Conference on Information and Knowledge Management, pages 729–738.
  5. 5.Jeffrey Dalton, Chenyan Xiong, Vaibhav Kumar, and Jamie Callan. 2020. Cast-19: A dataset for conversational information seeking. In Proceedings of the International Conference on Research and Development in Information Retrieval, pages 1985–1988.
  6. 6.Yang Deng, Yuexiang Xie, Yaliang Li, Min Yang, Nan Du, Wei Fan, Kai Lei, and Ying Shen. 2019. Multitask learning with multi-view attention for answer selection and knowledge base question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 6318–6325.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the International Conference of the North American Chapter of the Association for Computational Linguistics, pages 4171–4186.
  8. 8.Ahmed Elgohary, Denis Peskov, and Jordan Boyd-Graber. 2019. Can you unpack that? learning to rewrite questions-in-context. In Proceedings of the International Conference on Empirical Methods in Natural Language Processing, pages 5918–5924.
  9. 9.Hanyin Fang, Fei Wu, Zhou Zhao, Xinyu Duan, Yueting Zhuang, and Martin Ester. 2016. Community-based question answering via heterogeneous social network learning. In Proceedings of the AAAI Conference on Artificial Intelligence, 1.
  10. 10.Darryl Hannan, Akshay Jain, Mohit Bansal, li, and li. 2020. Manymodalqa: Modality disambiguation and qa over diverse inputs. Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):7879–7886.
  11. 11.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the International conference on computer vision and pattern recognition, pages 770–778.
  12. 12.Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Mueller, Francesco Piccinno, and Julian Eisenschlos. 2020. Tapas: Weakly supervised table parsing via pre-training. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 4320–4333.
  13. 13.Chenchen Jing, Yuwei Wu, Xiaoxun Zhang, Yunde Jia, and Qi Wu. 2020. Overcoming language priors in vqa via decomposed linguistic representations. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 11181–11188.
  14. 14.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the International Conference on Empirical Methods in Natural Language Processing, pages 6769–6781.
  15. 15.Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Jonghyun Choi, Ali Farhadi, and Hannaneh Hajishirzi. 2017. Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension. In Proceedings of the IEEE Conference on Computer Vision and Pattern recognition, pages 4999–5007.
  16. 16.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 6086–6096.
  17. 17.Jie Lei, Licheng Yu, Mohit Bansal, and Tamara Berg. 2018. Tvqa: Localized, compositional video question answering. In Proceedings of the International Conference on Empirical Methods in Natural Language Processing, pages 1369–1379.
  18. 18.Yongqi Li, Wenjie Li, and Liqiang Nie. 2022. Dynamic graph reasoning for conversational open-domain question answering. ACM Transactions on Information Systems, 40(4):1–24.
  19. 19.Liqiang Nie, Meng Wang, Zhengjun Zha, Guangda Li, and Tat-Seng Chua. 2011. Multimedia answering: enriching text qa with media information. In Proceedings of the International Conference on Research and Development in Information Retrieval, pages 695–704.
  20. 20.Chen Qu, Liu Yang, Cen Chen, Minghui Qiu, W Bruce Croft, and Mohit Iyyer. 2020. Open-retrieval conversational question answering. In Proceedings of the International conference on research and development in Information Retrieval, pages 539–548.
  21. 21.Siva Reddy, Danqi Chen, Christopher D. Manning, li, and li. 2019. Coqa: A conversational question answering challenge. Transactions of the Association for Computational Linguistics, 7:249–266.
  22. 22.Pengjie Ren, Zhumin Chen, Zhaochun Ren, Evangelos Kanoulas, Christof Monz, and Maarten de Rijke. 2021. Conversations with search engines: Serp-based conversational response generation.
  23. 23.Amrita Saha, Vardaan Pahuja, Mitesh Khapra, Karthik Sankaranarayanan, and Sarath Chandar. 2018a. Complex sequential question answering: Towards learning to converse over linked question answer pairs with a knowledge graph. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 705–713. AAAI.
  24. 24.Amrita Saha, Vardaan Pahuja, Mitesh M. Khapra, Karthik Sankaranarayanan, and Sarath Chandar. 2018b. Complex sequential question answering: Towards learning to converse over linked question answer pairs with a knowledge graph. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 705–713.
  25. 25.Amrita Saha, Vardaan Pahuja, Mitesh M. Khapra, Karthik Sankaranarayanan, and Sarath Chandar. 2018c. Complex sequential question answering: Towards learning to converse over linked question answer pairs with a knowledge graph. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 705–713.
  26. 26.Sanket Shah, Anand Mishra, Naganand Yadati, and Partha Pratim Talukdar. 2019. Kvqa: Knowledge-aware visual question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 8876–8884.
  27. 27.Tao Shen, Xiubo Geng, Tao Qin, Daya Guo, Duyu Tang, Nan Duan, Guodong Long, and Daxin Jiang. 2019. Multi-task learning for conversational question answering over a large-scale knowledge base. In Proceedings of the International Conference on Empirical Methods in Natural Language Processing, pages 2442–2451.
  28. 28.Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi, and Jonathan Berant. 2021. Multimodal qa: complex question answering over text, tables and images. In International Conference on Learning Representations.
  29. 29.Svitlana Vakulenko, Shayne Longpre, Zhucheng Tu, and Raviteja Anantha. 2021. Question rewriting for conversational question answering. In Proceedings of the International Conference on Web Search and Data Mining, pages 355–363.
  30. 30.Nikos Voskarides, Dan Li, Pengjie Ren, Evangelos Kanoulas, and Maarten de Rijke. 2020. Query resolution for conversational search with limited supervision. In Proceedings of the International Conference on Research and Development in Information Retrieval, pages 921–930.
  31. 31.Shuohang Wang, Mo Yu, Xiaoxiao Guo, Zhiguo Wang, Tim Klinger, Wei Zhang, Shiyu Chang, Gerry Tesauro, Bowen Zhou, and Jing Jiang. 2018. R 3: Reinforced ranker-reader for open-domain question answering. In Proceedings of the AAAI Conference on Artificial Intelligence.
  32. 32.Shi Yu, Jiahua Liu, Jingqin Yang, Chenyan Xiong, Paul Bennett, Jianfeng Gao, and Zhiyuan Liu. 2020. Few-shot generative conversational query rewriting. In Proceedings of the International conference on research and development in Information Retrieval, pages 1933–1936.

Citation

MLA
Li, Y., et al. “MMCoQA: Conversational Question Answering over Text, Tables, and Images”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 4220–31, https://doi.org/10.18653/v1/2022.acl-long.290.
APA
Li, Y., Li, W., & Nie, L. (2022). MMCoQA: Conversational Question Answering over Text, Tables, and Images. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4220–4231. https://doi.org/10.18653/v1/2022.acl-long.290
Chicago
Li, Y., W. Li, and L. Nie. 2022. “MMCoQA: Conversational Question Answering over Text, Tables, and Images”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4220–31. https://doi.org/10.18653/v1/2022.acl-long.290.
Harvard
Li, Y., Li, W. and Nie, L. (2022) “MMCoQA: Conversational Question Answering over Text, Tables, and Images”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 4220–4231. Available at: https://doi.org/10.18653/v1/2022.acl-long.290.
Vancouver
1. Li Y, Li W, Nie L (2022) MMCoQA: Conversational Question Answering over Text, Tables, and Images. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 4220–4231

BibTeX

@inproceedings{li-etal-2022-mmcoqa,
    title = "{MMC}o{QA}: Conversational Question Answering over Text, Tables, and Images",
    author = "Li, Yongqi  and
      Li, Wenjie  and
      Nie, Liqiang",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.290/",
    doi = "10.18653/v1/2022.acl-long.290",
    pages = "4220--4231"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/