T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Large Language Model Signals for Science Question Answering

Lei WangYi HuJiabang HeXing XuNing LiuHui LiuHeng Tao Shen

article2024AAAI107 citations

Proposes a data-generation and mixing framework that uses large language model reasoning signals to teach smaller multimodal models how to solve complex scientific questions, achieving a new state of the art on ScienceQA.

Listen

Complex scientific question answering requires artificial intelligence systems to interpret multimodal information—such as text, diagrams, and maps—while performing multi-step reasoning. Traditional approaches rely on training smaller models using human-annotated step-by-step explanations (known as chain-of-thought rationales). However, manual annotation is labor-intensive, expensive, and frequently lacks the broader external knowledge required to solve complex problems accurately.

The article demonstrates a new training approach called T-SciQ, which uses large language models to generate high-quality reasoning explanations and systematically teach compact student models to solve multimodal science questions.

The researchers developed a three-stage framework using an advanced teacher model to generate two distinct types of explanations: standard step-by-step reasoning for straightforward questions and plan-based reasoning that breaks complex problems into simpler subtasks. The framework applies a data-mixing strategy that uses a validation set to select the optimal explanation type for each skill category. These combined signals are then used to train smaller student models (under 1 billion parameters, over 200 times smaller than the teacher model) across a standard multimodal benchmark consisting of 21,208 science questions, as well as six additional language reasoning benchmarks.

The analysis produced several key findings. First, the primary student model achieved a new state-of-the-art accuracy of 96.18% on the multimodal benchmark, outperforming human performance (88.40%), the leading multimodal baseline (91.68%), and large few-shot models like GPT-4 (82.69%). Second, student models consistently outperformed baselines trained on human-annotated explanations across different model sizes and architectures, yielding absolute gains of 4.5% to 6.84%. Third, combining standard and plan-based explanations delivered superior accuracy compared to using either explanation style in isolation. Finally, the framework generalized effectively across six diverse text reasoning benchmarks, substantially improving performance in arithmetic, commonsense, and logic tasks.

These results demonstrate that compact, cost-efficient artificial intelligence models can surpass human benchmarks and massive commercial systems when trained on structured, model-generated explanations. Organizations can drastically lower operational deployment costs and latency by replacing human annotation pipelines with synthetic teaching data while achieving superior multi-step reasoning and open-world knowledge integration.

Decision-makers should consider adopting synthetic reasoning generation and dynamic data-mixing strategies when deploying smaller, task-specific models. For immediate next steps, technical teams should explore parameter-efficient fine-tuning techniques and evaluate different foundation models as teachers to optimize training costs.

The primary limitation noted in the article is the current reliance on full fine-tuning across specific model architectures and proprietary teacher model interfaces. However, given the consistent outperformance across diverse question categories, model architectures, and task benchmarks, there is high confidence in the robustness and practical effectiveness of the proposed approach.

Cover for T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Large Language Model Signals for Science Question Answering

Abstract

Large Language Models (LLMs) have recently demonstrated exceptional performance in various Natural Language Processing (NLP) tasks. They have also shown the ability to perform chain-of-thought (CoT) reasoning to solve complex problems. Recent studies have explored CoT reasoning in complex multimodal scenarios, such as the science question answering task, by fine-tuning multimodal models with high-quality human-annotated CoT rationales. However, collecting high-quality CoT rationales is usually time-consuming and costly. Besides, the annotated rationales are hardly accurate due to the external essential information missed. To address these issues, we propose a novel method termed T-SciQ that aims at teaching science question answering with LLM signals. The T-SciQ approach generates high-quality CoT rationales as teaching signals and is advanced to train much smaller models to perform CoT reasoning in complex modalities. Additionally, we introduce a novel data mixing strategy to produce more effective teaching data samples for simple and complex science question answer problems. Extensive experimental results show that our T-SciQ method achieves a new state-of-the-art performance on the ScienceQA benchmark, with an accuracy of 96.18%. Moreover, our approach outperforms the most powerful fine-tuned baseline by 4.5%. The code is publicly available at https://github.com/T-SciQ/T-SciQ.

Table of Contents

  • Introduction
  • Related Work
  • Our T-SciQ Approach
  • Overview
  • Generating Teaching Data
  • Mixing Teaching Data
  • Fine-Tuning
  • Experiment
  • Experimental Setup
  • Main Results
  • Further Analysis
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — T-SciQ Framework Overview

    model/method

    T-SciQ (Teaching Science Question answering with LLM signals) is a framework designed to impart chain-of-thought (CoT) reasoning capabilities to compact multimodal student models (with parameter counts below 1B1\text{B}, such as UnifiedQA and Multimodal-CoT) using synthetic rationales generated by large language models (LLMs). The framework addresses two fundamental limitations of human-annotated CoT rationales: the labor-intensive annotation bottleneck and the frequent absence of essential open-world factual knowledge in human annotations.

    T-SciQ operates in three main stages:

    1. Teaching Data Generation: An LLM teacher (e.g., GPT-3.5 text-davinci-003) synthesizes two distinct forms of reasoning signals: standard chain-of-thought (QA-CoT) for direct problem solving, and hierarchical plan-based chain-of-thought (QA-PCoT) for multi-step problem decomposition.
    2. Data Mixing: A skill-level validation selection procedure evaluates whether QA-CoT or QA-PCoT yields higher answer accuracy for each specific question skill, constructing an optimal mixed training set.
    3. Two-Stage Student Fine-Tuning: The multimodal student model is trained in two sequential stages: rationale generation teaching (learning to generate the reasoning chain from multimodal inputs) followed by answer inference teaching (learning to predict the final answer given multimodal inputs augmented with the generated rationale).
  2. Knowl 2 — QA-CoT and Plan-Based QA-PCoT Prompting Mechanisms

    model/method

    T-SciQ generates two distinct styles of synthetic reasoning rationales using large language models:

    1. Standard CoT Generation (QA-CoT): For a question instance with question text XqX_q, context XcX_c (or N/A if absent), options XoX_o, and ground-truth answer label AA used as a guidance hint, the prompt is structured as: Question: [Xq]. Context: [Xc]. Options: [Xo]. Correct Answer: [A]. Please give me a detailed explanation. The LLM outputs an explanation incorporating open-world knowledge to produce teaching signal Ti,cotT_{i,\text{cot}}.

    2. Plan-Based CoT Generation (QA-PCoT): To handle complex multi-step reasoning, a 3-step zero-shot prompting pipeline decomposes problems across each question skill SS:

    • Step 1 (Lecture Generation): Given QA pairs belonging to skill SS, the LLM is prompted: Skill: [S]. QA pairs: [Xq, A] ... based on the problems above, please give a general lecture on the [S] type of question in one sentence. This produces domain lecture LL.
    • Step 2 (Plan Generation): The LLM receives the skill, lecture, and QA pairs: Skill: [S]. Lecture: [L]. QA pairs: [Xq, A] ... Based on the lecture above and these problems, let's understand these problems and devise a general and brief plan step by step to solve these problems (begin with 1, 2, 3...). This yields problem-solving plan PP.
    • Step 3 (Rationale Generation): For each instance, the LLM carries out the plan: Skill: [S]. Lecture: [L]. Plan: [P]. QA pair: [Xq, A]. Based on the lecture, the plan and the problem, please carry out the plan and solve the problem step by step (begin with 1, 2, 3...). This produces plan-based teaching signal Ti,pcotT_{i,\text{pcot}}.
  3. Knowl 3 — Skill-Level Validation Data Mixing Algorithm

    algorithm

    To combine the advantages of QA-CoT on direct problems and QA-PCoT on complex multi-step problems, T-SciQ evaluates rationale efficacy per skill on the validation dataset before selecting the training signal for that skill.

    Input: Validation set DvalD_{val} partitioned into skills SS, candidate rationales TcotT_{\text{cot}} and TpcotT_{\text{pcot}}, answer generation model FasF_a^s trained on human-annotated seed data, training set DtrainD_{train}.
    Output: Mixed training dataset DT-SciQD_{T\text{-}SciQ}.
    Initialize DT-SciQ←∅D_{T\text{-}SciQ} \leftarrow \emptyset
    for each skill s∈Ss \in S do
        Ecot(s)←0E_{\text{cot}}(s) \leftarrow 0
        Epcot(s)←0E_{\text{pcot}}(s) \leftarrow 0
        for each validation instance (Xi,la,Xi,v,Ai)∈Dval(X_{i,la}, X_{i,v}, A_i) \in D_{val} belonging to skill ss do
            A^i,cot←Fas(Xi,la,Xi,v,Ti,cot)\hat{A}_{i,\text{cot}} \leftarrow F_a^s(X_{i,la}, X_{i,v}, T_{i,\text{cot}})
            if A^i,cot≠Ai\hat{A}_{i,\text{cot}} \neq A_i then
                Ecot(s)←Ecot(s)+1E_{\text{cot}}(s) \leftarrow E_{\text{cot}}(s) + 1
            end if
            A^i,pcot←Fas(Xi,la,Xi,v,Ti,pcot)\hat{A}_{i,\text{pcot}} \leftarrow F_a^s(X_{i,la}, X_{i,v}, T_{i,\text{pcot}})
            if A^i,pcot≠Ai\hat{A}_{i,\text{pcot}} \neq A_i then
                Epcot(s)←Epcot(s)+1E_{\text{pcot}}(s) \leftarrow E_{\text{pcot}}(s) + 1
            end if
        end for
        if Epcot(s)<Ecot(s)E_{\text{pcot}}(s) < E_{\text{cot}}(s) then
            T∗(s)←PCoTT^*(s) \leftarrow \text{PCoT}
        else
            T∗(s)←CoTT^*(s) \leftarrow \text{CoT}
        end if
        for each training instance i∈Dtraini \in D_{train} belonging to skill ss do
            if T∗(s)==PCoTT^*(s) == \text{PCoT} then
                DT-SciQ←DT-SciQ∪{(Xi,la,Xi,v,Ti,pcot,Ai)}D_{T\text{-}SciQ} \leftarrow D_{T\text{-}SciQ} \cup \{(X_{i,la}, X_{i,v}, T_{i,\text{pcot}}, A_i)\}
            else
                DT-SciQ←DT-SciQ∪{(Xi,la,Xi,v,Ti,cot,Ai)}D_{T\text{-}SciQ} \leftarrow D_{T\text{-}SciQ} \cup \{(X_{i,la}, X_{i,v}, T_{i,\text{cot}}, A_i)\}
            end if
        end for
    end for
    return DT-SciQD_{T\text{-}SciQ}
  4. Knowl 4 — Two-Stage Multimodal Student Training Objectives and Architecture

    model/method

    T-SciQ fine-tunes a multimodal student model via two decoupled training stages using the mixed dataset DT-SciQD_{T\text{-}SciQ}:

    1. Rationale Generation Teaching: Given initial text input Xi,la1X_{i,la}^1 (containing question text, context, and options) and visual feature representation Xi,vX_{i,v}, the rationale generation model FrF_r parameterized by θr\theta_r autoregressively generates the assigned teaching rationale sequence Ti=(Ti,1,…,Ti,NTi)T_i = (T_{i,1}, \dots, T_{i,N_{T_i}}) of length NTiN_{T_i}: p(Ti∣Xi,la1,Xi,v)=∏j=1NTipθr(Ti,j∣Xi,la1,Xi,v,Ti,<j)p(T_i \mid X_{i,la}^1, X_{i,v}) = \prod_{j=1}^{N_{T_i}} p_{\theta_r}(T_{i,j} \mid X_{i,la}^1, X_{i,v}, T_{i,<j})

    2. Answer Inference Teaching: An augmented language input Xi,la2=[Xi,la1;Ti]X_{i,la}^2 = [X_{i,la}^1; T_i] is constructed by concatenating the teaching rationale to the initial text input. The answer inference model FaF_a parameterized by θa\theta_a autoregressively generates the target answer label sequence Ai=(Ai,1,…,Ai,NAi)A_i = (A_{i,1}, \dots, A_{i,N_{A_i}}) of length NAiN_{A_i}: p(Ai∣Xi,la2,Xi,v)=∏j=1NAipθa(Ai,j∣Xi,la2,Xi,v,Ai,<j)p(A_i \mid X_{i,la}^2, X_{i,v}) = \prod_{j=1}^{N_{A_i}} p_{\theta_a}(A_{i,j} \mid X_{i,la}^2, X_{i,v}, A_{i,<j})

    Model Architecture: The backbone employs a language Transformer encoder and a visual Transformer (e.g., DETR or CLIP) to extract multimodal representations, which are merged via a gated fusion mechanism before decoding with a Transformer decoder. The rationale generation module FrF_r and answer inference module FaF_a share identical underlying model architecture parameters but are trained on distinct input-output sequences. At inference time, rationales predicted by FrF_r are concatenated with test inputs and fed into FaF_a to predict final answers.

  5. Knowl 5 — Main ScienceQA Benchmark Results

    empirical result

    The T-SciQ method was evaluated on the multimodal multiple-choice ScienceQA benchmark (21,208 questions partitioned into 12,726 training, 4,241 validation, and 4,241 test instances across natural science NAT, social science SOC, language science LAN, text context TXT, image context IMG, no context NO, grades 1–6 G1-6, and grades 7–12 G7-12).

    Could not parse LaTeX table

    Multimodal-T-SciQLarge_{\text{Large}} (738M parameters) achieves a new state-of-the-art accuracy of 96.18%, outperforming the human benchmark (88.40%), the strongest fine-tuned baseline Multimodal-CoTLarge_{\text{Large}} (91.68% by +4.50%), the instruction-tuned LLaVA (90.92% by +5.26%), and the few-shot LLM framework Chameleon (86.54% by +9.64%).

  6. Knowl 6 — Ablation Study on LLM Teaching Signals and Data Mixing

    empirical result

    An ablation study on ScienceQA compares the contribution of individual LLM reasoning signals (QA-CoT and QA-PCoT) against the mixed T-SciQ dataset across both Base (223M) and Large (738M) student model scales.

    Could not parse LaTeX table

    Both individual LLM-generated signals outperform training exclusively on human annotations (85.99% and 88.56% vs 84.91% for Base; 93.44% and 94.11% vs 91.68% for Large). Furthermore, combining QA-CoT and QA-PCoT via skill-level data mixing achieves the highest overall accuracy (91.75% Base, 96.18% Large), demonstrating the complementary nature of direct reasoning chains and structured problem decomposition.

  7. Knowl 7 — Ablation on Visual Feature Extractors

    empirical result

    The performance of Multimodal-T-SciQBase_{\text{Base}} on the ScienceQA test set was evaluated across different visual feature backbones and teaching signals.

    Could not parse LaTeX table

    Incorporating visual features produces substantial improvements over language-only modeling (e.g., from 87.24% to 91.75% for mixed T-SciQ). DETR achieves the strongest performance among visual extractors (91.75% for mixed T-SciQ vs 90.90% with CLIP and 91.44% with ResNet-50), and is therefore adopted as the default visual feature extractor.

  8. Knowl 8 — Evaluation on General NLP Reasoning Benchmarks

    empirical result

    To assess generalizability beyond multimodal science questions, the T-SciQ teaching signal pipeline was evaluated on six diverse text reasoning datasets covering arithmetic (AQuA), symbolic reasoning (Coin Flip), commonsense reasoning (CommonSenseQA, StrategyQA), and logical reasoning (Date Understanding, Tracking Shuffled Objects), and compared against Reason-Teacher (Ho et al., 2022).

    Could not parse LaTeX table

    T-SciQ outperforms Reason-Teacher across 5 of the 6 datasets by large margins (+50.78% on AQuA, +28.93% on Date Understanding, +5.84% on Shuffled Objects, +14.00% on CommonSenseQA, and +21.72% on StrategyQA) while matching performance on Coin Flip (98.67%).

  9. Knowl 9 — Impact of Base LLM Teachers and Data Proportion on Student Performance

    empirical result

    Ablations examining teacher model choice, training data composition, and learning dynamics on Multimodal-T-SciQBase_{\text{Base}} show:

    1. Teacher LLM Combinations: Testing all 9 pairwise combinations of teacher models (text-davinci-002, text-davinci-003, and ChatGPT) for generating QA-CoT and QA-PCoT signals revealed that all 9 configurations outperformed the human-annotated baseline (84.91%), with accuracies ranging from 87.83% to 91.75%. The combination of text-davinci-003 for QA-CoT and text-davinci-003 for QA-PCoT yielded the peak accuracy of 91.75%, while ChatGPT for QA-CoT paired with text-davinci-003 for QA-PCoT achieved 91.25%.
    2. Proportion of Synthetic Teaching Data: Systematically varying the proportion of LLM-generated T-SciQ training data from 0% to 100% (replacing human-annotated signals) demonstrated a strictly monotonic increase in student model accuracy as the fraction of synthetic data grew.
    3. Training Dynamics: Comparing accuracy across training epochs between Multimodal-CoTBase_{\text{Base}} and Multimodal-T-SciQBase_{\text{Base}} demonstrated that T-SciQ achieved higher accuracy starting from the very first epoch and maintained this margin throughout training.

Coverage note — None was omitted; all primary methods, algorithms, experimental formulations, benchmark comparisons, ablations, and analytical findings from the paper have been extracted as self-contained knowls.

References

  1. 1.Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, 6077–6086.
  2. 2.Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877– 1901.
  3. 3.Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020. End-to-End Object Detection with Transformers. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I, 213–229.
  4. 4.Chen, T.; Kornblith, S.; Swersky, K.; Norouzi, M.; and Hinton, G. E. 2020. Big self-supervised models are strong semi-supervised learners. Advances in neural information processing systems, 33: 22243–22255.
  5. 5.Chen, W.; Ma, X.; Wang, X.; and Cohen, W. W. 2022. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. arXiv preprint arXiv:2211.12588.
  6. 6.Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  7. 7.Dalvi, B.; Jansen, P.; Tafjord, O.; Xie, Z.; Smith, H.; Pipatanangkura, L.; and Clark, P. 2021. Explaining answers with entailment trees. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP).
  8. 8.Fu, Y.; Peng, H.; Ou, L.; Sabharwal, A.; and Khot, T. 2023. Specializing Smaller Language Models towards Multi-Step Reasoning. arXiv preprint arXiv:2301.12726.
  9. 9.Fu, Y.; Peng, H.; Sabharwal, A.; Clark, P.; and Khot, T. 2022. Complexity-based prompting for multi-step reasoning. arXiv preprint arXiv:2210.00720.
  10. 10.Gao, P.; Jiang, Z.; You, H.; Lu, P.; Hoi, S. C.; Wang, X.; and Li, H. 2019. Dynamic fusion with intra-and inter-modality attention flow for visual question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 6639–6648.
  11. 11.Geva, M.; Khashabi, D.; Segal, E.; Khot, T.; Roth, D.; and Berant, J. 2021. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9: 346–361.
  12. 12.He, J.; Wang, L.; Hu, Y.; Liu, N.; Liu, H.; Xu, X.; and Shen, H. T. 2023. ICL-D3IE: In-Context Learning with Diverse Demonstrations Updating for Document Information Extraction. In IEEE/CVF International Conference on Computer Vision, 19428–19437.
  13. 13.He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, 770–778. IEEE Computer Society.
  14. 14.Ho, N.; Schmid, L.; and Yun, S.-Y. 2022. Large Language Models Are Reasoning Teachers. arXiv preprint arXiv:2212.10071.
  15. 15.Hsieh, C.-Y.; Li, C.-L.; Yeh, C.-K.; Nakhost, H.; Fujii, Y.; Ratner, A.; Krishna, R.; Lee, C.-Y.; and Pfister, T. 2023. Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes. arXiv preprint arXiv:2305.02301.
  16. 16.Hu, Z.; Lan, Y.; Wang, L.; Xu, W.; Lim, E.-P.; Lee, R. K.-W.; Bing, L.; and Poria, S. 2023. LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models. arXiv preprint arXiv:2304.01933.
  17. 17.Huang, J.; Gu, S. S.; Hou, L.; Wu, Y.; Wang, X.; Yu, H.; and Han, J. 2022. Large language models can self-improve. arXiv preprint arXiv:2210.11610.
  18. 18.Jansen, P. A.; Wainwright, E.; Marmorstein, S.; and Morrison, C. T. 2018. Worldtree: A corpus of explanation graphs for elementary science questions supporting multi-hop inference. arXiv preprint arXiv:1802.03052.
  19. 19.Kembhavi, A.; Seo, M.; Schwenk, D.; Choi, J.; Farhadi, A.; and Hajishirzi, H. 2017. Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4999–5007.
  20. 20.Khashabi, D.; Min, S.; Khot, T.; Sabharwal, A.; Tafjord, O.; Clark, P.; and Hajishirzi, H. 2020. Unifiedqa: Crossing format boundaries with a single qa system. arXiv preprint arXiv:2005.00700.
  21. 21.Khot, T.; Trivedi, H.; Finlayson, M.; Fu, Y.; Richardson, K.; Clark, P.; and Sabharwal, A. 2022. Decomposed prompting: A modular approach for solving complex tasks. arXiv preprint arXiv:2210.02406.
  22. 22.Kim, J.-H.; Jun, J.; and Zhang, B.-T. 2018. Bilinear attention networks. Advances in neural information processing systems, 31.
  23. 23.Kim, W.; Son, B.; and Kim, I. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning, 5583–5594. PMLR.
  24. 24.Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2022. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916.
  25. 25.Li, B.; Lv, C.; Zhou, Z.; Zhou, T.; Xiao, T.; Ma, A.; and Zhu, J. 2022a. On Vision Features in Multimodal Machine Translation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6327–6337.
  26. 26.Li, L. H.; Yatskar, M.; Yin, D.; Hsieh, C.-J.; and Chang, K.-W. 2019. Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557.
  27. 27.Li, Y.; Lin, Z.; Zhang, S.; Fu, Q.; Chen, B.; Lou, J.-G.; and Chen, W. 2022b. On the advance of making language models better reasoners. arXiv preprint arXiv:2206.02336.
  28. 28.Ling, W.; Yogatama, D.; Dyer, C.; and Blunsom, P. 2017. Program induction by rationale generation: Learning to solve and explain algebraic word problems. arXiv preprint arXiv:1705.04146.
  29. 29.Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023. Visual instruction tuning. arXiv preprint arXiv:2304.08485.
  30. 30.Lu, P.; Mishra, S.; Xia, T.; Qiu, L.; Chang, K.-W.; Zhu, S.-C.; Tafjord, O.; Clark, P.; and Kalyan, A. 2022a. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35: 2507–2521.
  31. 31.Lu, P.; Peng, B.; Cheng, H.; Galley, M.; Chang, K.-W.; Wu, Y. N.; Zhu, S.-C.; and Gao, J. 2023. Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models. arXiv preprint arXiv:2304.09842.
  32. 32.Lu, P.; Qiu, L.; Chang, K.-W.; Wu, Y. N.; Zhu, S.-C.; Rajpurohit, T.; Clark, P.; and Kalyan, A. 2022b. Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning. arXiv preprint arXiv:2209.14610.
  33. 33.Lu, P.; Qiu, L.; Chen, J.; Xia, T.; Zhao, Y.; Zhang, W.; Yu, Z.; Liang, X.; and Zhu, S.-C. 2021. Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning. arXiv preprint arXiv:2110.13214.
  34. 34.Magister, L. C.; Mallinson, J.; Adamek, J.; Malmi, E.; and Severyn, A. 2022. Teaching small language models to reason. arXiv preprint arXiv:2212.08410.
  35. 35.Nye, M.; Andreassen, A. J.; Gur-Ari, G.; Michalewski, H.; Austin, J.; Bieber, D.; Dohan, D.; Lewkowycz, A.; Bosma, M.; Luan, D.; et al. 2021. Show your work: Scratchpads for intermediate computation with language models. arXiv preprint arXiv:2112.00114.
  36. 36.OpenAI. 2022. Introducing chatgpt. https://openai.com/blog/chatgpt. Accessed: 2022-11-30.
  37. 37.OpenAI. 2023. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774.
  38. 38.Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, 8748–8763. PMLR.
  39. 39.Rubin, O.; Herzig, J.; and Berant, J. 2021. Learning to retrieve prompts for in-context learning. arXiv preprint arXiv:2112.08633.
  40. 40.Sampat, S. K.; Yang, Y.; and Baral, C. 2020. Visuo-Lingustic Question Answering (VLQA) Challenge. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings (EMNLP), 4606–4616.
  41. 41.Talmor, A.; Herzig, J.; Lourie, N.; and Berant, J. 2018. Commonsenseqa: A question answering challenge targeting commonsense knowledge. arXiv preprint arXiv:1811.00937.
  42. 42.Thoppilan, R.; De Freitas, D.; Hall, J.; Shazeer, N.; Kulshreshtha, A.; Cheng, H.-T.; Jin, A.; Bos, T.; Baker, L.; Du, Y.; et al. 2022. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239.
  43. 43.Tian, Q.; Zhu, H.; Wang, L.; Li, Y.; and Lan, Y. 2023. R3 Prompting: Review, Rephrase and Resolve for Chain-of-Thought Reasoning in Large Language Models under Noisy Context. arXiv preprint arXiv:2310.16535.
  44. 44.Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Roziere, B.; Goyal, N.; Hambro, E.; ` Azhar, F.; et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  45. 45.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems 30, 5998–6008.
  46. 46.Wang, L.; Xu, W.; Lan, Y.; Hu, Z.; Lan, Y.; Lee, R. K.-W.; and Lim, E.-P. 2023. Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. arXiv preprint arXiv:2305.04091.
  47. 47.Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; and Zhou, D. 2022a. Rationale-augmented ensembles in language models. arXiv preprint arXiv:2207.00747.
  48. 48.Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; and Zhou, D. 2022b. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  49. 49.Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Chi, E.; Le, Q.; and Zhou, D. 2022a. Chain of Thought Prompting Elicits Reasoning in Large Language Models. ArXiv preprint, abs/2201.11903.
  50. 50.Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Chi, E.; Le, Q.; and Zhou, D. 2022b. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  51. 51.Yu, Z.; Yu, J.; Cui, Y.; Tao, D.; and Tian, Q. 2019. Deep modular co-attention networks for visual question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 6281–6290.
  52. 52.Zhang, R.; Han, J.; Zhou, A.; Hu, X.; Yan, S.; Lu, P.; Li, H.; Gao, P.; and Qiao, Y. 2023a. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199.
  53. 53.Zhang, Z.; Zhang, A.; Li, M.; and Smola, A. 2022. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493.
  54. 54.Zhang, Z.; Zhang, A.; Li, M.; Zhao, H.; Karypis, G.; and Smola, A. 2023b. Multimodal chain-of-thought reasoning in language models. arXiv preprint arXiv:2302.00923.
  55. 55.Zhou, D.; Schärli, N.; Hou, L.; Wei, J.; Scales, N.; Wang, X.; Schuurmans, D.; Bousquet, O.; Le, Q.; and Chi, E. 2022. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625.

Citation

MLA
Wang, L., et al. “T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Mixed Large Language Model Signals for Science Question Answering”. arXiv, 2023, http://arxiv.org/abs/2305.03453v4.
APA
Wang, L., Hu, Y., He, J., Xu, X., Liu, N., Liu, H., & Shen, H. T. (2023). T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Mixed Large Language Model Signals for Science Question Answering. arXiv. http://arxiv.org/abs/2305.03453v4
Chicago
Wang, L., Y. Hu, J. He, et al. 2023. “T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Mixed Large Language Model Signals for Science Question Answering”. arXiv. http://arxiv.org/abs/2305.03453v4.
Harvard
Wang, L. et al. (2023) “T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Mixed Large Language Model Signals for Science Question Answering”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.03453v4.
Vancouver
1. Wang L, Hu Y, He J, Xu X, Liu N, Liu H, Shen HT (2023) T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Mixed Large Language Model Signals for Science Question Answering. arXiv

BibTeX

@article{wang2023sciq,
  title = {T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Mixed Large Language Model Signals for Science Question Answering},
  author = {Wang, Lei and Hu, Yi and He, Jiabang and Xu, Xing and Liu, Ning and Liu, Hui and Shen, Heng Tao},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.03453v4},
  eprint = {2305.03453}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF