MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering

Fangyu LiuFrancesco PiccinnoSyrine KricheneChenxi PangKenton LeeMandar JoshiYasemin AltunNigel CollierJulian Martin Eisenschlos

article2023ACL193 citations

Proposes a pretraining framework combining chart derendering and math reasoning tasks that improves visual language models, boosting performance on ChartQA and PlotQA by up to 20% while generalizing effectively to diverse visual document tasks.

Listen

Visual language data, such as charts, plots, diagrams, and infographics, are essential for communicating complex information across documents, reports, and websites. Despite advances in vision-language artificial intelligence, existing models struggle to interpret these visual assets because standard natural-image training fails to capture layout organization, numerical extraction, and mathematical reasoning. Addressing this limitation is critical for automating data analysis and building reliable multimodal systems.

The article demonstrates a pretraining method called MATCHA (Math reasoning and Chart derendering pretraining) designed to improve an image-to-text model's ability to process visual language. The approach continually trains an existing base vision model, Pix2Struct, using a balanced mixture of tasks: 40% chart derendering (converting visual plots back into underlying data tables or rendering code), 40% mathematical reasoning (solving text-based arithmetic and comparison problems converted into images), and 20% webpage screenshot parsing.

The findings show that MATCHA establishes new state-of-the-art performance across several standard benchmarks. On chart question answering benchmarks, MATCHA achieves an overall accuracy of 64.2% on ChartQA and 91.5% on PlotQA, outperforming previous top models that lack access to underlying source tables by 8.2% and up to 19% respectively. In fact, MATCHA surpasses baseline systems that were provided with the actual underlying data tables. The model also sets a new benchmark on chart summarization and transfers effectively beyond plots, improving average performance across other document and interface understanding tasks by 2.3%.

These results demonstrate that visual layout deconstruction and mathematical reasoning are the core capabilities needed for visual language understanding. By training a relatively compact 300-million-parameter model on these targeted tasks, MATCHA significantly outperforms substantially larger general-purpose vision models (such as the 17-billion-parameter PaLI) on chart reasoning while reducing computational demands and eliminating reliance on brittle optical character recognition pipelines.

For practical implementation, organizations should adopt chart derendering and numerical reasoning objectives when building automated chart-processing and document-intelligence workflows. Future work should explore hybrid approaches—such as pairing chart-to-table translation modules with external calculators or symbolic program executors—to solve high-precision arithmetic where pure neural computation still makes errors.

Key limitations include difficulty with highly complex calculations requiring exact numerical precision, weaker performance on visual plot attributes (such as identifying colors or shapes) compared to massive web-scale vision models, and reliance on single-run evaluations due to computational constraints. Readers should maintain moderate confidence in the reported metrics while exercising caution in deployment scenarios that require strict arithmetic precision.

arXiv: 2212.09662
Cover for MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering

Abstract

Visual language data such as plots, charts, and infographics are ubiquitous in the human world. However, state-of-the-art vision-language models do not perform well on these data. We propose MATCHA (Math reasoning and Chart derendering pretraining) to enhance visual language models' capabilities in jointly modeling charts/plots and language data. Specifically we propose several pretraining tasks that cover plot deconstruction and numerical reasoning which are the key capabilities in visual language modeling. We perform the MATCHA pretraining starting from Pix2Struct, a recently proposed image-to-text visual language model. On standard benchmarks such as PlotQA and ChartQA, the MATCHA model outperforms state-of-the-art methods by as much as nearly 20%. We also examine how well the MATCHA pretraining transfers to domains such as screenshots, textbook diagrams, and document figures and observe overall improvement, verifying the usefulness of MATCHA pretraining on broader visual language tasks.12

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Chart Derendering
  • 3.2 Math Reasoning
  • 4 Experiment
  • 4.1 Experimental Setups
  • 4.2 Main Results
  • 4.3 Results on Pix2Struct Tasks
  • 5 Analyses and Discussions
  • 5.1 Ablation Study
  • 5.2 Fine-grained Analysis and Error Analysis
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • References
  • A More Details on Datasets Used
  • B Details of Baselines

Knowls

  1. Knowl 1 — MATCHA’s joint visual-language pretraining design

    model/method

    MATCHA is a continual-pretraining method for Pix2Struct’s image-to-text Transformer. It adds two objectives to the original Pix2Struct screenshot-parsing objective:

    • Chart derendering: map a chart image to its underlying data table or rendering code, teaching layout understanding, number extraction, and organization of visual elements.
    • Math reasoning: map an image containing a rendered mathematical or numerical-reasoning problem to its answer, teaching operations over extracted numbers.
    • Screenshot parsing: continue predicting simplified HTML from partially masked website screenshots to reduce catastrophic forgetting of Pix2Struct’s original capability.

    The workflow illustration on page 2 shows these objectives as two complementary branches: chart images are decoded into code or tables, while rendered questions or passages are decoded into numerical answers. Because MATCHA has no separate textual-input branch, text queries are rendered as part of the input image; for chart question answering, the question is placed as a header above the chart.

  2. Knowl 2 — Chart derendering data and supervision

    model/method

    MATCHA constructs chart-derendering supervision from independent chart-to-code and chart-to-table pairs. For chart-to-code pairs, the authors crawl appropriately licensed GitHub IPython notebooks, save figures together with the code block immediately preceding each figure, and train the model to reconstruct that code. These code snippets can be noisy because preceding code may contain irrelevant statements or omit earlier statements needed to create the chart.

    For chart-to-table pairs, the authors create synthetic charts from web-crawled Wikipedia tables by randomly varying the plotting library, chart type, style, colors, whether values are displayed, and text fonts and sizes. The synthetic corpus uses matplotlib or seaborn and bar, line, or pie charts. It is supplemented with PlotQA-generated chart-table pairs and approximately 20,000 real-world chart-table pairs collected from ChartQA sources: Statista, Pew, Our World in Data, and OECD.

    To avoid test leakage, only ChartQA and PlotQA training-set chart-table pairs are used for pretraining; test charts and tables are excluded. The chart-to-code example shown in the paper’s page-2 visual includes an Airbus A380 chart, plotting code, and its underlying table, illustrating the intended reconstruction targets.

  3. Knowl 3 — Rendered mathematical reasoning pretraining

    model/method

    MATCHA uses two textual numerical-reasoning datasets as image-to-text pretraining data. MATH supplies synthetic arithmetic and comparison questions, while DROP supplies reading-comprehension questions over paragraphs that require extracting relevant numbers and performing discrete operations. DROP contains 96,000 question-answer pairs over 6,700 paragraphs.

    The authors render each MATH input as an image. For DROP, the paragraph and question are concatenated and rendered as one image. The target is the answer, not an intermediate reasoning trace. The selected MATH modules cover arithmetic operations such as addition, subtraction, multiplication, division, mixed operations, and multiple-operation variants, as well as comparison operations such as closest value, pair comparison, sorting, and kth-largest-value queries, including composed variants.

    The two datasets provide complementary supervision: MATH offers large amounts of operation-specific numerical data, whereas DROP more closely resembles downstream visual question answering because information extraction and numerical reasoning must be performed jointly.

  4. Knowl 4 — MATCHA pretraining mixture and optimization setup

    experimental setup

    The final MATCHA pretraining mixture samples examples with the following rates and reported dataset sizes:

    • MATH: 20%, 2M examples.
    • DROP: 20%, 96K examples.
    • Chart-to-code from GitHub: 4%, 23M examples.
    • Synthetic chart-to-table pairs created by the authors: 12%, 270K examples.
    • ChartQA chart-to-table pairs: 12%, 22K examples.
    • PlotQA chart-to-table pairs: 12%, 224K examples.
    • Pix2Struct screenshot parsing: 20%, 80M examples.

    Thus, math reasoning contributes 40% of sampling, chart derendering contributes 40%, and screenshot parsing contributes 20%. The chart-to-code rate was reduced from an initially equal allocation because its noisy supervision caused training instability; the freed probability mass was assigned to chart-to-table tasks.

    Pix2Struct is pretrained for 100,000 steps with batch size 512 and maximum sequence length 192; the final MATCHA checkpoint is selected at step 90,000 using validation exact match. Downstream fine-tuning uses batch size 256 and maximum sequence length 128, runs for 10,000 steps on ChartQA and Chart-to-Text, and runs for 20,000 steps on PlotQA. MATCHA and Pix2Struct use 64 GCP-TPUv3 devices.

  5. Knowl 5 — Evaluation tasks and scoring protocol

    experimental setup

    MATCHA is evaluated on chart question answering, chart summarization, and non-chart visual-language tasks. ChartQA contains augmented and human-written subsets; PlotQA contains v1 and v2 subsets; Chart-to-Text contains Pew and Statista subsets. The augmented ChartQA and PlotQA-v1 data are more extractive, whereas ChartQA-human and PlotQA-v2 require more reasoning. Chart-to-Text-Pew summaries are automatically extracted, while Statista summaries are human written.

    The reported fine-tuning-set sizes are: ChartQA-human, 4.8K tables and 9.6K pairs; ChartQA-machine, 17.1K tables and 23.1K pairs; PlotQA-v1, 224K tables and 8M pairs; PlotQA-v2, 224K tables and 29M pairs; Chart-to-Text-Pew, 9K tables and 9K pairs; and Chart-to-Text-Statista, 35K tables and 35K pairs.

    ChartQA and PlotQA use relaxed correctness, meaning exact matching with tolerance for up to 5% numerical error. Chart-to-Text uses BLEU4. MATCHA receives chart images and rendered questions rather than gold underlying tables. For comparison, several baselines are explicitly given gold tables or use OCR-extracted tables.

  6. Knowl 6 — MATCHA’s main chart and plot benchmark results

    empirical result

    MATCHA substantially improves over Pix2Struct and other chart models without access to gold tables. The reported scores are:

    • MATCHA: ChartQA augmented 90.2, ChartQA human 38.2, ChartQA average 64.2; PlotQA v1 92.3, PlotQA v2 90.7, PlotQA average 91.5; Chart-to-Text Pew 12.2, Statista 39.4, average 25.8; overall average 60.5.
    • Pix2Struct: ChartQA augmented 81.6, human 30.5, average 56.0; PlotQA v1 73.2, v2 71.9, average 72.5; Chart-to-Text Pew 10.3, Statista 38.0, average 24.2; overall average 50.9.
    • PaLI-17B at resolution 588: ChartQA augmented 64.9, human 30.4, average 47.6; PlotQA v1 64.5, v2 15.2, average 39.8; Chart-to-Text Pew 11.2, Statista 41.4, average 26.3; overall average 37.9.
    • VL-T5 with gold tables: PlotQA v1 96.4, v2 84.7, average 90.6, and ChartQA average 59.1.

    MATCHA improves over the strongest no-gold-table baseline, Pix2Struct, by 8.2 points on ChartQA and by 19.0 points on the PlotQA average. It outperforms all compared systems on PlotQA-v2, the numerically demanding subset, including systems with gold tables. VL-T5 with a gold table remains approximately 4 points better on PlotQA-v1, whose synthetic extractive questions are simpler. MATCHA is better than Pix2Struct on both Chart-to-Text subsets and establishes the best reported Pew score, although PaLI-17B at resolution 588 scores higher on Statista. Across all reported setups, MATCHA is the strongest model without gold-table access and has an overall average approximately 10 points above Pix2Struct.

  7. Knowl 7 — Transfer beyond charts and plots

    empirical result

    Replacing Pix2Struct’s initial checkpoint with MATCHA improves performance on most of the additional Pix2Struct visual-language tasks. The reported scores for Pix2Struct versus MATCHA are:

    • ChartQA: 56.0 versus 64.2.
    • AI2D textbook-diagram QA: 40.9 versus 42.6.
    • OCR-VQA: 69.4 versus 68.9.
    • RefExp: 92.2 versus 94.2.
    • Widget Captioning: 133.1 versus 137.7.
    • Screen2Words: 107.0 versus 106.2.
    • TextCaps: 88.0 versus 92.4.
    • DocVQA: 72.1 versus 74.2.
    • InfoVQA: 38.2 versus 37.2.

    The average across all tasks increases from 77.4 to 79.7, while the average excluding ChartQA increases from 80.1 to 81.7. The latter is a 1.6-point improvement outside the chart domain; across all tasks, the improvement is 2.3 points. This transfer includes textbook diagrams, screenshots, and scanned documents, indicating that the pretraining signal is not confined to plots and charts.

  8. Knowl 8 — Ablation evidence for each pretraining component

    empirical result

    A 50,000-step ablation study fine-tunes each checkpoint on ChartQA. The full mixture obtains 88.6 on ChartQA-augmented, 37.4 on ChartQA-human, and 63.0 averaged across the two subsets. Removing components produces the following scores:

    • No math reasoning: 88.2, 33.0, and 60.6.
    • No chart derendering: 83.7, 34.4, and 59.1.
    • No screenshot parsing: 87.8, 34.9, and 61.4.
    • No MATH dataset: 88.2, 36.7, and 62.5.
    • No DROP dataset: 88.2, 34.3, and 61.3.
    • No real-world chart-table pairs: 87.4, 34.5, and 61.0.
    • No chart-to-code data: 89.1, 34.6, and 61.9.

    Removing chart derendering causes the largest component-level decrease, approximately 4 points overall; removing math reasoning decreases the average by 2.4 points, and removing screenshot parsing decreases it by 1.6 points. Math reasoning matters more for the human subset, while chart derendering matters more for the augmented subset. Removing DROP causes a 1.7-point average decrease versus 0.5 points for removing MATH, which the authors associate with DROP’s closer match to downstream question answering. Removing real-world chart-table pairs decreases the overall score by about 2 points, especially on the human subset, and removing chart-to-code data decreases the score by about 1.1 points, mainly on human questions.

    Continuing Pix2Struct pretraining with only its original screenshot-parsing objective improves ChartQA from 56.0 to 57.0 after 50,000 steps, whereas adding the MATCHA objectives produces 64.2 under the corresponding full-pretraining comparison.

  9. Knowl 9 — Fine-grained strengths and remaining error modes

    empirical result

    The authors manually classify 100 ChartQA test examples into three possibly overlapping challenge categories after excluding 7 annotation errors. Complex data extraction occurs in 55.9% of examples, numerical reasoning in 45.2%, and plot-attribute questions involving color, shape, or location in 7.5%.

    The category-level accuracy chart on page 8 reports the following values for PaLI at resolution 588, Pix2Struct, and MATCHA, respectively: data extraction 51.9, 69.2, and 76.9; math reasoning 26.2, 23.8, and 31.0; and plot attributes 28.6, 0.0, and 14.3. MATCHA improves over Pix2Struct in all three categories and exceeds PaLI on data extraction and math reasoning, but it remains behind PaLI on plot attributes.

    Among 100 MATCHA errors, after excluding 21 annotation errors, 48.3% involve math reasoning, 43.4% involve data extraction, and 8.0% involve plot attributes. Thus, numerical reasoning remains the dominant failure mode. MATCHA still struggles with sophisticated multi-step calculations and high-precision arithmetic; the paper’s example where all models fail requires recognizing that 6.67+5.8+5.63=18.1<18.186.67+5.8+5.63=18.1<18.18.

  10. Knowl 10 — Limitations of MATCHA

    limitation

    MATCHA’s error analysis shows that complex numerical reasoning and high-precision calculation remain unresolved. The authors also question whether performing arithmetic entirely in neural weight space is preferable to approaches that delegate computation to calculators, programs, or other external tools.

    MATCHA underperforms PaLI on plot attributes such as colors, shapes, and locations. The authors conjecture that this reflects MATCHA’s lack of PaLI’s massive grounded image-text pretraining with rich semantic associations. Chart-to-code pretraining supplies only limited attribute supervision because many plot characteristics are implicit defaults of plotting libraries rather than explicitly represented in the code.

    The reported experimental numbers come from single runs. Although the authors evaluate many datasets and ablations, they acknowledge that multiple runs would be needed to quantify variance given sufficient computational resources. Finally, the method is evaluated on only some kinds of visual language; other systems, such as comics and manga, may use distinct visual vocabularies or grammars.

Coverage note — Secondary baseline descriptions, individual illustrative case studies, references, ethics statements, and related-work discussion were omitted because they do not add load-bearing contributed method, theory, or evaluation content beyond the included results and limitations.

References

  1. 1.Mubashara Akhtar, Oana Cocarascu, and Elena Simperl. 2023. Reading and reasoning over chart images for evidence-based automated fact-checking. In Findings of the Association for Computational Linguistics: EACL 2023, pages 399–414, Dubrovnik, Croatia. Association for Computational Linguistics.
  2. 2.Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016. Neural module networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 39–48.
  3. 3.Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen. 2022. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. arXiv preprint arXiv:2211.12588.
  4. 4.Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. 2023. Pali: A jointly-scaled multilingual language-image model. In The Eleventh International Conference on Learning Representations.
  5. 5.Zhoujun Cheng, Tianbao Xie, Peng Shi, Chengzu Li, Rahul Nadkarni, Yushi Hu, Caiming Xiong, Dragomir Radev, Mari Ostendorf, Luke Zettlemoyer, et al. 2023. Binding language models in symbolic languages. In The Eleventh International Conference on Learning Representations.
  6. 6.Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. 2021. Unifying vision-and-language tasks via text generation. In International Conference on Machine Learning, pages 1931–1942. PMLR.
  7. 7.Neil Cohn. 2013. The Visual Language of Comics: Introduction to the Structure and Cognition of Sequential Images. A&C Black.
  8. 8.Brian L. Davis, B. Morse, Bryan Price, Chris Tensemeyer, Curtis Wigington, and Vlad I. Morariu. 2023. End-to-end document recognition and understanding with dessurt. In Computer Vision – ECCV 2022 Workshops: Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part IV, page 280–296, Berlin, Heidelberg. Springer-Verlag.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  10. 10.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations.
  11. 11.Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2368–2378, Minneapolis, Minnesota. Association for Computational Linguistics.
  12. 12.Julian Eisenschlos, Syrine Krichene, and Thomas Müller. 2020. Understanding tables with intermediate pre-training. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 281–296, Online. Association for Computational Linguistics.
  13. 13.Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2022. PAL: Program-aided language models. arXiv preprint arXiv:2211.10435.
  14. 14.Mor Geva, Ankit Gupta, and Jonathan Berant. 2020. Injecting numerical reasoning skills into language models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 946–958, Online. Association for Computational Linguistics.
  15. 15.Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Eisenschlos. 2020. TaPas: Weakly supervised table parsing via pre-training. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4320–4333, Online. Association for Computational Linguistics.
  16. 16.Robert E Horn. 1998. Visual language. MacroVu Inc. Washington.
  17. 17.Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei. 2022. LayoutLMv3: Pre-training for document ai with unified text and image masking. In Proceedings of the 30th ACM International Conference on Multimedia, MM ’22, page 4083–4091, New York, NY, USA. Association for Computing Machinery.
  18. 18.Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. 2017. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2901–2910.
  19. 19.Kushal Kafle, Brian Price, Scott Cohen, and Christopher Kanan. 2018. DVQA: Understanding data visualizations via question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5648–5656.
  20. 20.Samira Ebrahimi Kahou, Vincent Michalski, Adam Atkinson, Ákos Kádár, Adam Trischler, and Yoshua Bengio. 2017. FigureQA: An annotated figure dataset for visual reasoning. arXiv preprint arXiv:1710.07300.
  21. 21.Shankar Kantharaj, Rixie Tiffany Leong, Xiang Lin, Ahmed Masry, Megh Thakkar, Enamul Hoque, and Shafiq Joty. 2022. Chart-to-text: A large-scale benchmark for chart summarization. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4005–4023, Dublin, Ireland. Association for Computational Linguistics.
  22. 22.Jihyung Kil, Soravit Changpinyo, Xi Chen, Hexiang Hu, Sebastian Goodman, Wei-Lun Chao, and Radu Soricut. 2022. PreSTU: Pre-training for scene-text understanding. arXiv preprint arXiv:2209.05534.
  23. 23.Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park. 2022. OCR-free document understanding transformer. In European Conference on Computer Vision, pages 498–517. Springer.
  24. 24.Kenton Lee, Mandar Joshi, Iulia Turc, Hexiang Hu, Fangyu Liu, Julian Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova. 2023. Pix2Struct: Screenshot parsing as pretraining for visual language understanding. In Proceedings of the 40th International Conference on Machine Learning.
  25. 25.Matan Levy, Rami Ben-Ari, and Dani Lischinski. 2022. Classification-regression for chart comprehension. In European Conference on Computer Vision, pages 469–484. Springer.
  26. 26.Fangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy, Nigel Collier, and Desmond Elliott. 2021. Visually grounded reasoning across languages and cultures. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10467–10485, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  27. 27.Fangyu Liu, Julian Martin Eisenschlos, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Wenhu Chen, Nigel Collier, and Yasemin Altun. 2023. DePlot: One-shot visual language reasoning by plot-to-table translation. In Findings of the Association for Computational Linguistics: ACL 2023. Association for Computational Linguistics.
  28. 28.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  29. 29.Ahmed Masry, Do Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. 2022. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. In Findings of the Association for Computational Linguistics: ACL 2022, pages 2263–2279, Dublin, Ireland. Association for Computational Linguistics.
  30. 30.Nitesh Methani, Pritha Ganguly, Mitesh M Khapra, and Pratyush Kumar. 2020. PlotQA: Reasoning over scientific plots. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1527–1536.
  31. 31.Xinyu Pi, Qian Liu, Bei Chen, Morteza Ziyadi, Zeqi Lin, Qiang Fu, Yan Gao, Jian-Guang Lou, and Weizhu Chen. 2022. Reasoning like program executors. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 761–779, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  32. 32.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67.
  33. 33.David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli. 2019. Analysing mathematical reasoning abilities of neural models. In International Conference on Learning Representations.
  34. 34.Alane Suhr, Mike Lewis, James Yeh, and Yoav Artzi. 2017. A corpus of natural language for visual reasoning. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 217–223, Vancouver, Canada. Association for Computational Linguistics.
  35. 35.Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. 2019. A corpus for reasoning about natural language grounded in photographs. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6418–6428, Florence, Italy. Association for Computational Linguistics.
  36. 36.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  37. 37.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  38. 38.Yuhuai Wu, Felix Li, and Percy Liang. 2022. Insights into pre-training via simpler synthetic tasks. In Advances in Neural Information Processing Systems.
  39. 39.Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020. LayoutLM: Pre-training of text and layout for document image understanding. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 1192–1200.

Citation

MLA
Liu, F., et al. “MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 12756–70, https://doi.org/10.18653/v1/2023.acl-long.714.
APA
Liu, F., Piccinno, F., Krichene, S., Pang, C., Lee, K., Joshi, M., Altun, Y., Collier, N., & Eisenschlos, J. (2023). MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12756–12770. https://doi.org/10.18653/v1/2023.acl-long.714
Chicago
Liu, F., F. Piccinno, S. Krichene, et al. 2023. “MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12756–70. https://doi.org/10.18653/v1/2023.acl-long.714.
Harvard
Liu, F. et al. (2023) “MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 12756–12770. Available at: https://doi.org/10.18653/v1/2023.acl-long.714.
Vancouver
1. Liu F, Piccinno F, Krichene S, Pang C, Lee K, Joshi M, Altun Y, Collier N, Eisenschlos J (2023) MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 12756–12770

BibTeX

@inproceedings{liu-etal-2023-matcha,
    title = "{M}at{C}ha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering",
    author = "Liu, Fangyu  and
      Piccinno, Francesco  and
      Krichene, Syrine  and
      Pang, Chenxi  and
      Lee, Kenton  and
      Joshi, Mandar  and
      Altun, Yasemin  and
      Collier, Nigel  and
      Eisenschlos, Julian",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.714/",
    doi = "10.18653/v1/2023.acl-long.714",
    pages = "12756--12770"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/