An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models

Fatemeh ShiriXiao-Yu GuoMona FarXin YuReza HafYuan-Fang Li

article2024EMNLP76 citations

Presents the Spatial-MM benchmark to expose critical weaknesses in large multimodal models, showing that while symbolic aids like bounding boxes improve performance, models still fail on human-perspective viewpoints and gain no benefit from chain-of-thought prompting on complex spatial questions.

Listen

Modern artificial intelligence systems increasingly rely on Large Multimodal Models (LMMs) to interpret both visual images and text. While these models perform well on broad vision-and-language tasks, their ability to understand spatial arrangements and reason through complex physical relationships remains unreliable. To address this critical gap, the article introduces Spatial-MM, a new benchmark designed to evaluate spatial reasoning across both single-step and multi-hop scenarios. The evaluation assesses leading commercial and open-source models—including GPT-4o, GPT-4 Vision, Gemini 1.5 Pro, and LLaVA-1.5—across different spatial viewpoints, visual grounding techniques, and structured reasoning paths.

The experimental findings show significant limitations in how current models handle spatial data. First, models perform markedly worse when required to answer questions from the perspective of an entity inside the image (such as an observed human) compared to the standard camera viewpoint; for example, GPT-4o's accuracy drops from over 63% from the camera view to 27.5% from an internal human perspective. Second, traditional text-based chain-of-thought prompting fails to improve spatial accuracy in multi-hop questions, with models often performing worse than under standard prompting. Third, detailed failure analysis demonstrates that when multi-hop reasoning fails, an incorrect spatial step is responsible roughly 91% of the time, whereas non-spatial steps rarely cause errors. Conversely, providing structured visual grounding—such as synthesized bounding boxes or scene graphs—substantially improves accuracy across all evaluated models.

These results demonstrate that while current AI systems excel at basic object recognition, they cannot be trusted to independently manage complex spatial relationships or perspective-shifting tasks. Relying on raw multimodal models for spatial decision-making in fields like robotics, autonomous systems, or spatial navigation creates substantial operational and safety risks. To mitigate these risks, systems requiring spatial reasoning should integrate explicit visual scaffolding, such as automated scene graphs or localized bounding boxes, rather than relying on chain-of-thought textual prompting alone. Future development should focus on expanding larger-scale multi-hop spatial training data and refining models' intrinsic perspective-taking capabilities.

Cover for An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models

Abstract

Large Multimodal Models (LMMs) have achieved strong performance across a range of vision and language tasks. However, their spatial reasoning capabilities are under-investigated. In this paper, we construct a novel VQA dataset, Spatial-MM, to comprehensively study LMMs' spatial understanding and reasoning capabilities. Our analyses on object-relationship and multi-hop reasoning reveal several important findings. Firstly, bounding boxes and scene graphs, even synthetic ones, can significantly enhance LMMs' spatial reasoning. Secondly, LMMs struggle more with questions posed from the human perspective than the camera perspective about the image. Thirdly, chain of thought (CoT) prompting does not improve model performance on complex multi-hop questions involving spatial relations. Lastly, our perturbation analysis on GQA-spatial reveals that LMMs are much stronger at basic object detection than complex spatial reasoning. We believe our new benchmark dataset and in-depth analyses can spark further research on LMMs spatial reasoning.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Large Multimodal Models
  • 2.2 Spatial relationship Benchmarks
  • 2.3 Multi-hop Reasoning
  • 3 The Spatial-MM Benchmark
  • 3.1 Spatial-Obj
  • 3.2 Spatial-CoT
  • 3.3 Human performance
  • 4 Data Enrichment to Improve Spatial Reasoning
  • 5 Experiments
  • 5.1 Dataset
  • 5.2 LMMs
  • 5.3 Evaluation Metrics
  • 5.4 The Effect of Bounding Boxes and Scene Graphs
  • 5.5 The Effect of Human and Camera Perspectives
  • 5.6 Analysis based on Complex Multi-hop VQAs
  • 5.7 Analysis of Perturbations on GQA-spatial
  • 6 Conclusion
  • 7 Limitation
  • Acknowledgments
  • References
  • A LMMs
  • B Spatial-MM Benchmark Samples
  • C Pipeline
  • D Bounding Box Example
  • E Spatial Relationships Types
  • F Preposition analysis
  • G Prompts

Knowls

  1. Knowl 1 — Spatial-MM tests diverse spatial relations and viewpoints

    experimental setup

    Spatial-MM is a visual question-answering benchmark for spatial understanding, built from natural images collected from the Internet. Its Spatial-Obj subset contains 2,000 multiple-choice questions about one or two objects, including yes/no and wh-type questions. Three annotators created question-answer pairs, and another ten annotators reviewed them for correctness or ambiguity. Questions cover 36 spatial relations, including left/right, front/behind, above/below, inside/outside, touching, and facing direction. The questions span four visually challenging patterns: object localization, orientation and direction, viewpoint, and positional or relational context. Both camera and human viewpoints are represented. In a human quality check, annotators found the correct answer clearly identifiable in 98% of 300 sampled Spatial-Obj items.

  2. Knowl 2 — Spatial-CoT provides multi-hop questions with annotated reasoning paths

    data/table

    Spatial-CoT contains 310 open-ended, spatial-aware multi-hop questions about Internet images. GPT-4o first generated 800 candidate question-answer pairs; 178 were removed for lacking a target spatial relation, and human annotators discarded another 312 for insufficient multi-hop complexity. GPT-4o also drafted a reasoning path for each retained item, which three annotators could keep, remove, add to, or modify. They tagged each step as spatial (S) or non-spatial (NS). Every final path has at least two steps and at least one spatial step; 67% of all steps are spatial and 33% non-spatial. The question paths require two hops in 34% of cases, three hops in 39%, and four or more hops in 27%. In a human clarity assessment of 100 sampled items, annotators identified the correct answer in 99% of cases.

  3. Knowl 3 — GPT-4o-generated boxes and scene graphs supply visual grounding

    model/method

    The paper evaluates two ways of augmenting an image and question with spatial grounding. For bounding boxes, key objects are extracted from the question and GPT-4o is prompted to locate them in the image. Each box uses normalized image coordinates [xmin⁡,ymin⁡,xmax⁡,ymax⁡][x_{\min}, y_{\min}, x_{\max}, y_{\max}] in [0,1][0,1], with the first pair marking the top-left and the second pair the bottom-right corner; the prompt requests fewer than ten boxes. For scene graphs, GPT-4o first produces a spatially aware caption considering the key objects’ locations, directions, orientations, and relations, then extracts relations among those objects from the image, caption, and question. The resulting boxes or relations are provided alongside the image and question for answer prediction. The experiments compare synthesized grounding with ground-truth boxes or scene graphs where available. For multi-hop evaluation, models are prompted to express their reasoning steps as object attributes and relations in a scene-graph-like format so the steps can be compared semantically with human-checked reference paths.

  4. Knowl 4 — Visual grounding improves accuracy, with benefits depending on task and grounding type

    empirical result

    On Spatial-Obj, adding scene graphs improved GPT-4 Vision accuracy from 50.80% to 60.08% on one-object questions and from 52.85% to 61.77% on two-object questions. The reported gains are 9.28 and 8.92 percentage points, respectively. Across the models tested, boxes were generally more useful for one-object questions and scene graphs more useful for two-object questions on Spatial-Obj. On GQA-spatial, synthesized boxes raised overall accuracy for all four models: GPT-4o from 56.30% to 81.67%, Gemini 1.5 Pro from 18.79% to 43.29%, LLaVA-1.5 from 24.75% to 74.78%, and MiniGPT-v2 from 19.43% to 48.20%. The mean increase across those four comparisons is about 32.2 percentage points. Synthesized boxes also outperformed ground-truth boxes on overall GQA-spatial accuracy for each of those models; for example, GPT-4o scored 81.67% with synthesized boxes and 71.91% with ground-truth boxes. Thus, spatial grounding often helps, but whether boxes or relations help more varies by dataset and question type.

  5. Knowl 5 — Models perform worse from a depicted person’s viewpoint

    empirical result

    The benchmark compares questions explicitly asked from the camera’s viewpoint with questions asked from a human viewpoint inside the image. Human-perspective items make up 31% of Spatial-Obj and 46% of Spatial-CoT; the remainder use the camera perspective. All five evaluated models scored lower on the human-perspective questions. GPT-4o, for example, scored 43.75% versus 73.60% on Spatial-Obj and 27.50% versus 63.25% on Spatial-CoT (human versus camera viewpoint). The same direction of difference appears for GPT-4 Vision (42.56% versus 54.81% on Spatial-Obj; 33.05% versus 50.14% on Spatial-CoT), LLaVA-1.5 (37.70% versus 46.04%; 22.91% versus 40.47%), Gemini 1.5 Pro (32.92% versus 44.07%; 25.81% versus 46.29%), and MiniGPT-v2 (31.67% versus 40.75%; 22.08% versus 39.77%). Explicitly naming the viewpoint in the question therefore did not eliminate the models’ difficulty with human-centered perspective.

  6. Knowl 6 — Chain-of-thought prompting does not reliably help multi-hop spatial QA

    empirical result

    On Spatial-CoT, conventional chain-of-thought (CoT) prompting did not consistently improve final-answer accuracy over standard prompting. Overall accuracy for GPT-4o fell from 60.78% to 56.91%, for Gemini 1.5 Pro from 49.02% to 33.33%, and for LLaVA-1.5 from 49.02% to 39.22%; GPT-4 Vision was an exception, increasing from 54.90% to 56.86%. By comparison, scene-graph augmentation produced overall accuracies of 63.27% for GPT-4o, 60.00% for Gemini 1.5 Pro, 54.90% for LLaVA-1.5, and 66.00% for GPT-4 Vision. Accuracy also declined as reference paths grew longer: questions with four or more hops averaged 12 percentage points lower accuracy than two-hop questions and 7.5 points lower than three-hop questions across the reported settings. The results indicate that ordinary verbal CoT is not a dependable aid for these spatial multi-hop questions, while explicit visual grounding—particularly scene graphs—can help.

  7. Knowl 7 — Incorrect answers are usually associated with spatial reasoning errors

    empirical result

    The authors assessed reasoning paths for questions each model answered incorrectly under standard prompting by comparing model-generated steps with the human-checked reference path. A step counted as incorrect if it did not semantically match a reference step or if a reference step was missing. On average, only 1.25% of these questions had a correct path despite an incorrect final answer. The reported path-error categories were 91.11% with at least one incorrect spatial step and 9.75% with at least one incorrect non-spatial step, suggesting that spatial-step errors are the more common problem. At the individual-model level, the spatial and non-spatial error percentages were, respectively, 92.55% and 6.91% for GPT-4o, 91.76% and 8.24% for GPT-4 Vision, 89.40% and 12.58% for LLaVA-1.5, and 90.73% and 11.26% for Gemini 1.5 Pro. In a separate assessment of step-level classification, spatial-step F1 was lower than non-spatial-step F1 by about 21 percentage points with standard prompting and 19 points with scene-graph augmentation.

  8. Knowl 8 — GQA-spatial perturbations expose reliance on object recognition over relation reasoning

    empirical result

    The perturbation study changes GQA-spatial’s two caption options, which normally differ in their spatial preposition. It tests adding a “None of the above” choice, negating the correct relation, swapping the objects, replacing the shared object in both options with one absent from the image, and replacing the object only in the correct option with an absent object. Replacing the object in both options yielded very high accuracy: Gemini 1.5 Pro scored 91.87% on one-object and 69.10% on two-object questions, while LLaVA-1.5 scored 90.69% and 54.98%. Changing only the object in the correct option lowered Gemini’s scores to 45.76% and 30.21%, and LLaVA’s to 61.55% and 52.58%. This contrast is consistent with models being able to use object-presence cues while struggling to resolve spatial relations. In the object-swapping condition with a “None” choice, two-object accuracy was only 2.78% for Gemini and 1.37% for LLaVA; the models overwhelmingly selected the original option B, including when it was incorrect. Negation was also challenging: Gemini’s accuracy with negation and “None” was 1.75% and 2.176% for one- and two-object items, while LLaVA’s was 41.38% and 8.93%. The perturbation results show that performance is sensitive to changes in wording and object identity, not merely to the spatial relation being tested.

  9. Knowl 9 — The effect of grounding varies across spatial relation categories

    empirical result

    On GQA-spatial, the relation-specific analysis grouped items into front/behind, left/right, and top/bottom; the reported counts were 26, 1,032, and 393, respectively. Scene-graph augmentation improved GPT-4 Vision’s accuracy over standard prompting in all three groups: front/behind from 23.71% to 77.32%, left/right from 18.18% to 75.38%, and top/bottom from 63.10% to 100.00%. Gemini 1.5 Pro also improved with scene graphs in each group: 18.52% to 45.24%, 19.16% to 49.36%, and 17.05% to 33.59%, respectively. LLaVA-1.5 showed smaller and nonuniform gains: 29.57% to 41.83% for front/behind, 33.43% to 48.93% for left/right, and 20.10% to 22.90% for top/bottom. These results show that grounding effects differ both by model and by relation category; the particularly small front/behind sample also limits how broadly that category’s scores can be interpreted.

  10. Knowl 10 — Generating multi-hop questions and reference paths remains a scale limitation

    limitation

    Spatial-CoT question-answer pairs were initially generated with GPT-4 and then filtered or edited by human annotators; reference reasoning paths were generated through a similar model-assisted process and then reviewed by annotators. The authors note that this workflow reduces the need for annotators to write complex questions from scratch, but creating a larger labeled dataset remains difficult. The same scale constraint applies to producing reliable ground-truth reasoning paths.

Coverage note — The appendix’s full prompt templates and additional illustrative examples are omitted because they document implementation and examples rather than adding separate findings beyond the benchmark, grounding procedures, and analyses captured here.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  2. 2.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  3. 3.Boyuan Chen, Zhuo Xu, Sean Kirmani, Brian Ichter, Danny Driess, Pete Florence, Dorsa Sadigh, Leonidas Guibas, and Fei Xia. 2024. Spatialvlm: Endowing vision-language models with spatial reasoning capabilities. arXiv preprint arXiv:2401.12168.
  4. 4.Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechun Liu, Pengchuan Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong, and Mohamed Elhoseiny. 2023. Minigpt-v2: large language model as a unified interface for vision-language multi-task learning. arXiv preprint arXiv:2310.09478.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  6. 6.Gemini Team. 2024. Gemini: A family of highly capable multimodal models. Preprint, arXiv:2312.11805.
  7. 7.Drew A Hudson and Christopher D Manning. 2019. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6700–6709.
  8. 8.Amita Kamath, Jack Hessel, and Kai-Wei Chang. 2023. What’s" up" with vision-language models? investigating their struggle with spatial reasoning. arXiv preprint arXiv:2310.19785.
  9. 9.Jihyung Kil, Farideh Tavazoee, Dongyeop Kang, and Joo-Kyung Kim. 2024. II-MMR: identifying and improving multi-modal multi-hop reasoning in visual question answering. CoRR, abs/2402.11058.
  10. 10.Xuanyu Lei, Zonghan Yang, Xinrui Chen, Peng Li, and Yang Liu. 2024a. Scaffolding coordinates to promote vision-language coordination in large multi-modal models. arXiv preprint arXiv:2402.12058.
  11. 11.Xuanyu Lei, Zonghan Yang, Xinrui Chen, Peng Li, and Yang Liu. 2024b. Scaffolding coordinates to promote vision-language coordination in large multi-modal models. Preprint, arXiv:2402.12058.
  12. 12.Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge, and Ying Shan. 2023a. Seed-bench: Benchmarking multimodal llms with generative comprehension. arXiv preprint arXiv:2307.16125.
  13. 13.Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge, and Ying Shan. 2023b. Seed-bench: Benchmarking multimodal llms with generative comprehension. arXiv preprint arXiv:2307.16125.
  14. 14.Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. 2023c. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pages 19730–19742. PMLR.
  15. 15.Kunchang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. 2023d. Videochat: Chat-centric video understanding. CoRR, abs/2305.06355.
  16. 16.Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan. 2023. Video-llava: Learning united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122.
  17. 17.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023a. Improved baselines with visual instruction tuning.
  18. 18.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023b. Improved baselines with visual instruction tuning.
  19. 19.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023c. Visual instruction tuning.
  20. 20.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023d. Visual instruction tuning. In NeurIPS.
  21. 21.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  22. 22.Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. 2023e. Mmbench: Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281.
  23. 23.Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. 2023f. Mmbench: Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281.
  24. 24.Tengchao Lv, Yupan Huang, Jingye Chen, Lei Cui, Shuming Ma, Yaoyao Chang, Shaohan Huang, Wenhui Wang, Li Dong, Weiyao Luo, Shaoxiang Wu, Guoxin Wang, Cha Zhang, and Furu Wei. 2023. Kosmos-2.5: A multimodal literate model. CoRR, abs/2309.11419.
  25. 25.Michael R Lyu, Baishakhi Ray, Abhik Roychoudhury, Shin Hwei Tan, and Patanamon Thongtanunam. 2024. Automatic programming: Large language models and beyond. arXiv preprint arXiv:2405.02213.
  26. 26.Muhammad Maaz, Hanoona Abdul Rasheed, Salman H. Khan, and Fahad Shahbaz Khan. 2023. Video-chatgpt: Towards detailed video understanding via large vision and language models. CoRR, abs/2306.05424.
  27. 27.Cristiane Kutianski Marchi Fagundes, Kristin Stock, and Luciene Stamato Delazari. 2021. A cross-linguistic study of spatial location descriptions in new zealand english and brazilian portuguese natural language. Transactions in GIS, 25(6):3159–3187.
  28. 28.Minh-Vuong Nguyen, Linhao Luo, Fatemeh Shiri, Dinh Phung, Yuan-Fang Li, Thuy-Trang Vu, and Gholamreza Haffari. 2024. Direct evaluation of chain-of-thought in multi-hop reasoning with knowledge graphs. In ACL 2024.
  29. 29.OpenAI. 2023. GPT-4 technical report. arXiv preprint arXiv:2303.08774.
  30. 30.OpenAI. 2024. Hello gpt-4o.
  31. 31.Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. 2023. Kosmos-2: Grounding multimodal large language models to the world. CoRR, abs/2306.14824.
  32. 32.Archiki Prasad, Elias Stengel-Eskin, and Mohit Bansal. 2023. Rephrase, augment, reason: Visual grounding of questions for vision-language models. arXiv preprint arXiv:2310.05861.
  33. 33.Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie. 2024. Eyes wide shut? exploring the visual shortcomings of multimodal llms. arXiv preprint arXiv:2401.06209.
  34. 34.Jiaan Wang, Yunlong Liang, Fandong Meng, Zengkui Sun, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023. Is chatgpt a good nlg evaluator? a preliminary study. arXiv preprint arXiv:2303.04048.
  35. 35.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022a. Chain of thought prompting elicits reasoning in large language models. CoRR, abs/2201.11903.
  36. 36.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022b. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837.
  37. 37.Kaizhi Zheng, Xuehai He, and Xin Eric Wang. 2023a. Minigpt-5: Interleaved vision-and-language generation via generative vokens. CoRR, abs/2310.02239.
  38. 38.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023b. Judging llm-as-a-judge with mt-bench and chatbot arena. Preprint, arXiv:2306.05685.
  39. 39.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. 2022. Least-to-most prompting enables complex reasoning in large language models. CoRR, abs/2205.10625.
  40. 40.Jin Peng Zhou, Charles Staats, Wenda Li, Christian Szegedy, Kilian Q Weinberger, and Yuhuai Wu. 2024. Don’t trust: Verify–grounding llm quantitative reasoning with autoformalization. In ICLR 2024.

Citation

MLA
Shiri, F., et al. “An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 21440–55, https://doi.org/10.18653/v1/2024.emnlp-main.1195.
APA
Shiri, F., Guo, X.-Y., Far, M. G., Yu, X., Haf, R., & Li, Y.-F. (2024). An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 21440–21455. https://doi.org/10.18653/v1/2024.emnlp-main.1195
Chicago
Shiri, F., X.-Y. Guo, M. G. Far, X. Yu, R. Haf, and Y.-F. Li. 2024. “An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 21440–55. https://doi.org/10.18653/v1/2024.emnlp-main.1195.
Harvard
Shiri, F. et al. (2024) “An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 21440–21455. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.1195.
Vancouver
1. Shiri F, Guo X-Y, Far MG, Yu X, Haf R, Li Y-F (2024) An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 21440–21455

BibTeX

@inproceedings{shiri-etal-2024-empirical,
    title = "An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models",
    author = "Shiri, Fatemeh  and
      Guo, Xiao-Yu  and
      Far, Mona Golestan  and
      Yu, Xin  and
      Haf, Reza  and
      Li, Yuan-Fang",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.1195/",
    doi = "10.18653/v1/2024.emnlp-main.1195",
    pages = "21440--21455"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/