Abstract Visual Reasoning with Tangram Shapes
Anya JiNoriyuki KojimaNoah RushAlane SuhrWai Keen VongRobert D. HawkinsYoav Artzi
Introduces KILOGRAM, a large-scale dataset of tangram shapes paired with part-level segmentations and natural language descriptions to benchmark and improve abstract visual reasoning in multimodal models.
Human communication routinely relies on visual abstraction, allowing people to refer to ambiguous geometric shapes, such as tangram puzzles, by comparing them to familiar real-world concepts like animals or people. However, existing artificial intelligence vision-and-language models are primarily trained on photographic images and struggle to generalize abstract visual concepts. Evaluating and improving these capabilities has historically been hindered by the lack of large-scale, structured datasets, as prior cognitive science studies typically relied on tiny stimulus sets of only 10 to 20 shapes.
The article introduces KILOGRAM, a large-scale dataset designed to study abstract visual reasoning and part-whole decomposition in humans and computational models. The primary objective is to evaluate how well state-of-the-art multimodal models generalize language to abstract shapes and to assess whether decomposing shapes into annotated parts enhances reference resolution.
To build this resource, the authors curated and digitized 1,016 vector-graphic tangrams and collected 13,404 crowdsourced English annotations covering whole-shape descriptions, piece-level part segmentations, and part labels. The evaluation framework tested two prominent multimodal architectures, CLIP (separate visual and linguistic encoders) and ViLT (a single joint visio-linguistic encoder), across multiple input conditions in a reference game setup where models identified a target shape among distractors from text descriptions. Human performance across 217 participants served as an experimental baseline.
The article presents several key findings. First, off-the-shelf pre-trained models demonstrated poor zero-shot abstract reasoning, achieving only 11% to 19% accuracy in 10-way reference games, barely above the 10% random baseline and far below human accuracy of 48% to 63%. Second, task-specific fine-tuning substantially improved performance for both models to between 41% and 77% accuracy. Third, providing joint part names and color-coded visual segmentations yielded massive performance gains for the joint-encoding ViLT model, which reached 77.3% accuracy on held-out test data, significantly outperforming its text-only baseline (44.5%) and the separate-encoding CLIP model (46.5%). Finally, human evaluation revealed two distinct proficiency clusters, with top-performing humans achieving 83.8% accuracy when part-level information was available.
These findings indicate that off-the-shelf multimodal models possess high structural capacity for abstract reasoning but lack appropriate training alignments during pre-training. Systems using tightly integrated joint representations can resolve abstract ambiguities much more effectively when structured, part-level alignments connect textual phrases directly to visual sub-components. This demonstrates that deploying vision-and-language models in abstract, non-photographic domains requires targeted fine-tuning and fine-grained visual-linguistic grounding to prevent high failure rates.
Organizations developing multimodal interfaces should incorporate structured part-whole alignments and avoid relying on zero-shot pre-trained models for abstract visual reasoning tasks. Future development should explore interactive communication settings, generation and instruction-following tasks, and expansion to multilingual benchmarks to capture cross-cultural naming variations.
The conclusions are subject to limitations, including English-only crowdsourcing and the use of static, isolated descriptions rather than dynamic, interactive dialogue. Additionally, because distractors in synthetic reference games can sometimes be ambiguous, evaluation accuracy faces a natural ceiling below 100%. Nevertheless, the findings offer robust confidence that structured part decomposition is critical for improving abstract visual reasoning in multimodal systems.
- Paper: CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning, Justin Johnson et al. (2016). CLEVR’s diagnostic approach to compositional visual reasoning provides a useful foundation for understanding KILOGRAM’s evaluation of structured reasoning over shapes.
- Paper: ReferItGame: Referring to Objects in Photographs of Natural Scenes, Sahar Kazemzadeh et al. (2014). ReferItGame establishes the crowdsourced reference-game paradigm that helps frame KILOGRAM’s task of matching language descriptions to visual targets.
- Paper: VQA: Visual Question Answering, Stanislaw Antol et al. (2015). VQA’s abstract-scene dataset offers an earlier example of testing visual reasoning on non-photographic images, clarifying the gap KILOGRAM addresses with tangrams.
No sufficiently relevant recommendations were found.
