Abstract Visual Reasoning with Tangram Shapes

Anya JiNoriyuki KojimaNoah RushAlane SuhrWai Keen VongRobert D. HawkinsYoav Artzi

article2022EMNLP62 citations

Introduces KILOGRAM, a large-scale dataset of tangram shapes paired with part-level segmentations and natural language descriptions to benchmark and improve abstract visual reasoning in multimodal models.

Listen

Human communication routinely relies on visual abstraction, allowing people to refer to ambiguous geometric shapes, such as tangram puzzles, by comparing them to familiar real-world concepts like animals or people. However, existing artificial intelligence vision-and-language models are primarily trained on photographic images and struggle to generalize abstract visual concepts. Evaluating and improving these capabilities has historically been hindered by the lack of large-scale, structured datasets, as prior cognitive science studies typically relied on tiny stimulus sets of only 10 to 20 shapes.

The article introduces KILOGRAM, a large-scale dataset designed to study abstract visual reasoning and part-whole decomposition in humans and computational models. The primary objective is to evaluate how well state-of-the-art multimodal models generalize language to abstract shapes and to assess whether decomposing shapes into annotated parts enhances reference resolution.

To build this resource, the authors curated and digitized 1,016 vector-graphic tangrams and collected 13,404 crowdsourced English annotations covering whole-shape descriptions, piece-level part segmentations, and part labels. The evaluation framework tested two prominent multimodal architectures, CLIP (separate visual and linguistic encoders) and ViLT (a single joint visio-linguistic encoder), across multiple input conditions in a reference game setup where models identified a target shape among distractors from text descriptions. Human performance across 217 participants served as an experimental baseline.

The article presents several key findings. First, off-the-shelf pre-trained models demonstrated poor zero-shot abstract reasoning, achieving only 11% to 19% accuracy in 10-way reference games, barely above the 10% random baseline and far below human accuracy of 48% to 63%. Second, task-specific fine-tuning substantially improved performance for both models to between 41% and 77% accuracy. Third, providing joint part names and color-coded visual segmentations yielded massive performance gains for the joint-encoding ViLT model, which reached 77.3% accuracy on held-out test data, significantly outperforming its text-only baseline (44.5%) and the separate-encoding CLIP model (46.5%). Finally, human evaluation revealed two distinct proficiency clusters, with top-performing humans achieving 83.8% accuracy when part-level information was available.

These findings indicate that off-the-shelf multimodal models possess high structural capacity for abstract reasoning but lack appropriate training alignments during pre-training. Systems using tightly integrated joint representations can resolve abstract ambiguities much more effectively when structured, part-level alignments connect textual phrases directly to visual sub-components. This demonstrates that deploying vision-and-language models in abstract, non-photographic domains requires targeted fine-tuning and fine-grained visual-linguistic grounding to prevent high failure rates.

Organizations developing multimodal interfaces should incorporate structured part-whole alignments and avoid relying on zero-shot pre-trained models for abstract visual reasoning tasks. Future development should explore interactive communication settings, generation and instruction-following tasks, and expansion to multilingual benchmarks to capture cross-cultural naming variations.

The conclusions are subject to limitations, including English-only crowdsourcing and the use of static, isolated descriptions rather than dynamic, interactive dialogue. Additionally, because distractors in synthetic reference games can sometimes be ambiguous, evaluation accuracy faces a natural ceiling below 100%. Nevertheless, the findings offer robust confidence that structured part decomposition is critical for improving abstract visual reasoning in multimodal systems.

No sufficiently relevant recommendations were found.

Cover for Abstract Visual Reasoning with Tangram Shapes

Abstract

We introduce KILOGRAM, a resource for studying abstract visual reasoning in humans and machines. Drawing on the history of tangram puzzles as stimuli in cognitive science, we build a richly annotated dataset that, with >1k distinct stimuli, is orders of magnitude larger and more diverse than prior resources. It is both visually and linguistically richer, moving beyond whole shape descriptions to include segmentation maps and part labels. We use this resource to evaluate the abstract visual reasoning capacities of recent multi-modal models. We observe that pre-trained weights demonstrate limited abstract reasoning, which dramatically improves with fine-tuning. We also observe that explicitly describing parts aids abstract reasoning for both humans and models, especially when jointly encoding the linguistic and visual inputs.

Table of Contents

  • 1 Introduction
  • 2 Background and Related Work
  • 3 Data Collection
  • 3.1 Collecting Tangram Puzzles
  • 3.2 Whole-Part Annotation
  • 3.3 Standard Data Splits
  • 4 Data Analysis
  • 5 Visual Reasoning with Tangrams
  • 5.1 Reference Game Generation
  • 5.2 Models
  • 5.3 Experimental Conditions
  • 5.4 Implementation Details
  • 5.5 Estimating Human Performance
  • 5.6 Results and Analysis
  • 6 Discussion
  • 7 Limitations
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Examples from KILOGRAM
  • A.2 Collecting Tangrams
  • A.3 Crowdsourcing Qualifications and Survey
  • A.4 Dense Annotation Sampling
  • A.5 Example Inputs for Experimental Conditions
  • A.6 Human Performance Baseline Details
  • A.7 Model-specific Implementation Details
  • A.8 Random Generation of Reference Games
  • A.9 Reproducibility Checklist

Knowls

  1. Knowl 1 — KILOGRAM is a large, part-annotated tangram resource

    data/table

    KILOGRAM contains 1,016 abstract tangram shapes: 1,004 digitized from published tangram puzzles and 12 added from prior studies. Each shape is represented as vector graphics made from its seven constituent puzzle pieces, enabling annotations of both the whole shape and its parts. The resource contains 13,404 English annotations, with multiple annotations per shape, and includes whole-shape descriptions, part descriptions, and a segmentation of puzzle pieces into parts. This scale and part-level structure support studying both human variation in abstract visual descriptions and models’ whole–part visual reasoning.

  2. Knowl 2 — Two-stage procedure for collecting whole-shape and part annotations

    experimental setup

    Crowdworkers first saw a tangram in grayscale and completed “This shape, as a whole, looks like ____.” They then selected one or more of its seven puzzle pieces and described the selected part or parts using “The part(s) you selected look(s) like ____.” The interface colored selected pieces to link each part description to its segmentation; workers could delete annotations, mark a part UNKNOWN, and add pieces to an existing part. Workers had to assign every piece before submitting, producing a complete segmentation. The 297 U.S.-based Mechanical Turk workers passed a qualification task, had at least a 98% prior HIT acceptance rate, and could annotate a given tangram only once and no more than 200 distinct tangrams overall. The initial collection obtained 10,053 annotations for 1,004 shapes, at least 10 per shape. A further 74 shapes—62 selected to cover observed annotation variability and 12 from earlier studies—received at least 50 annotations each, with a mean of 53.66, to better estimate their annotation distributions.

  3. Knowl 3 — Dataset statistics and learning splits

    data/table

    Across KILOGRAM’s 13,404 annotations for 1,016 tangrams, the mean whole-shape description length is 2.28±1.622.28\pm1.62 tokens and the mean part-description length is 1.31±0.771.31\pm0.77 tokens. The vocabulary sizes are 3,031 for whole-shape descriptions, 3,110 for part descriptions, and 4,522 overall (after lowercasing and stemming for vocabulary counts). Shapes have a mean of 3.63±1.283.63\pm1.28 annotated parts, and each part groups a mean of 1.93±1.201.93\pm1.20 puzzle pieces. For analysis, FULL uses 10–11 annotations per shape (mean 10.11), DENSE contains all annotations for the 74 densely annotated shapes, and DENSE10 contains only the sparse annotations for those same 74 shapes. For learning, splits are by tangram: 692 training shapes, 125 development shapes, 125 test shapes, and 74 test-dense shapes; all densely annotated shapes are in test-dense.

  4. Knowl 4 — Operational measures of naming divergence and segmentation agreement

    definition

    KILOGRAM measures whole-shape naming divergence (SND) and part naming divergence (PND) by comparing annotations of the same tangram. Let NN be the number of annotations, and let x(j)=(x1(j),…,xMj(j))x^{(j)}=(x^{(j)}_1,\ldots,x^{(j)}_{M_j}) be the token sequence for annotation jj, of length MjM_j. For token occurrence ii in annotation jj, define its weight as the proportion of the other annotations that do not contain that token. The divergence of annotation jj is the mean of these weights across its tokens; the tangram’s divergence is the mean across its NN annotations. SND uses each whole-shape description as x(j)x^{(j)}. PND applies the same calculation to the concatenation of all part names in each annotation.

    Part segmentation agreement (PSA) compares pairs of segmentations over the same tangram’s puzzle pieces. For each pair, construct a matrix whose rows and columns represent the parts in the two segmentations and whose entries count shared puzzle pieces. Find a maximum-weight one-to-one assignment between parts; the pair’s score is the resulting total number of matched pieces. PSA is the mean of these scores over all annotation pairs for the tangram. Thus, larger SND or PND indicates more variable naming, while larger PSA indicates greater overlap under the paper’s matching-based measure.

  5. Knowl 5 — Human annotations show high naming variability and related whole- and part-level consensus

    empirical result

    On the FULL set, mean SND is 0.91±0.110.91\pm0.11, mean PND is 0.76±0.190.76\pm0.19, and mean PSA is 5.30±0.625.30\pm0.62. Thus, whole-shape descriptions are more divergent on average than part names, while segmentation agreement varies across shapes. Only one tangram has perfect whole-shape naming consensus. In a manual sample of 250 annotations, 30.8% evoked human-like concepts, 31.2% other animate concepts, and 38.0% non-animate concepts.

    Whole-shape and part-name divergence are moderately positively correlated across tangrams (r(1014)=.531r(1014)=.531, p≪.001p\ll.001). Naming divergence is weakly negatively correlated with segmentation agreement: SND–PSA r(1014)=−.216r(1014)=-.216, p≪.001p\ll.001, and PND–PSA r(1014)=−.165r(1014)=-.165, p≪.001p\ll.001. For the 74 densely annotated shapes, rankings of the shapes’ measures closely track rankings from their sparse annotations: DENSE10 versus DENSE Spearman correlations are .78 for SND, .87 for PND, and .76 for PSA, all with p≪.001p\ll.001. The authors report that the selected dense shapes cover the observed range of these measures.

  6. Knowl 6 — Reference-game benchmark tests recognition of abstract shapes from descriptions

    model/method

    In the reference-game task, a model receives a textual description and a set of k=10k=10 tangram images, then selects the image with the highest text–image compatibility score. Evaluation games are sampled so that images do not repeat, distractors do not have identical whole-shape text annotations to the target, and all images have the same number of annotated parts. These restrictions remove duplicate targets and reduce shortcuts such as distinguishing images by part count.

    The WHOLE text condition uses only the whole-shape description. PARTS adds the part names using the template “<whole shape> with <part>, <part>, …, and <part>.” BLACK images show all pieces in one color; COLOR images color each annotated part differently, with colors ordered to correspond to part names in the text. The principal conditions are WHOLE+BLACK, PARTS+BLACK, WHOLE+COLOR, and PARTS+COLOR. In PARTS+COLOR, the text’s part names are color-coded to match the corresponding image parts. An AUG training variant generates examples for every subset of an annotation’s part names, coloring only the included parts and leaving the others black; part names are shuffled so color indicates their position in the text rather than a fixed semantic identity.

  7. Knowl 7 — CLIP and ViLT are evaluated both pretrained and fine-tuned on KILOGRAM

    experimental setup

    The benchmark compares CLIP, which separately encodes text and image and scores them by dot-product similarity, with ViLT, which jointly encodes text and image and uses its matching-classification head as the compatibility score. For fine-tuning, each reference game supplies a k×kk\times k text–image matching matrix, using a randomly sampled description for each distractor; a symmetric cross-entropy loss trains matching in both text-to-image and image-to-text directions. Reported results ensemble three models by element-wise multiplication of their outputs.

    CLIP uses ViT-B/32 and is fine-tuned with Adam, learning rate 5×10−85\times10^{-8}, weight decay 10−610^{-6}, for up to 200 epochs with patience 50. ViLT uses AdamW, learning rate 10−410^{-4}, weight decay 10−210^{-2}, a cosine schedule with one epoch of warm-up, and up to 30 epochs with patience 10. Both select checkpoints by image-prediction accuracy on a non-augmented validation set from the training data. The models have 151.2M trainable parameters for CLIP and 87.4M for ViLT.

  8. Knowl 8 — Fine-tuning sharply improves accuracy; ViLT excels when text and part colors align

    data/table

    The table reports reference-game accuracy (percent) for pretrained (PT) and fine-tuned (FT) models on development and held-out test games. Human performance was measured on development games only. Random selection among 10 images gives a 10% baseline. Fine-tuning substantially raises accuracy for both models; the largest results occur when part names and corresponding part colors are both available to ViLT.

    Condition CLIP PT CLIP FT ViLT PT ViLT FT Human
    Development accuracy (%)
    WHOLE+BLACK 16.1 43.3 12.9 40.9 47.7
    PARTS+BLACK 16.4 45.3 12.5 45.7 49.1
    WHOLE+COLOR 15.9 40.8 11.7 41.0 49.5
    PARTS+COLOR 15.0 45.4 10.7 75.2 63.0
    PARTS+COLOR+AUG – 47.6 – 72.2 –
    Held-out test accuracy (%)
    WHOLE+BLACK 17.9 42.5 13.1 44.5 –
    PARTS+BLACK 18.6 45.8 13.3 50.3 –
    WHOLE+COLOR 18.1 41.4 12.8 44.8 –
    PARTS+COLOR 17.0 46.5 11.7 77.3 –
    PARTS+COLOR+AUG – 50.2 – 74.4 –
  9. Knowl 9 — Part information helps most when linguistic labels correspond to visual segmentation

    empirical result

    On fine-tuned models, adding part names to whole-shape descriptions improves both systems relative to WHOLE+BLACK on the held-out test set: CLIP increases from 42.5% to 45.8%, and ViLT from 44.5% to 50.3%. Coloring parts without naming them (WHOLE+COLOR) provides no comparable benefit. Providing both part names and matching part colors (PARTS+COLOR) raises fine-tuned ViLT to 77.3% test accuracy, compared with 50.3% for PARTS+BLACK; fine-tuned CLIP reaches 46.5% in PARTS+COLOR, close to its 45.8% with PARTS+BLACK. In development analyses that gradually add aligned part information, the probability assigned to the correct image rises for both models but with diminishing returns. The returns diminish more quickly for CLIP; ViLT continues benefiting until roughly four parts are supplied.

    The AUG condition, which trains on every subset of an annotation’s part labels, improves CLIP relative to its non-augmented PARTS+COLOR results (50.2% versus 46.5% on test) but lowers ViLT’s result (74.4% versus 77.3%). The authors suggest, as a hypothesis rather than a demonstrated mechanism, that partial annotations may make part-to-name correspondence harder for ViLT to learn.

  10. Knowl 10 — Reference games and human baseline impose limits on interpretation

    limitation

    Descriptions in KILOGRAM were elicited for isolated tangrams, not in interaction with a particular set of distractors. They therefore may omit the context-specific detail a speaker would use to disambiguate a target, and randomly assembled games can contain multiple images that plausibly fit the same description. The attainable accuracy on these games may consequently be below 100%, and the benchmark does not measure the pragmatic coordination that occurs in interactive reference. Human baseline estimates are also preliminary: 217 participants completed 20 trials each, and participant accuracies showed distinct higher- and lower-performing groups in three of the four conditions. In PARTS+COLOR, the fitted groups had mean accuracies of 52.5% and 83.8%, respectively, making a single mean an incomplete summary of human performance. The study’s data and analyses are in English; other languages and cultures may support different abstractions.

Coverage note — The low-level tangram digitization heuristics and ancillary analyses of particular part-word distributions are omitted because they do not materially change the dataset’s core design, its variability measures, or the principal model-evaluation findings.

References

  1. 1.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015. VQA: Visual question answering. In IEEE International Conference on Computer Vision, pages 2425–2433.
  2. 2.Mark Atkinson, Gregory J Mills, and Kenny Smith. 2019. Social group effects on the emergence of communicative conventions and language complexity. Journal of Language Evolution, 4(1):1–18.
  3. 3.Adrian Bangerter, Eric Mayor, and Dominique Knutsen. 2020. Lexical entrainment without conceptual pacts? revisiting the matching task. Journal of Memory and Language, 114:104129.
  4. 4.Irving Biederman. 1987. Recognition-by-components: a theory of human image understanding. Psychological review, 94 2:115–147.
  5. 5.Steven Bird. 2004. Nltk: The natural language toolkit. ArXiv, cs.CL/0205028.
  6. 6.G. Bradski. 2000. The OpenCV Library. Dr. Dobb’s Journal of Software Tools.
  7. 7.Bailey Brashears and John Paul Minda. 2020. The effects of feature verbalizablity on category learning. In Proceedings of the 42nd Conference of the Cognitive Science Society.
  8. 8.Lucía Castillo, Kenny Smith, and Holly P Branigan. 2019. Interaction promotes the adaptation of referential conventions to the communicative context. Cognitive science, 43(8):e12780.
  9. 9.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick. 2015. Microsoft COCO captions: Data collection and evaluation server. CoRR, abs/1504.00325.
  10. 10.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. Uniter: Universal image-text representation learning. In European conference on computer vision, pages 104–120.
  11. 11.Herbert H Clark and Deanna Wilkes-Gibbs. 1986. Referring as a collaborative process. Cognition, 22(1):1–39.
  12. 12.Yael M. Cycowicz, D Friedman, M Rothstein, and Joan Gay Snodgrass. 1997. Picture naming by young children: norms for name agreement, familiarity, and visual complexity. Journal of experimental child psychology, 65 2:171–237.
  13. 13.Melissa C Duff, Julie Hengst, Daniel Tranel, and Neal J Cohen. 2006. Development of shared information in communication despite hippocampal amnesia. Nature neuroscience, 9(1):140–146.
  14. 14.Jon Andoni Duñabeitia, Davide Crepaldi, Antje S. Meyer, Boris New, Christos Pliatsikas, Eva Smolka, and Marc Brysbaert. 2018. Multipic: A standardized set of 750 drawings with norms for six european languages. Quarterly Journal of Experimental Psychology, 71:808 – 816.
  15. 15.Joost Elffers. 1977. Tangram: The Ancient Chinese Puzzle. Penguin Books.
  16. 16.Judith E. Fan, Daniel Yamins, and Nicholas B. Turk-Browne. 2015. Common object representations for visual recognition and production. Cognitive Science.
  17. 17.Alicia Fasquel, Angèle Brunellière, and Dominique Knutsen. 2022. A modified procedure for naming 332 pictures and collecting norms: Using tangram pictures in psycholinguistic studies. Behavior research methods.
  18. 18.Nicholas FitzGerald, Yoav Artzi, and Luke Zettlemoyer. 2013. Learning distributions over logical forms for referring expression generation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 1914–1925.
  19. 19.Jean E Fox Tree. 1999. Listening in on monologues and dialogues. Discourse processes, 27(1):35–53.
  20. 20.Dedre Gentner and Arthur B. Markman. 1997. Structure mapping in analogy and similarity. American Psychologist, 52:45–56.
  21. 21.Noah D. Goodman and Michael C. Frank. 2016. Pragmatic language interpretation as probabilistic inference. Trends in Cognitive Sciences, 20:818–829.
  22. 22.Chris Harris, Mike Stephens, et al. 1988. A combined corner and edge detector. In Alvey vision conference, volume 15, pages 10–5244. Citeseer.
  23. 23.Robert D. Hawkins, Michael C. Frank, and Noah D. Goodman. 2020. Characterizing the dynamics of learning in repeated reference games. Cognitive science, 44(6):e12845.
  24. 24.Judith Holler and Katie Wilkin. 2011. Co-speech gesture mimicry in the process of collaborative referring during face-to-face dialogue. Journal of Nonverbal Behavior, 35(2):133–153.
  25. 25.William S Horton and Richard J Gerrig. 2002. Speakers’ experiences and audience design: Knowing when and knowing how to adjust utterances to addressees. Journal of Memory and Language, 47(4):589–606.
  26. 26.William S Horton and Daniel G Slaten. 2012. Anticipating who will say what: The influence of speaker-specific memory associations on reference resolution. Memory & cognition, 40(1):113–126.
  27. 27.Michel Hupet, Xavier Seron, and Yves Chantraine. 1991. The effects of the codability and discriminability of the referents on the collaborative referring procedure. British Journal of Psychology, 82(4):449–462.
  28. 28.Alyssa Ibarra and Michael K Tanenhaus. 2016. The flexibility of conceptual pacts: Referring expressions dynamically shift to accommodate new conceptualizations. Frontiers in psychology, 7:561.
  29. 29.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning, pages 4904–4916. PMLR.
  30. 30.Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning, pages 5583–5594. PMLR.
  31. 31.Robert M Krauss and Sidney Weinheimer. 1964. Changes in reference phrases as a function of frequency of usage in social interaction: A preliminary study. Psychonomic Science, 1(1):113–114.
  32. 32.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. 2016. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International Journal of Computer Vision, 123:32–73.
  33. 33.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755.
  34. 34.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 32.
  35. 35.Gary Lupyan and Bodo Winter. 2018. Language is more abstract than you think, or, why aren’t languages more iconic? Philosophical Transactions of the Royal Society B: Biological Sciences, 373.
  36. 36.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L. Yuille, and Kevin Murphy. 2016. Generation and comprehension of unambiguous object descriptions. In Proceedings of the Conference on Computer Vision and Pattern Recognition, pages 11–20. IEEE.
  37. 37.William McCarthy, Robert X. D. Hawkins, Haoliang Wang, Cameron Holdaway, and Judith E. Fan. 2021. Learning to communicate about shared procedural abstractions. ArXiv, abs/2107.00077.
  38. 38.Douglas L. Medin, Robert L. Goldstone, and Dedre Gentner. 1993. Respects for similarity. Psychological Review, 100:254–278.
  39. 39.Margaret Mitchell, Kees van Deemter, and Ehud Reiter. 2010. Natural reference to objects in a visual domain. In Proceedings of the International Natural Language Generation Conference.
  40. 40.Tara Murfitt and Jan McAllister. 2001. The effect of production variables in monolog and dialog on comprehension by novel listeners. Language and Speech, 44(3):325–350.
  41. 41.Sonia K Murthy, Thomas L Griffiths, and Robert D Hawkins. 2022. Shades of confusion: Lexical uncertainty modulates ad hoc coordination in an interactive communication task. Cognition, 225:105152.
  42. 42.Vicente Ordonez, Girish Kulkarni, and Tamara Berg. 2011. Im2text: Describing images using 1 million captioned photographs. Advances in neural information processing systems, 24.
  43. 43.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR.
  44. 44.Michael F Schober and Herbert H Clark. 1989. Understanding by addressees and overhearers. Cognitive psychology, 21(2):211–232.
  45. 45.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565.
  46. 46.Roger N. Shepard. 1987. Toward a universal law of generalization for psychological science. Science, 237 4820:1317–23.
  47. 47.Todd Shore, Theofronia Androulakaki, and Gabriel Skantze. 2018. KTH tangrams: A dataset for research on alignment and conceptual pacts in task-oriented dialogue. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  48. 48.J Slocum. 2003. The Tangram Book: The Story of the Chinese Puzzle with over 2000 Puzzles to Solve. Sterling Publishing, New York.
  49. 49.Joan Gay Snodgrass and Mary Vanderwart. 1980. A standardized set of 260 pictures: norms for name agreement, image agreement, familiarity, and visual complexity. Journal of experimental psychology. Human learning and memory, 6 2:174–215.
  50. 50.Alane Suhr, Mike Lewis, James Yeh, and Yoav Artzi. 2017. A corpus of natural language for visual reasoning. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 217–223.
  51. 51.Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. 2019. A corpus for reasoning about natural language grounded in photographs. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 6418–6428.
  52. 52.Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross. 2022. Winoground: Probing vision and language models for visio-linguistic compositionality. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5238–5248.
  53. 53.Barbara Tversky and Kathleen Hemenway. 1984. Objects, parts, and categories. Journal of Experimental Psychology: General, 113:169–193.
  54. 54.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems.
  55. 55.Jette Viethen and Robert Dale. 2008. The use of spatial relations in referring expression generation. In Proceedings of the International Conference on Natural Language Generation.
  56. 56.Catherine Wong, William McCarthy, Gabriel Grand, Yoni Friedman, Joshua B. Tenenbaum, Jacob Andreas, Robert D. Hawkins, and Judith E. Fan. 2022. Identifying concept libraries from language about object structure. ArXiv, abs/2205.05666.
  57. 57.Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. 2016. Modeling context in referring expressions. In The European Conference on Computer Vision, pages 69–85.
  58. 58.Martin Zettersten and Gary Lupyan. 2020. Finding categories through words: More nameable features improve category learning. Cognition, 196:104135.

Citation

MLA
Ji, A., et al. “Abstract Visual Reasoning with Tangram Shapes”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 582–601, https://doi.org/10.18653/v1/2022.emnlp-main.38.
APA
Ji, A., Kojima, N., Rush, N., Suhr, A., Vong, W. K., Hawkins, R. D., & Artzi, Y. (2022). Abstract Visual Reasoning with Tangram Shapes. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 582–601. https://doi.org/10.18653/v1/2022.emnlp-main.38
Chicago
Ji, A., N. Kojima, N. Rush, et al. 2022. “Abstract Visual Reasoning with Tangram Shapes”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 582–601. https://doi.org/10.18653/v1/2022.emnlp-main.38.
Harvard
Ji, A. et al. (2022) “Abstract Visual Reasoning with Tangram Shapes”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 582–601. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.38.
Vancouver
1. Ji A, Kojima N, Rush N, Suhr A, Vong WK, Hawkins RD, Artzi Y (2022) Abstract Visual Reasoning with Tangram Shapes. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 582–601

BibTeX

@inproceedings{ji-etal-2022-abstract,
    title = "Abstract Visual Reasoning with Tangram Shapes",
    author = "Ji, Anya  and
      Kojima, Noriyuki  and
      Rush, Noah  and
      Suhr, Alane  and
      Vong, Wai Keen  and
      Hawkins, Robert  and
      Artzi, Yoav",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.38/",
    doi = "10.18653/v1/2022.emnlp-main.38",
    pages = "582--601"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/