TIPS: Text-Image Pretraining with Spatial awareness

Kevis-Kokitsi ManinisKaifeng ChenSoham GhoshArjun KarpurKoert ChenYe XiaBingyi CaoDaniel SalzGuangxing HanJan Dlabal

article2025ICLR41 citations

Introduces a spatially aware vision-language pretraining method that unites synthetic captions with masked image modeling, allowing general image-text models to match specialized self-supervised representations on dense tasks like segmentation and depth estimation.

Listen

Modern computer vision increasingly relies on large foundation models that can be deployed off the shelf across diverse applications without expensive task-specific fine-tuning. However, current representation learning approaches suffer from a persistent trade-off. Multimodal vision-language models, such as standard image-text encoders, excel at holistic classification and cross-modal retrieval by aligning images with web text, but they perform poorly on dense, spatially grounded tasks like depth estimation and object segmentation. Conversely, self-supervised vision models produce representations with strong spatial coherence but lack language alignment, preventing their direct use in vision-language applications. The article addresses this operational gap by introducing Text-Image Pretraining with Spatial awareness (TIPS), an approach designed to deliver unified representations that excel at both holistic multimodal tasks and dense visual predictions.

The core objective of the article is to demonstrate that integrating enhanced textual descriptions with self-supervised spatial objectives produces an image-text model that matches or exceeds self-supervised systems on dense tasks while maintaining state-of-the-art vision-language capabilities. To achieve this, the authors developed a dual-embedding architecture powered by two complementary mechanisms. First, to overcome the noise and spatial ambiguity of standard web captions, the framework augments training data with synthetically generated captions that explicitly describe object arrangements and background context, assigning separate transformer tokens to web and synthetic text. Second, the framework incorporates self-supervised self-distillation and masked image modeling losses during contrastive training to enforce spatial consistency across image patches. The authors validated this methodology by scaling a Vision Transformer model with 1.1 billion parameters on a curated dataset of roughly 117 million image-text pairs, evaluating frozen features across 8 distinct vision tasks spanning 16 benchmark datasets.

The evaluation demonstrates that TIPS establishes new performance benchmarks across both dense and multimodal tasks. In dense prediction benchmarks, TIPS achieved a semantic segmentation score of 83.6 mean Intersection over Union on PASCAL VOC and reduced monocular depth estimation error to 0.353 root mean squared error on NYUv2, matching or surpassing leading self-supervised systems like DINOv2 while substantially outperforming existing weakly supervised baselines. In multimodal benchmarks, the model secured top performance in 6 out of 7 image-text retrieval evaluations, including an image-to-text recall@1 of 74.0 on COCO and 93.8 on Flickr30K. In neural 3D reconstruction from single images, replacing standard baseline representations with TIPS features improved rendering quality by 0.62 decibels in peak signal-to-noise ratio. Furthermore, knowledge distillation experiments showed that compressed student variants—ranging down to compact small and base models—retained strong capabilities, with a 487-million-parameter variant achieving performance comparable to the primary 1.1-billion-parameter teacher.

These findings indicate that organizations no longer need to maintain separate vision backbones for spatial perception and multimodal reasoning. Consolidating these workloads into a single, off-the-shelf encoder architecture reduces deployment overhead, simplifies infrastructure maintenance, and lowers the operational costs associated with downstream fine-tuning. The results also show that synthetic captioning effectively solves the spatial supervision bottleneck in web-scraped data, challenging the assumption that dense vision tasks inherently require purely self-supervised or densely annotated training pipelines.

Based on these results, engineering and product teams should consider deploying the released TIPS models for pipelines requiring unified spatial understanding and multimodal search, particularly via the distilled lightweight variants to balance inference efficiency and accuracy. When retraining or adapting vision encoders, teams should incorporate synthetic, spatially descriptive captions alongside web data and combine masked modeling with contrastive objectives. Future development should explore expanding synthetic captioning to non-English datasets, testing the architecture under real-time compute constraints, and scaling spatial image-text pretraining to broader video and interactive robotics domains.

No sufficiently relevant recommendations were found.

Cover for TIPS: Text-Image Pretraining with Spatial awareness

Abstract

While image-text representation learning has become very popular in recent years, existing models tend to lack spatial awareness and have limited direct applicability for dense understanding tasks. For this reason, self-supervised image-only pretraining is still the go-to method for many dense vision applications (e.g. depth estimation, semantic segmentation), despite the lack of explicit supervisory signals. In this paper, we close this gap between image-text and self-supervised learning, by proposing a novel general-purpose image-text model, which can be effectively used off the shelf for dense and global vision tasks. Our method, which we refer to as Text-Image Pretraining with Spatial awareness (TIPS), leverages two simple and effective insights. First, on textual supervision: we reveal that replacing noisy web image captions by synthetically generated textual descriptions boosts dense understanding performance significantly, due to a much richer signal for learning spatially aware representations. We propose an adapted training method that combines noisy and synthetic captions, resulting in improvements across both dense and global understanding tasks. Second, on the learning technique: we propose to combine contrastive image-text learning with self-supervised masked image modeling, to encourage spatial coherence, unlocking substantial enhancements for downstream applications. Building on these two ideas, we scale our model using the transformer architecture, trained on a curated set of public images. Our experiments are conducted on 8 tasks involving 16 datasets in total, demonstrating strong off-the-shelf performance on both dense and global understanding, for several image-only and image-text tasks. Code and models are released at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 TIPS
  • 3.1 Enhancing Weak Supervision with Synthetic Image Captions
  • 3.2 Integrating Self-Distillation and Masking to Boost Image Features
  • 3.3 Scaling TIPS
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Results
  • 5 Conclusions
  • References
  • A Appendix
  • A.1 Additional experimental results
  • A.2 Additional implementation details
  • A.3 Dataset curation
  • A.4 Detailed evaluation protocols
  • A.4.1 Dense Image Tasks
  • A.4.2 Global Image Tasks
  • A.4.3 Multimodal Retrieval Tasks
  • A.4.4 3D Vision Tasks.
  • A.5 Additional Qualitative Results
  • A.6 Dual embedding attention maps

Citation

MLA
Maninis, K.-K., et al. “TIPS: Text-Image Pretraining with Spatial Awareness”. arXiv, 2024, http://arxiv.org/abs/2410.16512v2.
APA
Maninis, K.-K., Chen, K., Ghosh, S., Karpur, A., Chen, K., Xia, Y., Cao, B., Salz, D., Han, G., Dlabal, J., Gnanapragasam, D., Seyedhosseini, M., Zhou, H., & Araujo, A. (2024). TIPS: Text-Image Pretraining with Spatial awareness. arXiv. http://arxiv.org/abs/2410.16512v2
Chicago
Maninis, K.-K., K. Chen, S. Ghosh, et al. 2024. “TIPS: Text-Image Pretraining with Spatial Awareness”. arXiv. http://arxiv.org/abs/2410.16512v2.
Harvard
Maninis, K.-K. et al. (2024) “TIPS: Text-Image Pretraining with Spatial awareness”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.16512v2.
Vancouver
1. Maninis K-K, Chen K, Ghosh S, et al (2024) TIPS: Text-Image Pretraining with Spatial awareness. arXiv

BibTeX

@article{maninis2024tips,
  title = {TIPS: Text-Image Pretraining with Spatial awareness},
  author = {Maninis, Kevis-Kokitsi and Chen, Kaifeng and Ghosh, Soham and Karpur, Arjun and Chen, Koert and Xia, Ye and Cao, Bingyi and Salz, Daniel and Han, Guangxing and Dlabal, Jan and Gnanapragasam, Dan and Seyedhosseini, Mojtaba and Zhou, Howard and Araujo, Andre},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.16512v2},
  eprint = {2410.16512}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/