Disentangling visual and written concepts in CLIP
Joanna MaterzynskaAntonio TorralbaDavid Bau
Proposes an orthogonal projection method to separate text-reading capabilities from visual object processing in CLIP's image encoder, effectively eliminating text artifacts in guided image generation and defending against typographic attacks.
Modern vision-language artificial intelligence models, such as CLIP, frequently confuse written text with visual objects. For example, placing a handwritten label on an object can cause a model to misclassify it entirely, and generating images from text prompts often introduces unwanted, rendered words into the generated picture. This vulnerability arises because real-world training data regularly pairs objects directly with their written labels on signs and packaging, creating deeply entangled internal representations. The article evaluates whether a model's understanding of written text can be separated from its perception of natural visual concepts, demonstrating a method to isolate or eliminate reading capabilities within the visual encoder.
To achieve this, the authors developed a linear projection method that isolates distinct visual and text subspaces using an orthogonality constraint. They created a dataset combining natural images, descriptive class labels, synthetic images of rendered text, and natural images overlaid with text, including both real English vocabulary and nonsense character strings. They trained two distinct linear transformations: a "learn to spell" model designed to isolate text-reading capabilities into a compact subspace, and a "forget to spell" model configured to suppress written text and retain purely visual representations. The approach was validated through cross-modal retrieval benchmarks, text-conditioned generative image synthesis, and classification tests against typographic adversarial attacks.
Key findings show that text processing can be successfully disentangled from general visual comprehension. First, the article demonstrates that text-reading capabilities can be compressed into a subspace as small as 64 dimensions while maintaining strong retrieval performance, achieving up to 90.3% retrieval accuracy on paired image-text tasks. Second, the "forget to spell" model reduced text detection in generated images by roughly 55% compared to the "learn to spell" model, significantly cleaning generative visual outputs. Third, in robustness evaluations against typographic attacks, the "forget to spell" projection increased classification accuracy from 49.4% in the baseline model to 77.2% on true object labels by suppressing misleading text overlays. Finally, experiments revealed that maintaining strict orthogonality during training was critical; omitting this constraint led to performance drops of up to 24% and caused generative image processes to collapse.
These findings indicate that foundational vision models do not require full retraining to address text-bias vulnerabilities and generative text artifacts. Instead, lightweight mathematical projections applied to pre-trained embeddings offer a computationally inexpensive and effective defense against typographic confusion and adversarial vulnerabilities. Organizations deploying vision-language systems can adopt these projection techniques to enhance model robustness and improve generative quality without incurring the high costs and extended timelines of training new models from scratch. Future work should focus on extending these projections to broader vocabularies and evaluating their performance across newer multimodal architectures.
While highly effective, the approach has notable limitations. The projections do not achieve total separation: the "forget to spell" model occasionally leaves faint, text-like textures in generated images, and the "learn to spell" model does not guarantee flawless rendering across all text prompts. Nevertheless, the consistent performance across both synthetic benchmarks and natural image evaluations provides strong confidence in the viability of subspace projection for disentangling visual and written concepts.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. Introduces the Contrastive Language-Image Pre-training (CLIP) architecture whose visual and written concept representations are directly analyzed and disentangled in the source.
- Paper: StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery, Or Patashnik et al. (2021). Demonstrates latent direction and subspace manipulation in CLIP embeddings to steer generative visual features without full retraining.
- Paper: Zero-Shot Text-to-Image Generation, Aditya Ramesh et al. (2021). Establishes foundational techniques for zero-shot text-to-image generation and CLIP-guided evaluation that motivate the study of typographic artifacts in multimodal synthesis.
- Paper: FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts, Yichen Gong et al. (2025). Exploits the typographic vulnerability of vision-language models identified and mitigated by the source to execute adversarial visual jailbreaks.
- Paper: ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval, Mengjun Cheng et al. (2022). Builds upon the interaction between visual appearance and embedded scene text by designing a unified dual-encoder framework for cross-modal retrieval.
- Paper: TextDiffuser: Diffusion Models as Text Painters, Jingye Chen et al. (2023). Addresses the inverse goal of rendering coherent written text within generated images using explicit layout guidance and character-aware diffusion objectives.
- Paper: On the Origins of Linear Representations in Large Language Models, Yibo Jiang et al. (2024). Provides formal theoretical foundations explaining why linear and orthogonal concept subspaces, such as those projected in the source, naturally emerge in representation learning.
- Paper: Improving CLIP Training with Language Rewrites, Lijie Fan et al. (2023). Extends the study of language overfitting and shortcut learning in CLIP by introducing automated text rewrites to improve representation robustness.
