Disentangling visual and written concepts in CLIP

Joanna MaterzynskaAntonio TorralbaDavid Bau

article2022CVPR65 citations

Proposes an orthogonal projection method to separate text-reading capabilities from visual object processing in CLIP's image encoder, effectively eliminating text artifacts in guided image generation and defending against typographic attacks.

Listen

Modern vision-language artificial intelligence models, such as CLIP, frequently confuse written text with visual objects. For example, placing a handwritten label on an object can cause a model to misclassify it entirely, and generating images from text prompts often introduces unwanted, rendered words into the generated picture. This vulnerability arises because real-world training data regularly pairs objects directly with their written labels on signs and packaging, creating deeply entangled internal representations. The article evaluates whether a model's understanding of written text can be separated from its perception of natural visual concepts, demonstrating a method to isolate or eliminate reading capabilities within the visual encoder.

To achieve this, the authors developed a linear projection method that isolates distinct visual and text subspaces using an orthogonality constraint. They created a dataset combining natural images, descriptive class labels, synthetic images of rendered text, and natural images overlaid with text, including both real English vocabulary and nonsense character strings. They trained two distinct linear transformations: a "learn to spell" model designed to isolate text-reading capabilities into a compact subspace, and a "forget to spell" model configured to suppress written text and retain purely visual representations. The approach was validated through cross-modal retrieval benchmarks, text-conditioned generative image synthesis, and classification tests against typographic adversarial attacks.

Key findings show that text processing can be successfully disentangled from general visual comprehension. First, the article demonstrates that text-reading capabilities can be compressed into a subspace as small as 64 dimensions while maintaining strong retrieval performance, achieving up to 90.3% retrieval accuracy on paired image-text tasks. Second, the "forget to spell" model reduced text detection in generated images by roughly 55% compared to the "learn to spell" model, significantly cleaning generative visual outputs. Third, in robustness evaluations against typographic attacks, the "forget to spell" projection increased classification accuracy from 49.4% in the baseline model to 77.2% on true object labels by suppressing misleading text overlays. Finally, experiments revealed that maintaining strict orthogonality during training was critical; omitting this constraint led to performance drops of up to 24% and caused generative image processes to collapse.

These findings indicate that foundational vision models do not require full retraining to address text-bias vulnerabilities and generative text artifacts. Instead, lightweight mathematical projections applied to pre-trained embeddings offer a computationally inexpensive and effective defense against typographic confusion and adversarial vulnerabilities. Organizations deploying vision-language systems can adopt these projection techniques to enhance model robustness and improve generative quality without incurring the high costs and extended timelines of training new models from scratch. Future work should focus on extending these projections to broader vocabularies and evaluating their performance across newer multimodal architectures.

While highly effective, the approach has notable limitations. The projections do not achieve total separation: the "forget to spell" model occasionally leaves faint, text-like textures in generated images, and the "learn to spell" model does not guarantee flawless rendering across all text prompts. Nevertheless, the consistent performance across both synthetic benchmarks and natural image evaluations provides strong confidence in the viability of subspace projection for disentangling visual and written concepts.

arXiv: 2206.07835
Cover for Disentangling visual and written concepts in CLIP

Abstract

The CLIP network measures the similarity between natural text and images; in this work, we investigate the entanglement of the representation of word images and natural images in its image encoder. First, we find that the image encoder has an ability to match word images with natural images of scenes described by those words. This is consistent with previous research that suggests that the meaning and the spelling of a word might be entangled deep within the network. On the other hand, we also find that CLIP has a strong ability to match nonsense words, suggesting that processing of letters is separated from processing of their meaning. To explicitly determine whether the spelling capability of CLIP is separable, we devise a procedure for identifying representation subspaces that selectively isolate or eliminate spelling capabilities. We benchmark our methods against a range of retrieval tasks, and we also test them by measuring the appearance of text in CLIP-guided generated images. We find that our methods are able to cleanly separate spelling capabilities of CLIP from the visual processing of natural images.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Terminology
  • 4. Visual comprehension
  • 5. Disentangling Text and Vision with Linear Projections
  • 6. Experiments
  • 7. Evaluation
  • 7.1. Text Generation
  • 7.2. Robustness
  • 8. Limitations
  • 9. Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Linear Subspace Projection for Disentangling Written and Visual Concepts in CLIP

    model/method

    To separate the processing of written text from visual semantics within a frozen CLIP embedding space Rd\mathbb{R}^d (d=512d = 512), a linear projection matrix W∈Rd×kW \in \mathbb{R}^{d \times k} is learned, where k≤dk \le d denotes a lower bottleneck dimension. Given precomputed CLIP feature vectors, projected representations are obtained via linear transformation v′=WTvv' = W^T v.

    The projection is trained on 5-tuples (xi,yi,xt,yt,xit)(x_i, y_i, x_t, y_t, x_{it}), where:

    • xi∈Rdx_i \in \mathbb{R}^d is the CLIP image embedding of a natural image.
    • yi∈Rdy_i \in \mathbb{R}^d is the CLIP text embedding of the class label prompt (e.g., "an image of a [class]").
    • xt∈Rdx_t \in \mathbb{R}^d is the CLIP image embedding of a synthetic image containing rendered text on a white background.
    • yt∈Rdy_t \in \mathbb{R}^d is the CLIP text embedding of the corresponding text string.
    • xit∈Rdx_{it} \in \mathbb{R}^d is the CLIP image embedding of the natural image xix_i with the text string from xtx_t rendered directly onto it.

    To preserve the geometric structure of the original representation and prevent representation collapse, training enforces an orthogonality regularization term:

    R(W)=∥I−WWT∥R(W) = \|I - WW^T\|

    where II is the identity matrix and ∥⋅∥\|\cdot\| is the matrix norm.

  2. Knowl 2 — Contrastive Loss Objectives for Learn-to-Spell and Forget-to-Spell Projections

    equation

    The projection matrix WW is optimized using symmetric cross-entropy contrastive losses LkL_k applied to pairs of projected embeddings, combined with an orthogonality penalty R(W)=∥I−WWT∥R(W) = \|I - WW^T\| scaled by hyperparameter γ=0.5\gamma = 0.5.

    Six pairwise contrastive objectives are defined over the 5-tuple embeddings (xi,yi,xt,yt,xit)(x_i, y_i, x_t, y_t, x_{it}):

    • L1L_1: alignment between natural image xix_i and true class text yiy_i.
    • L2L_2: alignment between text-modified image xitx_{it} and true class text yiy_i.
    • L3L_3: alignment between synthetic text image xtx_t and text string yty_t.
    • L4L_4: alignment between text-modified image xitx_{it} and synthetic text image xtx_t.
    • L5L_5: alignment between text-modified image xitx_{it} and text string yty_t.
    • L6L_6: alignment between text-modified image xitx_{it} and original natural image xix_i.

    The general theoretical objectives are:

    Lspell=−L1−L2−L6+L3+L4+L5+γR(W)L_{\text{spell}} = -L_1 - L_2 - L_6 + L_3 + L_4 + L_5 + \gamma R(W)

    Lforget=L1+L2+L6−L3−L4−L5+γR(W)L_{\text{forget}} = L_1 + L_2 + L_6 - L_3 - L_4 - L_5 + \gamma R(W)

    Empirical ablation shows that the optimal configuration for isolating text capabilities ("learn to spell", bottleneck dimension k=64k = 64) minimizes L1+L3+L4+L5+γR(W)L_1 + L_3 + L_4 + L_5 + \gamma R(W). The optimal configuration for suppressing text processing ("forget to spell", bottleneck dimension k=256k = 256) optimizes L1+L2+L5+L6+γR(W)L_1 + L_2 + L_5 + L_6 + \gamma R(W) to prioritize visual concept matching over text-string alignment.

  3. Knowl 3 — Defense Against Typographic Attacks Using the Forget-to-Spell Projection

    empirical result

    The "forget to spell" projection provides defense against typographic adversarial attacks in zero-shot classification without fine-tuning the base CLIP model.

    In an evaluation dataset of 180 images comprising 20 distinct object categories subjected to 8 typographic attacks (natural images of objects with conflicting text labels visibly attached, such as an apple labeled with the text "iPad"):

    • Standard CLIP ViT-B/32 achieves a top-1 classification accuracy of only 49.4% on the true object labels because its image encoder strongly aligns with the written attack text.
    • The 256-dimensional "forget to spell" projection increases top-1 accuracy on the true object labels to 77.2%.

    The projection suppresses sensitivity to typographic attack text while preserving sensitivity to visual object features, demonstrating out-of-domain transfer from synthetic training overlays to physical natural text attacks.

  4. Knowl 4 — Controlling Text Artifacts in CLIP-Guided Generative Image Synthesis

    empirical result

    When generating images with VQGAN guided by CLIP text-image similarity, prompts frequently cause the generator to synthesize legible text strings of the prompt as visual artifacts.

    Replacing standard CLIP embeddings with projected representations controls this behavior:

    • The 64-dimensional "learn to spell" projection amplifies the rendering of legible characters and words corresponding to the prompt.
    • The 256-dimensional "forget to spell" projection removes rendered Latin letters, producing purely visual scene content.

    Text detection evaluated with EasyOCR (counting detections covering ≥10%\ge 10\% of image area with ≥2\ge 2 matching letters from prompt across 1,000 real-word prompts and 1,000 nonsense-string prompts) shows:

    • The detection rate in images generated with "learn to spell" exceeds standard CLIP by 25.43%.
    • The difference in text detection prevalence between "learn to spell" and "forget to spell" is 54.92% across all prompts.
  5. Knowl 5 — Intrinsic OCR and Cross-Modal Image-to-Image Matching in CLIP

    empirical result

    The frozen image encoder of CLIP ViT-B/32 possesses intrinsic OCR and semantic word-to-image association capabilities without utilizing the text encoder:

    1. Direct visual comprehension: When tasked with matching synthetic images of written category names (e.g., an image of the word "playground") against natural images of the corresponding visual classes using only image embeddings, CLIP achieves:
    • Places 365: 15.58% top-1 accuracy (random baseline: 0.10%; standard zero-shot image-to-text with prompt engineering: 39.47%).
    • ImageNet: 10.58% top-1 accuracy (random baseline: 0.27%; standard zero-shot image-to-text with prompt engineering: 63.36%).
    1. Nonsense string matching: Evaluating cross-modal retrieval between synthetic text images and text strings (1-out-of-20,000 retrieval pool):
    • Real English words (20,000 pairs): 76.38% Image-to-Text retrieval, 91.46% Text-to-Image retrieval.
    • Nonsense strings of 3–8 uniformly sampled Latin characters (20,000 pairs): 61.77% Image-to-Text retrieval, 79.19% Text-to-Image retrieval.

    This shows that CLIP's image encoder parses individual letter sequences independently of natural semantic training priors.

  6. Knowl 6 — Scene Text Recognition Benchmark on IIIT5K Across Projection Dimensions

    data/table

    Evaluating the "learn to spell" projection on the out-of-domain IIIT5K scene text recognition dataset demonstrates that a 128-dimensional regularized orthogonal subspace preserves full OCR performance compared to the original 512-dimensional CLIP embedding, while unregularized projections suffer severe degradation.

    Model Dimension Regularized IIIT5K 1K Accuracy (%) IIIT5K Full (1772) Accuracy (%)
    CLIP 512 - 69.43 63.00
    Learn to spell 128 ✓ 67.67 63.20
    Learn to spell 128 × 45.56 39.23
    Learn to spell 64 ✓ 64.56 61.17
    Learn to spell 64 × 44.80 39.00

    The regularized 128-dimensional projection achieves 63.20% on the full 1,772-word retrieval task (+0.20% over unprojected CLIP). Omitting the orthogonality constraint ∥I−WWT∥\|I - WW^T\| leads to a 23.97% drop in full-dataset accuracy at dimension 128 (down to 39.23%) and a 22.17% drop at dimension 64 (down to 39.00%).

  7. Knowl 7 — Necessity of Orthogonality Regularization for Subspace Stability

    empirical result

    Enforcing the orthogonality regularization term R(W)=∥I−WWT∥R(W) = \|I - WW^T\| with weight γ=0.5\gamma = 0.5 is necessary to maintain meaningful representations in linear projections of CLIP embeddings:

    • In the "learn to spell" model, removing R(W)R(W) causes a ∼10%\sim 10\% drop on validation image-text retrieval (falling from 90.30% to 76.78% on real word image-to-text retrieval, and from 84.77% to 74.38% on fake word retrieval). In text-to-image steering, the orthogonal model yields 17.5% more text detections than its unregularized counterpart.
    • In the "forget to spell" model, omitting R(W)R(W) results in catastrophic representation collapse: retrieval accuracy on natural images versus text-modified natural images drops from 97.84% to 23.48%, and VQGAN image generation collapses entirely to a uniform red background pattern.
  8. Knowl 8 — Synthetic and Natural Image-Text Tuple Dataset Construction

    experimental setup

    Projection matrices are trained using 5-tuples derived from ImageNet and an English word dictionary without updating the underlying CLIP ViT-B/32 backbone:

    • Vocabulary: A corpus of 202,587 lowercase English words of lengths 3–10 letters (182,329 train, 20,258 validation).
    • Nonsense strings: 50% of the dataset tuples use fake words generated by uniformly sampling a string length between 3 and 10 and randomly sampling characters from the Latin alphabet.
    • Data elements: For each natural image xix_i and class label yiy_i ("an image of [class label]"), a synthetic text image xtx_t is rendered as black text on a white background, and a modified image xitx_{it} is generated by rendering the text string onto xix_i.
    • Optimization: Projection matrix WW is trained for 1 epoch using the Adam optimizer with batch size 128, initial learning rate 10−410^{-4}, and a learning rate decay factor of 0.5 every 4,000 steps.
  9. Knowl 9 — Incomplete Elimination and Formation of Text in Projected Generative Outputs

    limitation

    The linear projection framework cannot achieve complete separation of text and visual semantics in generative tasks:

    • The "forget to spell" projection cannot perfectly prevent all typographic artifacts during image synthesis; residual text-like textures (frequently resembling non-Latin scripts or fragmented letter strokes) occasionally persist in the generated outputs.
    • The "learn to spell" projection cannot guarantee fully formed, legible, or orthographically correct typography for all text prompts, frequently producing partially distorted or incomplete letterforms.

Coverage note — None. All substantial contributed methods, formal formulations, empirical analyses (visual comprehension, controllable generative steering, typographic attack defense, out-of-domain OCR generalization), and limitations are included.

References

  1. 1.Guillaume Alain and Yoshua Bengio. Understanding intermediate layers using linear classifier probes. In ICLR Workshop, 2016. 2
  2. 2.David Bau, Alex Andonian, Audrey Cui, YeonHwan Park, Ali Jahanian, Aude Oliva, and Antonio Torralba. Paint by word. arXiv preprint arXiv:2103.10951, 2021. 2
  3. 3.Edo Collins, Raja Bala, Bob Price, and Sabine Süsstrunk. Editing in style: Uncovering the local semantics of gans. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5771–5780, 2020. 2
  4. 4.Katherine Crowson. VQGAN+CLIP. https://colab.research.google.com/drive/15UwYDsnNeldJFHJ9NdgYBYeo6xPmSelP, Jan. 2021. 2
  5. 5.Katherine Crowson. VQGAN+pooling. https://colab.research.google.com/drive/1ZAus_gn2RhTZWzOWUpPERNC0Q8OhZRTZ, Jan. 2021. 2, 6
  6. 6.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12873–12883, 2021. 2, 6
  7. 7.Ruth Fong and Andrea Vedaldi. Net2vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8730–8738, 2018. 2
  8. 8.Lore Goetschalckx, Alex Andonian, Aude Oliva, and Phillip Isola. Ganalyze: Toward visual definitions of cognitive image properties. In CVPR, pages 5744–5753, 2019. 2
  9. 9.Gabriel Goh, Nick Cammarata, Chelsea Voss, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah. Multimodal neurons in artificial neural networks. Distill, 6(3):e30, 2021. 1
  10. 10.Ankush Gupta, Andrea Vedaldi, and Andrew Zisserman. Synthetic data for text localisation in natural images. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2315–2324, 2016. 7
  11. 11.Erik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, and Sylvain Paris. Ganspace: Discovering interpretable gan controls. arXiv preprint arXiv:2004.02546, 2020. 2
  12. 12.Ali Jahanian, Lucy Chai, and Phillip Isola. On the "steerability" of generative adversarial networks. In ICLR, 2020. 2
  13. 13.JaidedAI. EasyOCR. https://github.com/JaidedAI/EasyOCR, 2021. 7
  14. 14.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8110–8119, 2020. 2
  15. 15.Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav). In International conference on machine learning, pages 2668–2677. PMLR, 2018. 2
  16. 16.Yoann Lemesle, Masataka Sawayama, Guillermo Valle-Perez, Maxime Adolphe, Hélène Sauzéon, and Pierre-Yves Oudeyer. Language-biased image classification: Evaluation based on semantic compositionality. In International Conference on Learning Representations, 2022. 1, 2
  17. 17.Anand Mishra, Karteek Alahari, and CV Jawahar. Scene text recognition using higher order language priors. In BMVC-British Machine Vision Conference. BMVA, 2012. 7
  18. 18.Ryan Murdock. The Big Sleep. https://colab.research.google.com/drive/1NCceX2mbiKOSlAd_o7IU7nA9UskKN5WR, Jan. 2021. 2
  19. 19.Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. Styleclip: Text-driven manipulation of stylegan imagery. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2085–2094, 2021. 2
  20. 20.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. arXiv preprint arXiv:2103.00020, 2021. 1, 2, 4
  21. 21.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. arXiv preprint arXiv:2102.12092, 2021. 2
  22. 22.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015. 3
  23. 23.Yujun Shen, Ceyuan Yang, Xiaoou Tang, and Bolei Zhou. Interfacegan: Interpreting the disentangled face representation learned by gans. IEEE transactions on pattern analysis and machine intelligence, 2020. 2
  24. 24.Zongze Wu, Dani Lischinski, and Eli Shechtman. Stylespace analysis: Disentangled controls for stylegan image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12863–12872, 2021. 2
  25. 25.Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 million image database for scene recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017. 3
  26. 26.Bolei Zhou, Yiyou Sun, David Bau, and Antonio Torralba. Interpretable basis decomposition for visual explanation. In Proceedings of the European Conference on Computer Vision (ECCV), pages 119–134, 2018. 2

Citation

MLA
Materzynska, J., et al. “Disentangling Visual and Written Concepts in CLIP”. arXiv, 2022, http://arxiv.org/abs/2206.07835v1.
APA
Materzynska, J., Torralba, A., & Bau, D. (2022). Disentangling visual and written concepts in CLIP. arXiv. http://arxiv.org/abs/2206.07835v1
Chicago
Materzynska, J., A. Torralba, and D. Bau. 2022. “Disentangling Visual and Written Concepts in CLIP”. arXiv. http://arxiv.org/abs/2206.07835v1.
Harvard
Materzynska, J., Torralba, A. and Bau, D. (2022) “Disentangling visual and written concepts in CLIP”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2206.07835v1.
Vancouver
1. Materzynska J, Torralba A, Bau D (2022) Disentangling visual and written concepts in CLIP. arXiv

BibTeX

@article{materzynska2022disentangling,
  title = {Disentangling visual and written concepts in CLIP},
  author = {Materzynska, Joanna and Torralba, Antonio and Bau, David},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2206.07835v1},
  eprint = {2206.07835}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE