Text-To-Concept (and Back) via Cross-Model Alignment

Mazda MoayeriKeivan RezaeiMaziar SanjabiSoheil Feizi

article2023ICML76 citations

Demonstrates that a simple linear layer can align arbitrary vision encoders with CLIP's feature space, enabling zero-shot classification, unsupervised concept bottleneck models, and two-way translation between internal representations and natural language.

Listen

Deep computer vision models create complex, high-dimensional representations that are rich in semantic information but difficult for humans to interpret directly. Connecting these internal features to human language typically requires massive multimodal datasets, expensive concept annotations, or retraining large networks from scratch. Consequently, smaller, specialized vision models deployed across organizations often operate as opaque black boxes whose rich latent capabilities remain underutilized.

The article demonstrates that diverse computer vision models represent images in fundamentally similar ways and can be aligned using a single, computationally inexpensive linear layer. Its primary objective is to show that linearly mapping off-the-shelf vision encoders to a multimodal CLIP space enables bidirectional communication between models and human language—converting text descriptions into concept vectors within an image model and translating internal model vectors back into readable text.

The researchers evaluated this framework using multiple standard architectures, including standard ResNets, robust ResNets, and Vision Transformers, trained with both supervised and self-supervised methods on the ImageNet dataset. They learned simple affine transformations to align these fixed, single-modality models to CLIP without altering the underlying models or requiring paired image-text training. The credibility of the approach was tested across several downstream tasks, including zero-shot classification, concept-based image retrieval, dataset shift diagnostics, and interpretable classification models, supported by quantitative benchmarks and a human validation study.

The article reports four central findings. First, diverse vision models align remarkably well through linear mapping, indicating that disparate architectures and training methods organize visual data into similar internal geometries. Second, aligning fixed, unimodal models to CLIP grants them strong zero-shot classification abilities; for example, a self-supervised model achieved 85% accuracy on an unseen 17-way categorization, and smaller models trained on roughly 0.3% of CLIP’s data volume occasionally matched or outperformed CLIP itself on specific tasks like color recognition. Third, the framework enables Concept Bottleneck Models without manual concept labeling, achieving 93.8% accuracy on the RIVAL10 benchmark while isolating the exact influence of individual concepts on final predictions. Fourth, translating internal classification vectors into language via a generative text model produced human-verified relevant descriptions in over 92% of evaluated cases.

These findings have major practical implications for model governance, development costs, and system transparency. Organizations can unlock zero-shot recognition, search, and explainability capabilities in smaller, highly efficient models without undertaking expensive multimodal data collection or intensive retraining. The approach lowers compute expenses and provides non-invasive diagnostic tools to detect operational risks, such as data distribution shifts or problematic spurious correlations, before deploying models in production.

Decision-makers should leverage linear alignment techniques to audit existing vision models, build interpretable classifiers for high-stakes domains, and diagnose dataset drift using human-understandable terms. When implementing this method, teams should prefer general regression objectives over task-specific cross-entropy alignment to preserve generalizability across diverse concepts.

The approach exhibits limitations when evaluating fine-grained character recognition tasks like optical character recognition and relies partly on the quality of CLIP text encodings and generative language prompts. Nevertheless, confidence in the core findings remains high, as the underlying linear alignment consistently succeeds across diverse architectures and requires only minimal, scalable optimization.

Moayeri et al (2023).pdf

No sufficiently relevant recommendations were found.

Cover for Text-To-Concept (and Back) via Cross-Model Alignment

Abstract

We observe that the mapping between an image’s representation in one model to its representation in another can be learned surprisingly well with just a linear layer, even across diverse models. Building on this observation, we propose text-to-concept, where features from a fixed pretrained model are aligned linearly to the CLIP space, so that text embeddings from CLIP’s text encoder become directly comparable to the aligned features. With text-to-concept, we convert fixed off-the-shelf vision encoders to surprisingly strong zero-shot classifiers for free, with accuracy at times even surpassing that of CLIP, despite being much smaller models and trained on a small fraction of the data compared to CLIP. We show other immediate use-cases of text-to-concept, like building concept bottleneck models with no concept supervision, diagnosing distribution shifts in terms of human concepts, and retrieving images satisfying a set of text-based constraints. Lastly, we demonstrate the feasibility of concept-to-text, where vectors in a model’s feature space are decoded by first aligning to the CLIP before being fed to a GPT-based generative model. Our work suggests existing deep models, with presumably diverse architectures and training, represent input samples relatively similarly, and a two-way communication across model representation spaces and to humans (through language) is viable.

Table of Contents

  • 1. Introduction
  • 2. Review of Literature
  • 3. Model Alignment
  • 4. Text to Concept
  • 4.1. Method Details
  • 4.2. Zero-Shot Classification
  • 5. Additional Applications of Text-to-Concept
  • 5.1. Concept-Bottleneck Networks for Free
  • 5.2. Concept-Based Dataset Summarization and Distribution Shift Diagnosis
  • 5.3. Concept Logic for Image Retrieval
  • 6. Concept-to-Text
  • 7. Conclusion
  • 8. Acknowledgements
  • References
  • A. Cross-Model Alignment
  • A.1. Optimizing Linear Transformation
  • A.2. Alternate Objectives for Alignment to CLIP
  • B. PC Alignment
  • C. Models
  • D. Optimizing Linear Alignment
  • E. Prompts for Text-to-concept
  • F. Zero-shot Classification
  • G. Concept Bottleneck Models
  • H. Concept Logic
  • I. Concept-to-text
  • J. Limitations
  • K. Additional Related Works

Knowls

  1. Knowl 1 — Affine alignment maps fixed vision representations between model spaces

    model/method

    Let fs:X→Rdsf_s:\mathcal{X}\to\mathbb{R}^{d_s} and ft:X→Rdtf_t:\mathcal{X}\to\mathbb{R}^{d_t} be fixed source and target image encoders, where X\mathcal{X} is the image domain and ds,dtd_s,d_t are their feature dimensions. The method learns only an affine map from source to target features, using paired representations of the same images from an unlabeled calibration set DtrainD_{\mathrm{train}}:

    (W∗,b∗)=arg⁡min⁡W∈Rds×dt, b∈Rdt1∣Dtrain∣∑x∈Dtrain∥W⊤fs(x)+b−ft(x)∥22.(W^*,b^*)=\arg\min_{W\in\mathbb{R}^{d_s\times d_t},\,b\in\mathbb{R}^{d_t}}\frac{1}{|D_{\mathrm{train}}|}\sum_{x\in D_{\mathrm{train}}}\left\|W^\top f_s(x)+b-f_t(x)\right\|_2^2.

    The encoders remain fixed; the optimization fits the linear transformation to reproduce the target encoder’s features, without labels or concept annotations. In the reported setup, calibration images came from ImageNet-1K training data. Before optimization, each representation space was rescaled so the variance of its elements was 4.54.5. The aligner was optimized with SGD (learning rate 0.010.01, momentum 0.90.9, weight decay 5×10−45\times10^{-4}), a cosine-annealing scheduler with Tmax⁡=200T_{\max}=200, and six epochs of training.

  2. Knowl 2 — Linear alignment retains substantial accuracy across diverse encoders

    data/table

    The table reports group-averaged results for ImageNet-1K models (all groups except CLIP were pretrained on ImageNet-1K). Each cell gives retained accuracy / R2R^2 when mapping source representations (rows) into target representations (columns). Retained accuracy is the accuracy of the source encoder followed by the learned map and target classifier, divided by the target model’s unaligned accuracy; R2R^2 measures the linear feature-regression fit. The groups are supervised ResNets (Sup RN), robust ResNets (Robust RN), supervised vision transformers (Sup ViT), self-supervised ResNets (SS RN), self-supervised vision transformers (SS ViT), and CLIP models.

    Source \\backslash target Sup RN Robust RN Sup ViT SS RN SS ViT CLIP
    Sup RN 0.98/0.77 1.03/0.59 0.82/0.35 0.86/0.47 0.79/0.77 0.80/0.67
    Robust RN 0.87/0.59 0.98/0.79 0.76/0.36 0.80/0.53 0.73/0.79 0.73/0.68
    Sup ViT 1.01/0.38 1.08/0.36 0.96/0.54 0.93/0.34 0.86/0.69 0.68/0.60
    SS RN 0.88/0.60 0.94/0.65 0.77/0.42 0.90/0.76 0.78/0.85 0.79/0.73
    SS ViT 0.93/0.54 0.97/0.56 0.83/0.38 0.81/0.51 0.88/0.87 0.82/0.71
    CLIP 0.72/0.42 0.76/0.43 0.58/0.23 0.55/0.39 0.46/0.72 0.80/0.77

    The results show that even encoders with different architectures and training objectives can be aligned well enough to preserve much of a target model’s classification accuracy; some retained-accuracy values exceed 11 because the aligned source features can work better with the target head than the target’s own features. Alignment into CLIP space retained 0.680.68–0.820.82 of CLIP accuracy across the non-CLIP source groups. In the reverse direction, CLIP-to-other-group retained accuracy was lower, consistent with the paper’s observation that mapping from a lower-dimensional space into a higher-dimensional one is more difficult.

  3. Knowl 3 — Text-to-concept makes text embeddings comparable to non-CLIP image features

    model/method

    Text-to-concept converts a text description into a concept vector that can be compared by cosine similarity with image features from a fixed, non-CLIP vision encoder. First, template prompts are filled with the concept text and encoded by a CLIP text encoder; the resulting vectors are averaged to form one concept vector. The study’s default ImageNet prompts are the templates used in CLIP’s original ImageNet zero-shot procedure. For object-agnostic contexts such as “in a tree,” the authors also form a refined vector by averaging embeddings across templates and class names. Next, the fixed image encoder is aligned to CLIP’s image space using the affine feature-regression method and paired unlabeled images. Because CLIP’s image and text representations share a space, cosine similarity between an aligned image feature and the text-derived concept vector provides the concept score. Once the aligner is trained, new text concepts require no further image collection or training. Concept vectors are not always accurate; the paper notes that prompt refinement or selecting more relevant positive and negative image examples can improve them.

  4. Knowl 4 — Aligned encoders gain zero-shot recognition of unseen categorizations

    empirical result

    The authors evaluated text-to-concept zero-shot classification by comparing aligned image features with CLIP text embeddings of candidate class descriptions and selecting the most similar class, without labeled examples from those candidate classes. The evaluated off-the-shelf encoders were roughly 25 million parameters and trained on ImageNet-1K; the CLIP ViT-B/16 text-and-image baseline was roughly 80 million parameters and was used as a comparison. On coarse-grained ImageNet categorizations, self-supervised vision transformers reached 85% accuracy on a 17-way task, despite not being trained on those categories. The aligned encoders also performed strongly on several out-of-distribution coarse-category tasks, and some outperformed the CLIP baseline, particularly on color recognition. Character recognition was a marked weakness: most models were only marginally above random accuracy, although an adversarially trained ResNet was about twice as accurate as the other models on zero-shot MNIST; CLIP also struggled. The models did better than random on primitive concepts such as textures, colors, and shapes. These results establish that alignment can expose useful zero-shot concept structure in encoders not trained with text supervision, while performance depends strongly on the concept and evaluation task.

  5. Knowl 5 — Text-to-concept enables a concept bottleneck without concept supervision

    empirical result

    The paper builds a concept bottleneck model (CBM) for RIVAL10, a ten-class ImageNet-derived task whose images have 28 annotated attributes. The annotations were not used to train the attribute predictors: CLIP text embeddings of the attribute names supplied the concept vectors, and ImageNet-pretrained ResNet-50 features were aligned to CLIP ViT-B/16. Cosine similarities between each aligned image feature and the 28 attribute vectors served as the concept scores; a linear classifier trained on those scores predicted RIVAL10 classes. The resulting CBM achieved 93.8% classification accuracy, compared with 94.5% for a linear classifier given the ground-truth attribute labels. The text-derived scores predicted attributes with 0.8 AUROC overall, and 72% of attributes had AUROC of at least 0.75. Because class logits are linear functions of the concept scores, the model exposes the contribution of each named concept to a prediction without requiring concept labels during training.

  6. Knowl 6 — Concept similarities reveal an indoor-scene shift in ObjectNet

    empirical result

    To diagnose distribution shift in human-interpretable terms, the authors compared the distribution of image similarities to the text-derived concept “indoors” for ImageNet and ObjectNet images represented by an ImageNet-pretrained ResNet-50 aligned to CLIP. ObjectNet consists of images photographed in people’s homes. Its “indoors” similarity distribution was significantly shifted to the right relative to ImageNet, with significance assessed by a Kolmogorov–Smirnov test. The comparison demonstrates how tracking similarities to a bank of text-defined concepts can characterize a dataset or flag changes in an incoming data stream in terms of recognizable semantic attributes.

  7. Knowl 7 — Concept logic retrieves images satisfying separate positive and negative constraints

    algorithm

    For a fixed image corpus and a vision encoder aligned to CLIP, concept logic retrieves images by applying independent similarity filters rather than embedding a long multi-part query as one phrase. Each constraint is a triple (c,k,s)(c,k,s): text concept cc, nonnegative threshold scale kk in standard deviations, and sign s∈{+1,−1}s\in\{+1,-1\}. Encode each concept as a CLIP text-derived vector and compute its cosine similarity to every aligned image feature in the corpus. For s=+1s=+1, retain images whose similarity is at least kk standard deviations above that concept’s corpus mean; for s=−1s=-1, retain images whose similarity is at least kk standard deviations below the mean. Intersect the retained sets for all constraints to produce the results. The scale controls constraint strictness, and the negative sign implements negation without relying on the text encoder to interpret words such as “not.” The paper illustrates queries such as dog + beach + sunset and cat + orange − indoors.

  8. Knowl 8 — Concept-to-text decodes vectors from non-CLIP model spaces

    model/method

    Concept-to-text maps a vector from a fixed vision model’s feature space into CLIP space and then decodes it with ZeroCap, a GPT-2-based generator guided by a CLIP embedding. The method is applied not only to image embeddings but also to arbitrary semantic vectors, specifically classification-head vectors associated with ImageNet classes. Before alignment, the classification-head vectors are rescaled so their element variance matches that of the image representations used to train the aligner; this reduces the mismatch between the head vectors and the aligner’s training inputs. The vectors are then passed through the learned affine map into CLIP space and decoded with the prompt “Image of a ” and a one-word output length. The authors used default ZeroCap captioning settings and did not further tune the vision encoders, CLIP, or GPT-2. The decoded word can describe the class object or a related/common object, including a spurious correlate, rather than necessarily naming the class exactly.

  9. Knowl 9 — Human evaluation finds decoded class vectors usually describe their classes

    data/table

    Amazon Mechanical Turk annotators compared a collage of images from an ImageNet class with the one-word concept-to-text decoding of that model’s classification vector. The first reported rate counts captions judged relevant to the class images (any response except “unrelated”); the second counts captions judged similar to the main object (the two most affirmative responses); the third counts captions judged to describe the main object itself. The high first-column rates support the claim that aligned classification vectors often decode to class-relevant words, while the lower main-object rates show that a relevant word need not name the central object.

    Model Relevant to class images Similar to main object Describes main object
    Swin Small 94.48% 84.60% 69.51%
    ResNet-50 95.14% 88.37% 71.36%
    DINO ViT-S/8 92.18% 76.45% 60.47%
    Average 93.93% 83.14% 67.11%

    Two annotators rated each model–class pair. The study therefore supports the feasibility of decoding general vectors from non-CLIP models, while also indicating that decoded concepts may be broader or associated with co-occurring objects rather than the class’s principal object.

  10. Knowl 10 — Top principal components show approximate cross-model correspondence

    empirical result

    The authors also aligned model representations after separately centering each model’s features and projecting them onto their top 40 principal components. They fit an affine map between these reduced spaces and examined whether each target component was primarily associated with the source component of the same rank, allowing a band of five components on either side of the diagonal. Approximate one-to-one correspondence and preservation of component order were strongest among CLIP models and among supervised or self-supervised ResNets; the pattern was weaker among supervised vision transformers. This analysis suggests that, for several model groups, leading principal directions encode similarly ordered abstract information even when the full architectures or training objectives differ.

Coverage note — The paper’s alternative task-specific and proposed data-free alignment objectives, plus the detailed per-dataset zero-shot plots and individual retrieval outputs, are omitted because they are secondary variants or examples rather than additional core findings; alignment sample-efficiency observations are also not separated into a knowl.

References

  1. 1.Bansal, Y., Nakkiran, P., and Barak, B. Revisiting model stitching to compare neural representations. Advances in Neural Information Processing Systems, 34:225–236, 2021.
  2. 2.Barbu, A., Mayo, D., Alverio, J., Luo, W., Wang, C., Gutfreund, D., Tenenbaum, J. B., and Katz, B. Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In NeurIPS, 2019.
  3. 3.Beyer, L., Zhai, X., Royer, A., Markeeva, L., Anil, R., and Kolesnikov, A. Knowledge distillation: A good teacher is patient and consistent. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10915–10924, 2021.
  4. 4.Caron, M., Touvron, H., Misra, I., Jegou, H., Mairal, J., Bojanowski, P., and Joulin, A. Emerging properties in self-supervised vision transformers. In Proceedings of the International Conference on Computer Vision (ICCV), 2021.
  5. 5.Chen*, X., Xie*, S., and He, K. An empirical study of training self-supervised vision transformers. arXiv preprint arXiv:2104.02057, 2021.
  6. 6.Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., , and Vedaldi, A. Describing textures in the wild. In Proceedings of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2014.
  7. 7.Coates, A., Ng, A., and Lee, H. An analysis of single-layer networks in unsupervised feature learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, pp. 215–223. JMLR Workshop and Conference Proceedings, 2011.
  8. 8.Csiszarik, A., Kórösi-Szabó, P., Matszangosz, A., Papp, G., and Varga, D. Similarity and matching of neural network representations. Advances in Neural Information Processing Systems, 34:5656–5668, 2021.
  9. 9.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009.
  10. 10.Eichenberg, C., Black, S., Weinbach, S., Parcalabescu, L., and Frank, A. Magma - multimodal augmentation of generative models through adapter-based finetuning. ArXiv, abs/2112.05253, 2021.
  11. 11.Fel, T., Picard, A., Bethune, L., Boissin, T., Vigouroux, D., Colin, J., Cadene, R., and Serre, T. Craft: Concept recursive activation factorization for explainability. ArXiv, abs/2211.10154, 2022.
  12. 12.Felix, R., Kumar, B. V., Reid, I. D., and Carneiro, G. Multimodal cycle-consistent generalized zero-shot learning. In European Conference on Computer Vision, 2018.
  13. 13.Ghorbani, A., Wexler, J., Zou, J. Y., and Kim, B. Towards automatic concept-based explanations. In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alche-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/file/77d2afcb31f6493e350fca61764efb9a-Paper.pdf.
  14. 14.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  15. 15.Hinton, G. E., Vinyals, O., and Dean, J. Distilling the knowledge in a neural network. ArXiv, abs/1503.02531, 2015.
  16. 16.Ilharco, G., Wortsman, M., Carlini, N., Taori, R., Dave, A., Shankar, V., Namkoong, H., Miller, J., Hajishirzi, H., Farhadi, A., and Schmidt, L. Open clip, 7 2021.
  17. 17.Jain, S., Lawrence, H., Moitra, A., and Madry, A. Distilling model failures as directions in latent space. arXiv preprint arXiv:2206.14754, 2022.
  18. 18.Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav). In International conference on machine learning, pp. 2668–2677. PMLR, 2018.
  19. 19.Koh, P. W., Nguyen, T., Tang, Y. S., Mussmann, S., Pierson, E., Kim, B., and Liang, P. Concept bottleneck models. In International Conference on Machine Learning, pp. 5338–5348. PMLR, 2020.
  20. 20.Korchi, A. and Ghanou, Y. 2d geometric shapes dataset – for machine learning and pattern recognition. Data in Brief, 32:106090, 07 2020. doi: 10.1016/j.dib.2020.106090.
  21. 21.Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images. 2009.
  22. 22.LeCun, Y., Cortes, C., and Burges, C. Mnist handwritten digit database. ATT Labs [Online]. Available: http://yann.lecun.com/exdb/mnist, 2, 2010.
  23. 23.Lenc, K. and Vedaldi, A. Understanding image representations by measuring their equivariance and equivalence. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 991–999, 2015.
  24. 24.Liu, Z., Luo, P., Wang, X., and Tang, X. Deep learning face attributes in the wild. In Proceedings of International Conference on Computer Vision (ICCV), December 2015.
  25. 25.Merullo, J., Castricato, L., Eickhoff, C., and Pavlick, E. Linearly mapping from image to text space. arXiv preprint arXiv:2209.15162, 2022.
  26. 26.Moayeri, M., Pope, P. E., Balaji, Y., and Feizi, S. A comprehensive study of image classification model sensitivity to foregrounds, backgrounds, and visual attributes. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 19065–19075, 2022.
  27. 27.Mokady, R., Hertz, A., and Bermano, A. H. Clipcap: Clip prefix for image captioning. ArXiv, abs/2111.09734, 2021.
  28. 28.Moschella, L., Maiorca, V., Fumero, M., Norelli, A., Locatello, F., and Rodola, E. Relative representations enable zero-shot latent space communication. ArXiv, abs/2209.15430, 2022.
  29. 29.Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. Reading digits in natural images with unsupervised feature learning. In NIPS Workshop on Deep Learning and Unsupervised Feature Learning 2011, 2011. URL http://ufldl.stanford.edu/housenumbers/nips2011_housenumbers.pdf.
  30. 30.Oikarinen, T. and Weng, T.-W. Clip-dissect: Automatic description of neuron representations in deep vision networks. arXiv preprint arXiv:2204.10965, 2022.
  31. 31.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32, pp. 8024–8035. Curran Associates, Inc., 2019.
  32. 32.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language models are unsupervised multitask learners. 2019.
  33. 33.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. In ICML, 2021.
  34. 34.Salman, H., Ilyas, A., Engstrom, L., Kapoor, A., and Madry, A. Do adversarially robust imagenet models transfer better? Advances in Neural Information Processing Systems, 33:3533–3545, 2020.
  35. 35.Santurkar, S., Tsipras, D., and Madry, A. BREEDS: benchmarks for subpopulation shift. CoRR, abs/2008.04859, 2020. URL https://arxiv.org/abs/2008.04859.
  36. 36.Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al. Laion-5b: An open large-scale dataset for training next generation image-text models. arXiv preprint arXiv:2210.08402, 2022.
  37. 37.Tewel, Y., Shalev, Y., Schwartz, I., and Wolf, L. Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 17897–17907, 2021.
  38. 38.Thomee, B., Shamma, D. A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., and Li, L.-J. Yfcc100m: The new data in multimedia research. Communications of the ACM, 59(2):64–73, 2016.
  39. 39.Tsimpoukelli, M., Menick, J., Cabi, S., Eslami, S. M. A., Vinyals, O., and Hill, F. Multimodal few-shot learning with frozen language models. In Neural Information Processing Systems, 2021.
  40. 40.Wightman, R. Pytorch image models. https://github.com/rwightman/pytorch-image-models, 2019.
  41. 41.Wold, S., Esbensen, K., and Geladi, P. Principal component analysis. Chemometrics and intelligent laboratory systems, 2(1-3):37–52, 1987.
  42. 42.Xiao, H., Rasul, K., and Vollgraf, R. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. ArXiv, abs/1708.07747, 2017.
  43. 43.Xiao, K., Engstrom, L., Ilyas, A., and Madry, A. Noise or signal: The role of image backgrounds in object recognition. ArXiv preprint arXiv:2006.09994, 2020.
  44. 44.Yuksekgonul, M., Wang, M., and Zou, J. Y. Post-hoc concept bottleneck models. ArXiv, abs/2205.15480, 2022.
  45. 45.Zhai, X., Wang, X., Mustafa, B., Steiner, A., Keysers, D., Kolesnikov, A., and Beyer, L. Lit: Zero-shot transfer with locked-image text tuning. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 18102–18112, 2021.
  46. 46.Zhang, R., Madumal, P., Miller, T., Ehinger, K. A., and Rubinstein, B. I. P. Invertible concept-based explanations for cnn models with non-negative concept activation vectors. In AAAI Conference on Artificial Intelligence, 2020.
  47. 47.Zhu, J.-Y., Park, T., Isola, P., and Efros, A. A. Unpaired image-to-image translation using cycle-consistent adversarial networks. 2017 IEEE International Conference on Computer Vision (ICCV), pp. 2242–2251, 2017.

Citation

MLA
Moayeri, M., et al. “Text-To-Concept (and Back) via Cross-Model Alignment”. International Conference on Machine Learning, vol. 202, 2023, pp. 25037–60, https://proceedings.mlr.press/v202/moayeri23a.html.
APA
Moayeri, M., Rezaei, K., Sanjabi, M., & Feizi, S. (2023). Text-To-Concept (and Back) via Cross-Model Alignment. International Conference on Machine Learning, 202, 25037–25060. https://proceedings.mlr.press/v202/moayeri23a.html
Chicago
Moayeri, M., K. Rezaei, M. Sanjabi, and S. Feizi. 2023. “Text-To-Concept (and Back) via Cross-Model Alignment”. International Conference on Machine Learning 202: 25037–60. https://proceedings.mlr.press/v202/moayeri23a.html.
Harvard
Moayeri, M. et al. (2023) “Text-To-Concept (and Back) via Cross-Model Alignment”, International Conference on Machine Learning. PMLR, pp. 25037–25060. Available at: https://proceedings.mlr.press/v202/moayeri23a.html.
Vancouver
1. Moayeri M, Rezaei K, Sanjabi M, Feizi S (2023) Text-To-Concept (and Back) via Cross-Model Alignment. In: International Conference on Machine Learning. PMLR, pp 25037–25060

BibTeX

@InProceedings{pmlr-v202-moayeri23a,
  title = 	 {Text-To-Concept (and Back) via Cross-Model Alignment},
  author =       {Moayeri, Mazda and Rezaei, Keivan and Sanjabi, Maziar and Feizi, Soheil},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {25037--25060},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/moayeri23a/moayeri23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/moayeri23a.html},
  abstract = 	 {We observe that the mapping between an image’s representation in one model to its representation in another can be learned surprisingly well with just a linear layer, even across diverse models. Building on this observation, we propose text-to-concept, where features from a fixed pretrained model are aligned linearly to the CLIP space, so that text embeddings from CLIP’s text encoder become directly comparable to the aligned features. With text-to-concept, we convert fixed off-the-shelf vision encoders to surprisingly strong zero-shot classifiers for free, with accuracy at times even surpassing that of CLIP, despite being much smaller models and trained on a small fraction of the data compared to CLIP. We show other immediate use-cases of text-to-concept, like building concept bottleneck models with no concept supervision, diagnosing distribution shifts in terms of human concepts, and retrieving images satisfying a set of text-based constraints. Lastly, we demonstrate the feasibility of concept-to-text, where vectors in a model’s feature space are decoded by first aligning to the CLIP before being fed to a GPT-based generative model. Our work suggests existing deep models, with presumably diverse architectures and training, represent input samples relatively similarly, and a two-way communication across model representation spaces and to humans (through language) is viable.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/