keyword
text encoder
A text encoder is a neural network component in artificial intelligence and natural language processing that converts natural language text into dense numerical vectors, commonly known as embeddings. By transforming unstructured words, phrases, or full sentences into high-dimensional vector representations, a text encoder captures semantic meaning, syntactic structure, and contextual nuances in a format that computational models can process. These generated representations serve as foundational features for a wide variety of machine learning tasks, including text classification, semantic search, cross-modal alignment between text and visual data, and guiding generative models to produce media based on descriptive textual prompts.
4 items

Disentangling visual and written concepts in CLIP
Joanna Materzynska, Antonio Torralba, David Bau
Why you should read this
Proposes an orthogonal projection method to separate text-reading capabilities from visual object processing in CLIP's image encoder, effectively eliminating text artifacts in guided image generation and defending against typographic attacks.
The CLIP network measures the similarity between natural text and images; in this work, we investigate the entanglement of the representation of word images and natural images in its image encoder. First, we find that the image encoder has an ability to match word images with natural images of scenes described by those words. This is consistent with previous research that suggests that the meaning and the spelling of a word might be entangled deep within the network. On the other hand, we also find that CLIP has a strong ability to match nonsense words, suggesting that processing of letters is separated from processing of their meaning. To explicitly determine whether the spelling capability of CLIP is separable, we devise a procedure for identifying representation subspaces that selectively isolate or eliminate spelling capabilities. We benchmark our methods against a range of retrieval tasks, and we also test them by measuring the appearance of text in CLIP-guided generated images. We find that our methods are able to cleanly separate spelling capabilities of CLIP from the visual processing of natural images.
Added
2026-09-26

RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation with Natural Prompts
Han Liu, Yuhao Wu, Shixuan Zhai, Bo Yuan, Ning Zhang
Why you should read this
Develops a genetic-algorithm-based optimization framework that generates natural, stealthy adversarial text prompts capable of reliably producing target images across diverse text-to-image models in both white-box and black-box settings.
The field of text-to-image generation has made remarkable strides in creating high-fidelity and photorealistic images. As this technology gains popularity, there is a growing concern about its potential security risks. However, there has been limited exploration into the robustness of these models from an adversarial perspective. Existing research has primarily focused on untargeted settings, and lacks holistic consideration for reliability (attack success rate) and stealthiness (imperceptibility). In this paper, we propose RIATIG, a reliable and imperceptible adversarial attack against text-to-image models via inconspicuous examples. By formulating the example crafting as an optimization process and solving it using a genetic-based method, our proposed attack can generate imperceptible prompts for text-to-image generation models in a reliable way. Evaluation of six popular text-to-image generation models demonstrates the efficiency and stealthiness of our attack in both white-box and black-box settings. To allow the community to build on top of our findings, we’ve made the artifacts available1.
Added
2026-09-26

Incorporating Hierarchy into Text Encoder: a Contrastive Learning Approach for Hierarchical Text Classification
Zihan Wang, Peiyi Wang, Lianzhe Huang, Xin Sun, Houfeng Wang
Why you should read this
Proposes a hierarchy-guided contrastive learning framework that directly embeds taxonomic label relationships into the text encoder using modified Graphormer structures, removing the need for separate, redundant label representations during inference.
Hierarchical text classification is a challenging subtask of multi-label classification due to its complex label hierarchy. Existing methods encode text and label hierarchy separately and mix their representations for classification, where the hierarchy remains unchanged for all input text. Instead of modeling them separately, in this work, we propose Hierarchy-guided Contrastive Learning (HGCLR) to directly embed the hierarchy into a text encoder. During training, HGCLR constructs positive samples for input text under the guidance of the label hierarchy. By pulling together the input text and its positive sample, the text encoder can learn to generate the hierarchy-aware text representation independently. Therefore, after training, the HGCLR enhanced text encoder can dispense with the redundant hierarchy. Extensive experiments on three benchmark datasets verify the effectiveness of HGCLR.
Added
2026-09-26

IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
Hu Ye, Jun Zhang, Siyi Liu, Xiao Han, Wei Yang
Why you should read this
Introduces a lightweight, decoupled cross-attention adapter that equips pretrained text-to-image diffusion models with image prompting capabilities while maintaining full compatibility with text prompts and structural controls without retraining the base model.
Recent years have witnessed the strong power of large text-to-image diffusion models for the impressive generative capability to create high-fidelity images. However, it is very tricky to generate desired images using only text prompt as it often involves complex prompt engineering. An alternative to text prompt is image prompt, as the saying goes: "an image is worth a thousand words". Although existing methods of direct fine-tuning from pretrained models are effective, they require large computing resources and are not compatible with other base models, text prompt, and structural controls. In this paper, we present IP-Adapter, an effective and lightweight adapter to achieve image prompt capability for the pretrained text-to-image diffusion models. The key design of our IP-Adapter is decoupled cross-attention mechanism that separates cross-attention layers for text features and image features. Despite the simplicity of our method, an IP-Adapter with only 22M parameters can achieve comparable or even better performance to a fully fine-tuned image prompt model. As we freeze the pretrained diffusion model, the proposed IP-Adapter can be generalized not only to other custom models fine-tuned from the same base model, but also to controllable generation using existing controllable tools. With the benefit of the decoupled cross-attention strategy, the image prompt can also work well with the text prompt to achieve multimodal image generation. The project page is available at \url{this https URL}.
Added
2026-09-24
