Built independently by an author, for readers. Read the story and support ChapterPal

keyword

text detection

Text detection is a computer vision task that involves identifying and determining the precise spatial locations of written, printed, or rendered text within digital images or video frames. Often referred to as text localization, the process typically predicts bounding shapes, such as axis-aligned rectangles, rotated quadrilaterals, or polygon masks, that enclose text instances at the character, word, or line level across varied layouts and natural scenes. While text recognition focuses on converting visual patterns into machine-readable character strings, text detection is concerned solely with pinpointing where text exists in a visual input, serving as a foundational stage for optical character recognition, end-to-end text spotting, document intelligence, and multimodal visual reasoning systems.

5 items

Disentangling visual and written concepts in CLIP

Disentangling visual and written concepts in CLIP

Joanna Materzynska, Antonio Torralba, David Bau

OrganizationsHarvard UniversityMassachusetts Institute of Technology

Why you should read this

Proposes an orthogonal projection method to separate text-reading capabilities from visual object processing in CLIP's image encoder, effectively eliminating text artifacts in guided image generation and defending against typographic attacks.

The CLIP network measures the similarity between natural text and images; in this work, we investigate the entanglement of the representation of word images and natural images in its image encoder. First, we find that the image encoder has an ability to match word images with natural images of scenes described by those words. This is consistent with previous research that suggests that the meaning and the spelling of a word might be entangled deep within the network. On the other hand, we also find that CLIP has a strong ability to match nonsense words, suggesting that processing of letters is separated from processing of their meaning. To explicitly determine whether the spelling capability of CLIP is separable, we devise a procedure for identifying representation subspaces that selectively isolate or eliminate spelling capabilities. We benchmark our methods against a range of retrieval tasks, and we also test them by measuring the appearance of text in CLIP-guided generated images. We find that our methods are able to cleanly separate spelling capabilities of CLIP from the visual processing of natural images.

Added

2026-09-26

OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition

OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition

Jianqiang Wan, Sibo Song, Wenwen Yu, Yuliang Liu, Wenqing Cheng, Fei Huang, Xiang Bai, Cong Yao, Zhibo Yang

OrganizationsAlibaba GroupHuazhong University of Science and Technology

Why you should read this

Introduces OmniParser, a unified encoder-decoder framework that simultaneously handles text spotting, key information extraction, and table recognition by decoupling structured point sequences from region and text content generation.

Recently, visually-situated text parsing (VsTP) has experienced notable advancements, driven by the increasing demand for automated document understanding and the emergence of Generative Large Language Models (LLMs) capable of processing document-based questions. Various methods have been proposed to address the challenging problem of VsTP. However, due to the diversified targets and heterogeneous schemas, previous works usually design task-specific architectures and objectives for individual tasks, which inadvertently leads to modal isolation and complex workflow. In this paper, we propose a unified paradigm for parsing visually-situated text across diverse scenarios. Specifically, we devise a universal model, called OmniParser, which can simultaneously handle three typical visually-situated text parsing tasks: text spotting, key information extraction, and table recognition. In OmniParser, all tasks share the unified encoder-decoder architecture, the unified objective: point-conditioned text generation, and the unified input&output representation: prompt & structured sequences. Extensive experiments demonstrate that the proposed OmniParser achieves state-of-the-art (SOTA) or highly competitive performances on 7 datasets for the three visually-situated text parsing tasks, despite its unified, concise design. The code is available at AdvancedLiterateMachinery.

Added

2026-09-26

SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recognition

SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recognition

Mingxin Huang, Yuliang Liu, Zhenghao Peng, Chongyu Liu, Dahua Lin, Shenggao Zhu, Nicholas Jing Yuan, Kai Ding, Lianwen Jin

OrganizationsHuaweiIntSig Information Co., Ltd.Peng Cheng LaboratorySouth China University of TechnologyThe Chinese University of Hong Kong

Why you should read this

Presents an end-to-end scene text spotting framework that uses a Recognition Conversion mechanism to let recognition loss directly guide text localization, eliminating the need for character-level annotations or complex rectification modules on arbitrarily-shaped text.

End-to-end scene text spotting has attracted great attention in recent years due to the success of excavating the intrinsic synergy of the scene text detection and recognition. However, recent state-of-the-art methods usually incorporate detection and recognition simply by sharing the backbone, which does not directly take advantage of the feature interaction between the two tasks. In this paper, we propose a new end-to-end scene text spotting framework termed SwinTextSpotter. Using a transformer encoder with dynamic head as the detector, we unify the two tasks with a novel Recognition Conversion mechanism to explicitly guide text localization through recognition loss. The straightforward design results in a concise framework that requires neither additional rectification module nor character-level annotation for the arbitrarily-shaped text. Qualitative and quantitative experiments on multi-oriented datasets RoIC13 and ICDAR 2015, arbitrarily-shaped datasets Total-Text and CTW1500, and multi-lingual datasets ReCTS (Chinese) and VinText (Vietnamese) demonstrate SwinTextSpotter significantly outperforms existing methods. Code is available at https://github.com/mxin262/SwinTextSpotter.

Added

2026-09-26

Towards VQA Models That Can Read

Towards VQA Models That Can Read

Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, Marcus Rohrbach

OrganizationsGeorgia Institute of TechnologyMeta

Why you should read this

Introduces the TextVQA dataset and the LoRRA model to enable visual question answering systems to read and reason about text embedded in everyday images.

Studies have shown that a dominant class of questions asked by visually impaired users on images of their surroundings involves reading text in the image. But today's VQA models can not read! Our paper takes a first step towards addressing this problem. First, we introduce a new "TextVQA" dataset to facilitate progress on this important problem. Existing datasets either have a small proportion of questions about text (e.g., the VQA dataset) or are too small (e.g., the VizWiz dataset). TextVQA contains 45,336 questions on 28,408 images that require reasoning about text to answer. Second, we introduce a novel model architecture that reads text in the image, reasons about it in the context of the image and the question, and predicts an answer which might be a deduction based on the text and the image or composed of the strings found in the image. Consequently, we call our approach Look, Read, Reason & Answer (LoRRA). We show that LoRRA outperforms existing state-of-the-art VQA models on our TextVQA dataset. We find that the gap between human performance and machine performance is significantly larger on TextVQA than on VQA 2.0, suggesting that TextVQA is well-suited to benchmark progress along directions complementary to VQA 2.0.

Added

2026-09-16