Built independently by an author, for readers. Read the story and support ChapterPal

keyword

recurrent architecture

A recurrent architecture is a neural network design in which computational units or layers repeatedly apply the same operations across multiple steps, using outputs or internal states from prior steps as inputs for subsequent iterations. Unlike feedforward systems that process information in a single forward pass, recurrent architectures maintain an internal memory mechanism through cyclical connections, allowing the network to handle sequential inputs or progressively refine representations and predictions over time. By sharing parameters across iterations, these architectures efficiently model temporal dependencies, retain contextual information across dynamic sequence lengths, and iteratively update feature representations to improve task performance.

1 item

CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor

CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor

Shuyang Sun, Runjia Li, Philip Torr, Xiuye Gu, Siyang Li

OrganizationsGoogleUniversity of Oxford

Why you should read this

Presents a training-free recurrent framework that leverages frozen CLIP models to iteratively refine mask proposals and filter non-existent text queries, setting new state-of-the-art benchmarks in zero-shot open-vocabulary semantic and referring segmentation without sacrificing vocabulary breadth.

Existing open-vocabulary image segmentation methods require a fine-tuning step on mask labels and/or image-text datasets. Mask labels are labor-intensive, which limits the number of categories in segmentation datasets. Consequently, the vocabulary capacity of pre-trained VLMs is severely reduced after fine-tuning. However, without fine-tuning, VLMs trained under weak image-text supervision tend to make suboptimal mask predictions. To alleviate these issues, we introduce a novel recurrent framework that progressively filters out irrelevant texts and enhances mask quality without training efforts. The recurrent unit is a two-stage segmenter built upon a frozen VLM. Thus, our model retains the VLM’s broad vocabulary space and equips it with segmentation ability. Experiments show that our method outperforms not only the training-free counterparts, but also those fine-tuned with millions of data samples, and sets the new state-of-the-art records for both zero-shot semantic and referring segmentation. Concretely, we improve the current record by 28.8, 16.0, and 6.9 mIoU on Pascal VOC, COCO Object, and Pascal Context.

Added

2026-09-26