Built independently by an author, for readers. Read the story and support ChapterPal

keyword

non-local attention

Non-local attention is a deep learning mechanism that computes the representation of a specific position or token by aggregating features from across the entire input rather than restricting computations to a local neighborhood. Derived from the concept of non-local means in classical image processing, it calculates similarity weights between distant spatial or temporal features, allowing models to capture long-range dependencies and global context. Unlike standard local convolutions or window-constrained attention layers that only focus on adjacent regions, non-local attention enables neural networks to leverage self-similarity and recurring patterns distributed across distant areas, making it especially effective for tasks such as image restoration, enhancement, and dense visual recognition.

2 items

Transcending the Limit of Local Window: Advanced Super-Resolution Transformer with Adaptive Token Dictionary

Transcending the Limit of Local Window: Advanced Super-Resolution Transformer with Adaptive Token Dictionary

Leheng Zhang, Yawei Li, Xingyu Zhou, Xiaorui Zhao, Shuhang Gu

OrganizationsETH ZurichUniversity of Electronic Science and Technology of China

Why you should read this

Proposes an adaptive token dictionary for super-resolution Transformers that overcomes the receptive-field limits of window-based attention by dynamically grouping similar image tokens across the entire image to capture long-range dependencies.

Single Image Super-Resolution is a classic computer vision problem that involves estimating high-resolution (HR) images from low-resolution (LR) ones. Although deep neural networks (DNNs), especially Transformers for super-resolution, have seen significant advancements in recent years, challenges still remain, particularly in limited receptive field caused by window-based self-attention. To address these issues, we introduce a group of auxiliary Adaptive Token Dictionary to SR Transformer and establish an ATD-SR method. The introduced token dictionary could learn prior information from training data and adapt the learned prior to specific testing image through an adaptive refinement step. The refinement strategy could not only provide global information to all input tokens but also group image tokens into categories. Based on category partitions, we further propose a category-based self-attention mechanism designed to leverage distant but similar tokens for enhancing input features. The experimental results show that our method achieves the best performance on various single image super-resolution benchmarks.

Added

2026-09-26

ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive Coding

ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive Coding

Dailan He, Ziming Yang, Weikun Peng, Rui Ma, Hongwei Qin, Yan Wang

OrganizationsSenseTimeTsinghua University

Why you should read this

Presents ELIC, a learned image compression architecture that combines uneven space-channel contextual coding with efficient transform design to achieve state-of-the-art rate-distortion performance alongside fast inference, preview decoding, and progressive decoding.

Recently, learned image compression techniques have achieved remarkable performance, even surpassing the best manually designed lossy image coders. They are promising to be large-scale adopted. For the sake of practicality, a thorough investigation of the architecture design of learned image compression, regarding both compression performance and running speed, is essential. In this paper, we first propose uneven channel-conditional adaptive coding, motivated by the observation of energy compaction in learned image compression. Combining the proposed uneven grouping model with existing context models, we obtain a spatial-channel contextual adaptive model to improve the coding performance without damage to running speed. Then we study the structure of the main transform and propose an efficient model, ELIC, to achieve state-of-the-art speed and compression ability. With superior performance, the proposed model also supports extremely fast preview decoding and progressive decoding, which makes the coming application of learning-based image compression more promising.

Added

2026-09-26