Built independently by an author, for readers. Read the story and support ChapterPal

keyword

visual attention

Visual attention is the cognitive and computational process of selectively focusing perceptual and processing resources on specific, salient regions, objects, or features within a visual scene while filtering out irrelevant background information. In biological vision and cognitive science, it guides eye movements, gaze fixations, and visual search through an interplay of stimulus-driven bottom-up cues, such as contrast and motion, and goal-directed top-down factors, such as task objectives and semantic context. In artificial intelligence and computer vision, visual attention is implemented through computational mechanisms and neural network layers that dynamically compute spatial or feature-wise importance weights, enabling models to prioritize informative image areas for tasks such as object detection, image segmentation, visual question answering, and automated image captioning.

20 items

From Gaze to Guidance: Interpreting and Adapting to Users' Cognitive Needs with Multimodal Gaze-Aware AI Assistants

From Gaze to Guidance: Interpreting and Adapting to Users' Cognitive Needs with Multimodal Gaze-Aware AI Assistants

Valdemar Danry, Javier Hernandez, Andrew D. Wilson, Pattie Maes, Judith Amores

OrganizationsMassachusetts Institute of TechnologyMicrosoft

Why you should read this

Demonstrates that equipping multimodal AI assistants with egocentric gaze tracking allows them to pinpoint user comprehension difficulties, significantly improving information recall while reducing interaction effort.

Current LLM assistants are powerful at answering questions, but they have limited access to the behavioral context that reveals when and where a user is struggling. We present a gaze-grounded multimodal LLM assistant that uses egocentric video with gaze overlays to identify likely points of difficulty and target follow-up retrospective assistance. We instantiate this vision in a controlled study (n=36) comparing the gaze-aware AI assistant to a text-only LLM assistant. Compared to a conventional LLM assistant, the gaze-aware assistant was rated as significantly more accurate and personalized in its assessments of users' reading behavior and significantly improved people's ability to recall information. Users spoke significantly fewer words with the gaze-aware assistant, indicating more efficient interactions. Qualitative results underscored both perceived benefits in comprehension and challenges when interpretations of gaze behaviors were inaccurate. Our findings suggest that gaze-aware LLM assistants can reason about cognitive needs to improve cognitive outcomes of users.

Added

2026-09-29

The Secrets of Salient Object Segmentation

The Secrets of Salient Object Segmentation

Yin Li, Xiaodi Hou, Christof Koch, James M. Rehg, Alan L. Yuille

OrganizationsAllen Institute for Brain ScienceCalifornia Institute of TechnologyGeorgia Institute of TechnologyUniversity of California, Los Angeles

Why you should read this

Exposes critical design biases in salient object benchmarks and introduces a unified dataset with joint fixation and segmentation ground truth alongside a method that directly connects human visual attention to object segmentation.

In this paper we provide an extensive evaluation of fixation prediction and salient object segmentation algorithms as well as statistics of major datasets. Our analysis identifies serious design flaws of existing salient object benchmarks, called the dataset design bias, by over emphasizing the stereotypical concepts of saliency. The dataset design bias does not only create the discomforting disconnection between fixations and salient object segmentation, but also misleads the algorithm designing. Based on our analysis, we propose a new high quality dataset that offers both fixation and salient object segmentation ground-truth. With fixations and salient object being presented simultaneously, we are able to bridge the gap between fixations and salient objects, and propose a novel method for salient object segmentation. Finally, we report significant benchmark progress on three existing datasets of segmenting salient objects

Added

2026-09-25

Visual saliency based on multiscale deep features

Visual saliency based on multiscale deep features

Guanbin Li, Yizhou Yu

OrganizationsUniversity of Hong Kong

Why you should read this

Introduces a multiscale deep convolutional framework for visual saliency detection that integrates spatial coherence refinement across multiple segmentation levels and establishes state-of-the-art accuracy alongside the 4,447-image HKU-IS benchmark dataset.

Visual saliency is a fundamental problem in both cognitive and computational sciences, including computer vision. In this CVPR 2015 paper, we discover that a high-quality visual saliency model can be trained with multiscale features extracted using a popular deep learning architecture, convolutional neural networks (CNNs), which have had many successes in visual recognition tasks. For learning such saliency models, we introduce a neural network architecture, which has fully connected layers on top of CNNs responsible for extracting features at three different scales. We then propose a refinement method to enhance the spatial coherence of our saliency results. Finally, aggregating multiple saliency maps computed for different levels of image segmentation can further boost the performance, yielding saliency maps better than those generated from a single segmentation. To promote further research and evaluation of visual saliency models, we also construct a new large database of 4447 challenging images and their pixelwise saliency annotation. Experimental results demonstrate that our proposed method is capable of achieving state-of-the-art performance on all public benchmarks, improving the F-Measure by 5.0% and 13.2% respectively on the MSRA-B dataset and our new dataset (HKU-IS), and lowering the mean absolute error by 5.7% and 35.1% respectively on these two datasets.

Added

2026-09-25

Attention to Scale: Scale-Aware Semantic Image Segmentation

Attention to Scale: Scale-Aware Semantic Image Segmentation

Liang-Chieh Chen, Yi Yang, Jiang Wang, Wei Xu, Alan L. Yuille

OrganizationsBaiduJohns Hopkins UniversityUniversity of California, Los Angeles

Why you should read this

Proposes an attention mechanism that dynamically weights multi-scale features at each pixel, improving semantic image segmentation accuracy over standard pooling baselines while providing interpretable diagnostics of scale selection.

Incorporating multi-scale features in fully convolutional neural networks (FCNs) has been a key element to achieving state-of-the-art performance on semantic image segmentation. One common way to extract multi-scale features is to feed multiple resized input images to a shared deep network and then merge the resulting features for pixelwise classification. In this work, we propose an attention mechanism that learns to softly weight the multi-scale features at each pixel location. We adapt a state-of-the-art semantic image segmentation model, which we jointly train with multi-scale input images and the attention model. The proposed attention model not only outperforms average- and max-pooling, but allows us to diagnostically visualize the importance of features at different positions and scales. Moreover, we show that adding extra supervision to the output at each scale is essential to achieving excellent performance when merging multi-scale features. We demonstrate the effectiveness of our model with extensive experiments on three challenging datasets, including PASCAL-Person-Part, PASCAL VOC 2012 and a subset of MS-COCO 2014.

Added

2026-09-25

Deeply Supervised Salient Object Detection with Short Connections

Deeply Supervised Salient Object Detection with Short Connections

Qibin Hou, Ming-Ming Cheng, Xiao-Wei Hu, Ali Borji, Zhuowen Tu, Philip Torr

OrganizationsNankai UniversityUniversity of California, San DiegoUniversity of Central FloridaUniversity of Oxford

Why you should read this

Introduces short connections to deeply supervised skip-layer architectures to capture multi-scale features, achieving state-of-the-art salient object detection accuracy and fast inference across five standard benchmarks.

Recent progress on saliency detection is substantial, benefiting mostly from the explosive development of Convolutional Neural Networks (CNNs). Semantic segmentation and saliency detection algorithms developed lately have been mostly based on Fully Convolutional Neural Networks (FCNs). There is still a large room for improvement over the generic FCN models that do not explicitly deal with the scale-space problem. Holistically-Nested Edge Detector (HED) provides a skip-layer structure with deep supervision for edge and boundary detection, but the performance gain of HED on salience detection is not obvious. In this paper, we propose a new method for saliency detection by introducing short connections to the skip-layer structures within the HED architecture. Our framework provides rich multi-scale feature maps at each layer, a property that is critically needed to perform segment detection. Our method produces state-of-the-art results on 5 widely tested salient object detection benchmarks, with advantages in terms of efficiency (0.15 seconds per image), effectiveness, and simplicity over the existing algorithms.

Added

2026-09-25

Knowing When to Look: Adaptive Attention via a Visual Sentinel for Image Captioning

Knowing When to Look: Adaptive Attention via a Visual Sentinel for Image Captioning

Jiasen Lu, Caiming Xiong, Devi Parikh, Richard Socher

OrganizationsGeorgia Institute of TechnologySalesforceVirginia Tech

Why you should read this

Introduces an adaptive attention mechanism with a visual sentinel that allows image captioning models to dynamically decide when to attend to visual features versus relying on linguistic context, substantially improving caption generation accuracy on standard benchmarks.

Attention-based neural encoder-decoder frameworks have been widely adopted for image captioning. Most methods force visual attention to be active for every generated word. However, the decoder likely requires little to no visual information from the image to predict non-visual words such as "the" and "of". Other words that may seem visual can often be predicted reliably just from the language model e.g., "sign" after "behind a red stop" or "phone" following "talking on a cell". In this paper, we propose a novel adaptive attention model with a visual sentinel. At each time step, our model decides whether to attend to the image (and if so, to which regions) or to the visual sentinel. The model decides whether to attend to the image and where, in order to extract meaningful information for sequential word generation. We test our method on the COCO image captioning 2015 challenge dataset and Flickr30K. Our approach sets the new state-of-the-art by a significant margin.

Added

2026-09-24

Image Captioning with Semantic Attention

Image Captioning with Semantic Attention

Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, Jiebo Luo

OrganizationsAdobeUniversity of Rochester

Why you should read this

Proposes a semantic attention framework for image captioning that selectively fuses bottom-up concept proposals with top-down visual features in recurrent neural networks to generate more accurate descriptions.

Automatically generating a natural language description of an image has attracted interests recently both because of its importance in practical applications and because it connects two major artificial intelligence fields: computer vision and natural language processing. Existing approaches are either top-down, which start from a gist of an image and convert it into words, or bottom-up, which come up with words describing various aspects of an image and then combine them. In this paper, we propose a new algorithm that combines both approaches through a model of semantic attention. Our algorithm learns to selectively attend to semantic concept proposals and fuse them into hidden states and outputs of recurrent neural networks. The selection and fusion form a feedback connecting the top-down and bottom-up computation. We evaluate our algorithm on two public benchmarks: Microsoft COCO and Flickr30K. Experimental results show that our algorithm significantly outperforms the state-of-the-art approaches consistently across different evaluation metrics.

Added

2026-09-19

Salient Object Detection: A Benchmark

Salient Object Detection: A Benchmark

Ali Borji, Ming-Ming Cheng, Huaizu Jiang, Jia Li

OrganizationsBeihang UniversityNankai UniversityUniversity of Wisconsin-MilwaukeeXi'an Jiaotong University

Why you should read this

Establishes a comprehensive benchmark by evaluating forty models across six datasets, analyzing factors like center bias and scene complexity to expose key failure modes and guide future salient object detection research.

We extensively compare, qualitatively and quantitatively, 40 state-of-the-art models (28 salient object detection, 10 fixation prediction, 1 objectness, and 1 baseline) over 6 challenging datasets for the purpose of benchmarking salient object detection and segmentation methods. From the results obtained so far, our evaluation shows a consistent rapid progress over the last few years in terms of both accuracy and running time. The top contenders in this benchmark significantly outperform the models identified as the best in the previous benchmark conducted just two years ago. We find that the models designed specifically for salient object detection generally work better than models in closely related areas, which in turn provides a precise definition and suggests an appropriate treatment of this problem that distinguishes it from other problems. In particular, we analyze the influences of center bias and scene complexity in model performance, which, along with the hard cases for state-of-the-art models, provide useful hints towards constructing more challenging large scale datasets and better saliency models. Finally, we propose probable solutions for tackling several open problems such as evaluation scores and dataset bias, which also suggest future research directions in the rapidly-growing field of salient object detection.

Added

2026-09-18

Attention U-Net: Learning Where to Look for the Pancreas

Attention U-Net: Learning Where to Look for the Pancreas

Ozan Oktay, Jo Schlemper, Loic Le Folgoc, Matthew Lee, Mattias Heinrich, Kazunari Misawa, Kensaku Mori, Steven McDonagh, Nils Y Hammerla, Bernhard Kainz, Ben Glocker, Daniel Rueckert

OrganizationsAichi Cancer CenterBabylon HealthHeartFlowImperial College LondonNagoya UniversityUniversity of Lübeck

Why you should read this

Proposes attention gates for the U-Net architecture that automatically focus on target structures of varying shapes and sizes while suppressing irrelevant background regions, eliminating the need for multi-stage localization pipelines in medical image segmentation.

We propose a novel attention gate (AG) model for medical imaging that automatically learns to focus on target structures of varying shapes and sizes. Models trained with AGs implicitly learn to suppress irrelevant regions in an input image while highlighting salient features useful for a specific task. This enables us to eliminate the necessity of using explicit external tissue/organ localisation modules of cascaded convolutional neural networks (CNNs). AGs can be easily integrated into standard CNN architectures such as the U-Net model with minimal computational overhead while increasing the model sensitivity and prediction accuracy. The proposed Attention U-Net architecture is evaluated on two large CT abdominal datasets for multi-class image segmentation. Experimental results show that AGs consistently improve the prediction performance of U-Net across different datasets and training sizes while preserving computational efficiency. The code for the proposed architecture is publicly available.

Added

2026-09-11

Recurrent Models of Visual Attention

Recurrent Models of Visual Attention

Volodymyr Mnih, Nicolas Heess, Alex Graves, Koray Kavukcuoglu

OrganizationsGoogle

Why you should read this

Proposes a recurrent attention model trained with reinforcement learning that sequentially samples task-relevant image regions at high resolution, decoupling computational cost from total image size while outperforming standard convolutional networks on cluttered visual tasks.

Applying convolutional neural networks to large images is computationally expensive because the amount of computation scales linearly with the number of image pixels. We present a novel recurrent neural network model that is capable of extracting information from an image or video by adaptively selecting a sequence of regions or locations and only processing the selected regions at high resolution. Like convolutional neural networks, the proposed model has a degree of translation invariance built-in, but the amount of computation it performs can be controlled independently of the input image size. While the model is non-differentiable, it can be trained using reinforcement learning methods to learn task-specific policies. We evaluate our model on several image classification tasks, where it significantly outperforms a convolutional neural network baseline on cluttered images, and on a dynamic visual control problem, where it learns to track a simple object without an explicit training signal for doing so.

Added

2026-09-11

License

Published with permission

DRAW: A Recurrent Neural Network For Image Generation

DRAW: A Recurrent Neural Network For Image Generation

Karol Gregor, Ivo Danihelka, Alex Graves, Danilo Jimenez Rezende, Daan Wierstra

OrganizationsGoogle

Why you should read this

Demonstrates that teaching a neural network to build images incrementally by focusing attention on different regions—similar to how an artist sketches and refines a drawing—produces more realistic generated images than methods that try to create entire pictures in one step, suggesting that human-like sequential construction is key to capturing visual complexity.

This paper introduces the Deep Recurrent Attentive Writer (DRAW) neural network architecture for image generation. DRAW networks combine a novel spatial attention mechanism that mimics the foveation of the human eye, with a sequential variational auto-encoding framework that allows for the iterative construction of complex images. The system substantially improves on the state of the art for generative models on MNIST, and, when trained on the Street View House Numbers dataset, it generates images that cannot be distinguished from real data with the naked eye.

Added

2026-02-21

Squeeze-and-Excitation Networks

Squeeze-and-Excitation Networks

Jie Hu, Li Shen, Gang Sun

OrganizationsInstitute of Automation, Chinese Academy of SciencesInstitute of Software, Chinese Academy of SciencesMomentaUniversity of Chinese Academy of SciencesUniversity of MacauUniversity of Oxford

Why you should read this

Introduces a lightweight attention mechanism that adaptively recalibrates channel-wise feature responses by explicitly modeling interdependencies between channels.

Convolutional neural networks are built upon the convolution operation, which extracts informative features by fusing spatial and channel-wise information together within local receptive fields. In order to boost the representational power of a network, several recent approaches have shown the benefit of enhancing spatial encoding. In this work, we focus on the channel relationship and propose a novel architectural unit, which we term the "Squeeze-and-Excitation" (SE) block, that adaptively recalibrates channel-wise feature responses by explicitly modelling interdependencies between channels. We demonstrate that by stacking these blocks together, we can construct SENet architectures that generalise extremely well across challenging datasets. Crucially, we find that SE blocks produce significant performance improvements for existing state-of-the-art deep architectures at minimal additional computational cost. SENets formed the foundation of our ILSVRC 2017 classification submission which won first place and significantly reduced the top-5 error to 2.251%, achieving a ~25% relative improvement over the winning entry of 2016. Code and models are available at https://github.com/hujie-frank/SENet.

Added

2026-02-18

Show, Attend and Tell: Neural Image Caption Generation with Visual Attention

Show, Attend and Tell: Neural Image Caption Generation with Visual Attention

Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, Yoshua Bengio

OrganizationsUniversité de MontréalUniversity of Toronto

Why you should read this

Demonstrates how a novel attention-based model not only generates state-of-the-art image captions but also visually proves its ability to dynamically focus on salient objects within an image, offering a transparent and intuitive approach to understanding AI's "gaze."

Inspired by recent work in machine translation and object detection, we introduce an attention based model that automatically learns to describe the content of images. We describe how we can train this model in a deterministic manner using standard backpropagation techniques and stochastically by maximizing a variational lower bound. We also show through visualization how the model is able to automatically learn to fix its gaze on salient objects while generating the corresponding words in the output sequence. We validate the use of attention with state-of-the-art performance on three benchmark datasets: Flickr8k, Flickr30k and MS COCO.

Added

2026-02-11

License

Published with permission

Show, Attend and Tell: Neural Image Caption Generation with Visual Attention

Show, Attend and Tell: Neural Image Caption Generation with Visual Attention

Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, Yoshua Bengio

OrganizationsUniversité de MontréalUniversity of Toronto

Why you should read this

Demonstrates that attention can apply to spatial grids (images), introducing "hard" stochastic attention alongside "soft" deterministic attention.

Inspired by recent work in machine translation and object detection, we introduce an attention based model that automatically learns to describe the content of images. We describe how we can train this model in a deterministic manner using standard backpropagation techniques and stochastically by maximizing a variational lower bound. We also show through visualization how the model is able to automatically learn to fix its gaze on salient objects while generating the corresponding words in the output sequence. We validate the use of attention with state-of-the-art performance on three benchmark datasets: Flickr8k, Flickr30k and MS COCO.

Added

2026-01-28