State-of-the-Art in Visual Attention Modeling

Ali BorjiLaurent Itti

article2013TPAMI1,934 citations

Presents a comprehensive taxonomy and qualitative evaluation of nearly 65 computational visual attention models across 13 behavioral and structural criteria to clarify their mechanisms, limitations, and practical applications in computer vision and robotics.

Listen

Visual environments bombard observers with hundreds of millions of data bits per second, making complete real-time processing impossible for biological and artificial systems alike. Visual attention acts as a vital filtering mechanism that selects behaviorally relevant information and discards unnecessary background data. As computer vision, robotics, and cognitive systems expand into complex, real-time operating environments, building accurate computational models of visual selection has become critical for managing processing demands and enhancing autonomous decision-making.

The article systematically evaluates the computational state of the art in visual attention modeling by analyzing approximately 65 distinct models across cognitive, Bayesian, decision-theoretic, information-theoretic, graphical, spectral, and machine learning domains. It aims to provide a unified conceptual framework and establish a qualitative taxonomy across 13 core behavioral and computational criteria, clarifying how existing approaches predict where humans look.

The review classifies models based on foundational attributes such as stimulus-driven (bottom-up) versus goal-driven (top-down) processing, spatial versus spatio-temporal dynamics, and space-based versus object-based representations. It examines how these frameworks are trained and tested against empirical benchmarks, including static and video eye-tracking datasets, using point-based, region-based, and distribution-level evaluation metrics.

The analysis reveals several key findings regarding model capabilities and current limitations. First, the vast majority of existing computational frameworks remain stimulus-driven and space-based, relying primarily on low-level features such as color, intensity, and orientation, which capture only a minor fraction of human gaze allocation in real-world scenarios. Second, biologically inspired architectures, such as decision-theoretic and adaptive whitening formulations, consistently outperform purely heuristic models in fixation prediction accuracy while successfully replicating known behavioral phenomena. Third, baseline model evaluations are frequently distorted by systematic experimental artifacts; for example, trivial center-bias and border effects can artificially inflate performance metrics such as area under the curve scores, sometimes allowing a simple central Gaussian distribution to outperform sophisticated saliency algorithms. Finally, machine learning classifiers trained on high-level features like faces and text achieve strong predictive accuracy but often operate as data-dependent black boxes that lack biological interpretability.

These findings have direct operational implications for deploying vision systems in autonomous navigation, video compression, image quality assessment, and human-robot interaction. Relying heavily on bottom-up saliency introduces substantial performance risks in dynamic settings where mission-critical tasks and expectations govern visual prioritization. Furthermore, unstandardized evaluation metrics create false confidence regarding model reliability, risking suboptimal resource allocation when integrating these systems into real-world applications.

To advance the field, stakeholders and developers should prioritize creating unified, standardized benchmark datasets and evaluation protocols—such as shuffled metrics that neutralize center-bias—analogous to established challenges in object and face recognition. Development efforts should pivot toward formulating principled computational frameworks for task-driven, top-down attention and integrating temporal dynamics for interactive and virtual environments. Future research must also develop rigorous experimental criteria to validate biological plausibility and bridge the persistent gap between covert mental focus and overt eye movements.

  • Paper: A Model of Saliency-Based Visual Attention for Rapid Scene Analysis, L. Itti et al. (1998). This seminal paper introduced the foundational biologically inspired, bottom-up saliency map architecture that serves as the core reference baseline evaluated throughout the survey.
  • Paper: Graph-Based Visual Saliency, Jonathan Harel et al. (2006). This work establishes Graph-Based Visual Saliency (GBVS), a major milestone in bottom-up fixation prediction that the survey categorizes and compares extensively.
  • Paper: Global contrast based salient region detection, Ming-Ming Cheng et al. (2011). This paper presents regional and histogram-based global contrast algorithms that form key representative methods in the survey's taxonomy of salient region detection.
  • Paper: Attention mechanisms in computer vision: A survey, Meng-Hao Guo et al. (2021). This survey provides a comprehensive modern follow-up that traces how visual attention evolved from classical saliency modeling into deep learning mechanisms and vision transformers.
  • Paper: Recurrent Models of Visual Attention, Volodymyr Mnih et al. (2014). This work translates the concepts of sequential gaze deployment and foveated visual attention reviewed in the survey into a differentiable recurrent neural network framework.
  • Paper: Show, Attend and Tell: Neural Image Caption Generation with Visual Attention, Kelvin Xu et al. (2015). This foundational paper applies task-driven, top-down visual attention mechanisms to neural image caption generation by dynamically conditioning spatial focus on generated words.
  • Paper: Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering, Peter Anderson et al. (2018). This paper directly operationalizes the survey's theoretical distinction between bottom-up saliency and top-down task guidance within multimodal vision-language architectures.
  • Paper: Residual Attention Network for Image Classification, Fei Wang et al. (2017). This work embeds bottom-up and top-down attention mechanisms directly inside deep feedforward convolutional networks for image classification.
  • Paper: CBAM: Convolutional Block Attention Module, Sanghyun Woo et al. (2018). This study applies spatial and channel attention modules as lightweight components to refine intermediate feature maps across standard vision backbones.
  • Paper: Sanity Checks for Saliency Maps, Julius Adebayo et al. (2018). This paper establishes critical evaluation methodology by demonstrating that popular saliency methods often fail basic sanity checks regarding model parameters and training data.
  • Paper: Transformers in Vision: A Survey, Salman Khan et al. (2021). This survey examines the subsequent paradigm shift in computer vision where self-attention mechanisms replace convolutional feature extractors altogether.
Cover for State-of-the-Art in Visual Attention Modeling

Abstract

Modeling visual attention—particularly stimulus-driven, saliency-based attention—has been a very active research area over the past 25 years. Many different models of attention are now available which, aside from lending theoretical contributions to other fields, have demonstrated successful applications in computer vision, mobile robotics, and cognitive systems. Here we review, from a computational perspective, the basic concepts of attention implemented in these models. We present a taxonomy of nearly 65 models, which provides a critical comparison of approaches, their capabilities, and shortcomings. In particular, 13 criteria derived from behavioral and computational studies are formulated for qualitative comparison of attention models. Furthermore, we address several challenging issues with models, including biological plausibility of the computations, correlation with eye movement datasets, bottom-up and top-down dissociation, and constructing meaningful performance measures. Finally, we highlight current research trends in attention modeling and provide insights for future.

Table of Contents

  • 1 INTRODUCTION
  • 1.1 Definitions
  • 1.2 Origins
  • 1.3 Empirical Foundations
  • 1.4 Applications
  • 1.5 Statement and Organization
  • 2 CATEGORIZATION FACTORS
  • 2.1 Bottom-Up versus Top-Down Models
  • 2.1.1 Object Features
  • 2.1.2 Scene Context
  • 2.1.3 Task Demands
  • 2.2 Spatial versus Spatio-Temporal Models
  • 2.3 Overt versus Covert Attention
  • 2.4 Space-Based versus Object-Based Models
  • 2.5 Features
  • 2.6 Stimuli and Task Type
  • 2.7 Evaluation Measures
  • 2.8 Data Sets
  • 3 ATTENTION MODELS
  • 3.1 Cognitive Models (C)
  • 3.2 Bayesian Models (B)
  • 3.3 Decision Theoretic Models (D)
  • 3.4 Information Theoretic Models (I)
  • 3.5 Graphical Models (G)
  • 3.6 Spectral Analysis Models (S)
  • 3.7 Pattern Classification Models (P)
  • 3.8 Other Models (O)
  • 4 DISCUSSION
  • 5 SUMMARY AND CONCLUSION
  • ACKNOWLEDGMENTS
  • REFERENCES

Knowls

  1. Knowl 1 — Statistical Unification of Saliency Modeling Approaches

    theoretical result

    Computational visual attention and saliency models can be unified into three probabilistic classes based on the quantities they estimate:

    1. Bottom-Up Self-Information Models: Saliency is formulated as the rarity or self-information of local visual features xx, expressed as 1P(x)\frac{1}{P(x)}, −log⁡P(x)-\log P(x), or local entropy EX[−log⁡P(x)]\mathbb{E}_X[-\log P(x)]. Different models correspond to specific assumptions regarding P(x)P(x) (e.g., Gaussian distributions or non-parametric kernel density estimates).

    2. Top-Down Target-Conditioned Likelihood Models: Saliency is defined by the likelihood of features given target presence, log⁡P(x∣Y=1)\log P(x \mid Y = 1), where Y=1Y = 1 denotes the presence of a visual search target.

    3. Discriminant / Posterior Probability Models: Saliency is computed as the posterior target probability P(Y=1∣X)P(Y = 1 \mid X) or the log-likelihood ratio (log-odds) log⁡P(x∣Y=1)P(x∣Y=0)\log \frac{P(x \mid Y = 1)}{P(x \mid Y = 0)}.

    Under this unification, the Bayesian Saliency Using Natural Statistics (SUN) formulation:

    log⁡sz=−log⁡P(F=fz)+log⁡P(F=fz∣C=1)+log⁡P(C=1∣L=lz)\log s_z = -\log P(F = f_z) + \log P(F = f_z \mid C = 1) + \log P(C = 1 \mid L = l_z)

    (where FF is the feature, CC denotes target presence, and LL denotes pixel location) is a special case of the decision-theoretic discriminant saliency model under two assumptions:

    • The feature distribution in the absence of the target matches the feature distribution of natural scenes, P(F=fz∣C=0)=P(F=fz)P(F = f_z \mid C = 0) = P(F = f_z).
    • The absence of the target is uniformly distributed across all spatial image locations, P(C=0∣L=lz)=KP(C = 0 \mid L = l_z) = K for constant KK.

    Graphical models extend these three classes by introducing spatial and temporal dependency structures (such as Markov transition matrices or conditional random field clique potentials) between neighboring locations.

  2. Knowl 2 — General Problem Formulation of Visual Attention Modeling

    definition

    Let a set of NN images I={Ii}i=1N\mathcal{I} = \{I_i\}_{i=1}^N be viewed by KK human observers. For each observer k∈{1,…,K}k \in \{1, \dots, K\} viewing image IiI_i, let the recorded overt visual scanpath be represented by the vector of eye fixations:

    Lik={(pijk,tijk)}j=1nikL_i^k = \{(p_{ij}^k, t_{ij}^k)\}_{j=1}^{n_i^k}

    where pijk=(xijk,yijk)∈R2p_{ij}^k = (x_{ij}^k, y_{ij}^k) \in \mathbb{R}^2 denotes the 2D spatial coordinate of the jj-th fixation, tijkt_{ij}^k is its occurrence timestamp, and nikn_i^k is the total number of fixations made by observer kk on image IiI_i.

    The computational goal of visual attention modeling (specifically for stimulus-driven, overt attention) is to find a mapping function f∈Ff \in \mathcal{F} that maps an input visual stimulus IiI_i to a 2D scalar saliency map while minimizing the prediction error with respect to human fixations under a distance metric m∈Mm \in \mathcal{M}:

    min⁡f∈F∑k=1K∑i=1Nm(f(Iik),Lik)\min_{f \in \mathcal{F}} \sum_{k=1}^K \sum_{i=1}^N m\left(f(I_i^k), L_i^k\right)

  3. Knowl 3 — Taxonomy of Computational Visual Attention Models

    model/method

    Computational models of visual attention can be categorized across 13 behavioral and computational dimensions:

    • Attentional Pathway: Bottom-up (stimulus-driven, fast, involuntary, exogenous) versus Top-down (task-driven, slow, goal-directed, endogenous).
    • Temporal Horizon: Spatial-only (operating on isolated frames) versus Spatio-temporal (incorporating temporal filtering, motion channels, or action history).
    • Stimulus Domain: Static (still images) versus Dynamic (video sequences, interactive simulations); Synthetic (search arrays, Gabor grids) versus Natural scenes.
    • Task Paradigms: Free-viewing, Visual search (target-directed), or Interactive/task-driven environments (e.g., driving, block manipulation).
    • Representational Unit: Space-based (pixel/spatial grid coordinates) versus Object-/Feature-based (segmented regions, proto-objects, or feature detector tuning).
    • Extracted Feature Cues: Low-level channels (intensity contrast, color opponency, multi-scale orientation), motion/flicker contrasts, mid-level geometric cues (symmetry, junctions, curvature), and high-level semantic channels (faces, text, scene gist).
    • Core Computational Families:
      1. Cognitive models: Implementations rooted in the primate visual pathway (e.g., center-surround pyramids, lateral cortical inhibition).
      2. Bayesian models: Probabilistic integration of sensory likelihoods with scene gist or spatial priors.
      3. Decision-theoretic models: Saliency formulated as optimal hypothesis testing minimizing classification error.
      4. Information-theoretic models: Saliency formulated as maximizing information gain or local Shannon self-information.
      5. Graphical models: Probabilistic graphical representations (HMMs, DBNs, CRFs) capturing spatial and temporal fixation dynamics.
      6. Spectral analysis models: Frequency-domain filtering of Fourier or Quaternion spectra.
      7. Pattern classification models: Supervised machine learning architectures (e.g., SVMs) trained on human gaze or region annotations.
      8. Reinforcement learning / Behavioral models: Action-perception coordination frameworks optimizing behavioral reward.
  4. Knowl 4 — Decision-Theoretic Discriminant Saliency Formulation

    equation

    In the decision-theoretic framework, visual saliency is defined as optimal classification between a center window and its surrounding context to minimize classification error. Given a feature vector F={F1,…,Fd}F = \{F_1, \dots, F_d\} at spatial location ll, let Cl=1C_l = 1 denote samples drawn from the central region (foreground/target) and Cl=0C_l = 0 denote samples drawn from the surrounding annular region (background/distractors). The mutual information I(F;C)=∑i=1dI(Fi;C)I(F; C) = \sum_{i=1}^d I(F_i; C) is used to quantify discriminative power.

    The posterior probability of target presence given feature observation Fl=fzF_l = f_z is formulated as:

    P(Cl=1∣Fl=fz)=σ(log⁡P(Fl=fz∣Cl=1)P(Cl=1)P(Fl=fz∣Cl=0)P(Cl=0))P(C_l = 1 \mid F_l = f_z) = \sigma\left( \log \frac{P(F_l = f_z \mid C_l = 1) P(C_l = 1)}{P(F_l = f_z \mid C_l = 0) P(C_l = 0)} \right)

    where σ(x)=(1+e−x)−1\sigma(x) = (1 + e^{-x})^{-1} is the sigmoid logistic function.

    Applying feature independence assumptions and spatial location priors L=lzL = l_z, the log-likelihood ratio inside the sigmoid expands to:

    log⁡P(Fl=fz∣Cl=1)P(Fl=fz∣Cl=0)+log⁡P(Cl=1∣L=lz)P(Cl=0∣L=lz)\log \frac{P(F_l = f_z \mid C_l = 1)}{P(F_l = f_z \mid C_l = 0)} + \log \frac{P(C_l = 1 \mid L = l_z)}{P(C_l = 0 \mid L = l_z)}

  5. Knowl 5 — Evaluation Metrics for Saliency Prediction: Symmetric KL Divergence and NSS

    equation

    Two standard distribution- and point-based metrics used to evaluate estimated saliency maps (ESMs) against human eye fixation data are:

    Symmetric Kullback-Leibler (KL) Divergence: Let ti=1,…,Nt_i = 1, \dots, N be human saccade locations. The model's ESM is sampled at human fixation locations xi,humanx_{i,\text{human}} and at uniformly distributed random image locations xi,randomx_{i,\text{random}}, normalized to [0,1][0, 1], and binned into qq histogram bins. Letting HkH_k and RkR_k denote the normalized sample frequencies in bin kk for human and random fixations respectively, the symmetric KL divergence is:

    KL=12∑k=1q(Hklog⁡HkRk+Rklog⁡RkHk)KL = \frac{1}{2} \sum_{k=1}^q \left( H_k \log \frac{H_k}{R_k} + R_k \log \frac{R_k}{H_k} \right)

    Higher KL divergence reflects stronger predictive separation. KL divergence is invariant to any continuous monotonic transformation applied to the saliency map values.

    Normalized Scanpath Saliency (NSS): Given an estimated saliency map SS with spatial mean μS\mu_S and standard deviation σS\sigma_S, the NSS at a human eye fixation coordinate (xh,yh)(x_h, y_h) is computed as:

    NSS=S(xh,yh)−μSσSNSS = \frac{S(x_h, y_h) - \mu_S}{\sigma_S}

    An NSS score of 11 denotes that fixated locations exhibit saliency values one standard deviation above the scene average, while NSS≤0NSS \le 0 indicates performance equal to or worse than random selection. Unlike KL divergence, NSS is sensitive to linear and non-linear monotonic reparameterizations.

  6. Knowl 6 — Evaluation Metrics for Saliency Prediction: AUC, CC, and String Editing Distance

    model/method

    Saliency models are quantitatively evaluated against eye tracking ground truth using receiver operating characteristics, correlation coefficients, or sequence alignment:

    • Area Under the ROC Curve (AUC): The estimated saliency map (ESM) is treated as a continuous binary classifier to separate fixated pixels (ground-truth positives) from non-fixated pixels (negatives). Sweeping a decision threshold across all normalized saliency values traces a Receiver Operating Characteristic (ROC) curve of true positive rate versus false positive rate. AUC scores range from 0.50.5 (chance) to 1.01.0 (perfect prediction) and remain invariant under any monotonic increasing transformation of saliency values.

    • Linear Correlation Coefficient (CC): Measures the linear association between the continuous ESM SS and a ground-truth fixation map GG (constructed by convolving discrete human fixation locations with a 2D Gaussian kernel):

    CC(G,S)=∑x,y(G(x,y)−μG)(S(x,y)−μS)σG2σS2CC(G, S) = \frac{\sum_{x,y} (G(x, y) - \mu_G) (S(x, y) - \mu_S)}{\sqrt{\sigma_G^2 \sigma_S^2}}

    where μG,μS\mu_G, \mu_S and σG2,σS2\sigma_G^2, \sigma_S^2 are the respective spatial means and variances. CC yields a value in [−1,+1][-1, +1], where +1+1 represents a perfect linear match.

    • String Editing Distance: Evaluates the sequential temporal order of saccadic scanpaths. Fixation locations from human gaze and saliency-guided trajectories are clustered into discrete Regions of Interest (ROIs) labeled with unique characters, producing strings such as stringh\text{string}_h and strings\text{string}_s. The similarity index is computed as:

    Similarity=1−Ss∣strings∣\text{Similarity} = 1 - \frac{S_s}{|\text{string}_s|}

    where SsS_s is the minimum Levenshtein edit distance computed using unit costs for insertion, deletion, and character substitution.

  7. Knowl 7 — Saliency Evaluation Artifacts: Border Effects, Center-Bias, and Unshuffled AUC

    empirical result

    Standard saliency benchmark metrics are susceptible to boundary handling and spatial distribution biases:

    1. Edge and Border Effects: Zero-padding or invalid filter response artifacts at image boundaries artificially inflate scores. On a dummy, completely uniform saliency map (all ones):

      • A border width of 0 pixels yields an ROC AUC of 0.500.50 and a KL divergence of 0.000.00.
      • A 4-pixel black boundary increases the dummy ROC AUC to 0.620.62 and KL divergence to 0.120.12.
      • An 8-pixel black boundary increases the dummy ROC AUC to 0.730.73 and KL divergence to 0.250.25.
    2. Central Fixation Bias: Human observers fixate disproportionately near the center of images due to viewing geometry and photographer bias. As a result, a static, baseline 2D Gaussian blob placed at the image center achieves higher AUC and correlation scores on standard datasets (such as Bruce-Tsotsos and CRCN-ORIG) than many computational saliency models.

    3. Unshuffled AUC Metric: To neutralize center-bias and border artifacts during ROC computation, the positive sample set for a given image consists of all human fixations recorded on that image, while the negative sample set is constructed from the union of all human fixations recorded across all other images in the dataset (excluding fixations that coincide with the positive set). This preserves the global empirical fixation distribution in the negative baseline.

  8. Knowl 8 — Spectral Domain Saliency Computation Models

    model/method

    Spectral analysis saliency models compute visual conspicuity in the frequency domain by detecting statistical singularities or suppressing redundant spectral energy:

    • Spectral Residual (SR) Model: For an input image I(x)I(x), the Fourier transform F[I(x)]\mathcal{F}[I(x)] yields amplitude spectrum A(f)=∣F[I(x)]∣A(f) = |\mathcal{F}[I(x)]| and phase spectrum P(f)=∠F[I(x)]P(f) = \angle \mathcal{F}[I(x)]. The log spectrum is computed as L(f)=log⁡(A(f))\mathcal{L}(f) = \log(A(f)). The spectral residual R(f)\mathcal{R}(f) represents anomalous frequency components obtained by subtracting a smoothed log spectrum:

    R(f)=L(f)−hn(f)∗L(f)\mathcal{R}(f) = \mathcal{L}(f) - h_n(f) * \mathcal{L}(f)

    where hn(f)h_n(f) is an n×nn \times n local averaging spatial filter. The spatial saliency map S(x)S(x) is synthesized via the inverse Fourier transform F−1\mathcal{F}^{-1} combining the spectral residual with the original phase spectrum, followed by Gaussian smoothing filter g(x)g(x):

    S(x)=g(x)∗∣F−1[exp⁡(R(f)+iP(f))]∣2S(x) = g(x) * \left| \mathcal{F}^{-1}\left[\exp\left(\mathcal{R}(f) + i P(f)\right)\right] \right|^2

    • Phase Spectrum of Quaternion Fourier Transform (PQFT): Extends frequency analysis to spatio-temporal video data by encoding color opponency, intensity, and motion channels into a quaternion matrix prior to computing the phase spectrum transform.

    • Spectral Whitening (SW): Computes a windowed Fourier transform f(u,v)=F[w(I(x,y))]f(u,v) = \mathcal{F}[w(I(x,y))] and flattens the spectrum via normalization n(u,v)=f(u,v)∥f(u,v)∥n(u,v) = \frac{f(u,v)}{\|f(u,v)\|}. The saliency map is recovered as S(x,y)=g(u,v)∗∣F−1[n(u,v)]∣2S(x,y) = g(u,v) * \left| \mathcal{F}^{-1}[n(u,v)] \right|^2, which suppresses uniform background motion and redundant textures.

  9. Knowl 9 — Information Maximization (AIM) and Self-Resemblance (SDSR) Saliency Models

    model/method

    Information-theoretic models estimate saliency by evaluating the statistical unlikelihood of local image features relative to their surrounding context:

    • Attention based on Information Maximization (AIM): Saliency is formulated as the Shannon self-information of a local image patch XX, given by I(X)=−log⁡p(X)I(X) = -\log p(X). To make the probability density estimation tractable over M×NM \times N RGB patches, Independent Component Analysis (ICA) basis filters WW learned from natural image patches are applied to project the patch into independent 1D feature responses y=WXy = W X. Saliency is computed as the sum of log-likelihoods over 1D non-parametric density estimates pi(yi)p_i(y_i):

    I(X)=−∑i=1dlog⁡pi(wiTX)I(X) = -\sum_{i=1}^d \log p_i(w_i^T X)

    • Saliency Detection by Self-Resemblance (SDSR): Local image structure is represented via local regression steering kernels K(xl−xi)K(x_l - x_i) that compute local feature covariance matrices ClC_l. The saliency sis_i of a pixel feature matrix FiF_i relative to surrounding feature matrices FjF_j is computed as the inverse sum of matrix cosine similarities ρ(Fi,Fj)\rho(F_i, F_j):

    si=[∑j=1Nexp⁡(−1+ρ(Fi,Fj)σ2)]−1s_i = \left[ \sum_{j=1}^N \exp\left( \frac{-1 + \rho(F_i, F_j)}{\sigma^2} \right) \right]^{-1}

    where σ\sigma is a local weighting parameter, penalizing features that frequently resemble their surrounding neighborhood.

  10. Knowl 10 — Methodological Limitations and Open Challenges in Visual Attention Modeling

    limitation

    Current visual attention models face four primary theoretical and empirical limitations:

    1. Lack of Top-Down Task-Driven Principles: The vast majority of published models operate purely feed-forward on low-level bottom-up cues. Principled computational frameworks that integrate time-varying cognitive task demands, internal mental states, and closed-loop behavioral reward during complex, interactive behaviors (e.g., driving or video game play) remain scarce.

    2. Overt versus Covert Attention Dissociation: Existing models validate predicted saliency exclusively against overt eye movements (foveal fixations). However, human visual attention includes covert shifts (mental deployment of attention without eye movements) that precede saccades or operate independently, which cannot be measured or validated using standard fixation datasets alone.

    3. Evaluation Biases and Metric Inconsistencies: Saliency models frequently diverge in ranking depending on the chosen metric (e.g., AUC vs. NSS vs. KL). Center-bias present in ground-truth eye-tracking datasets enables simple spatial center-priors (such as a fixed Gaussian) to outperform feature-based models unless specifically corrected by metrics like Unshuffled AUC.

    4. Biological Plausibility versus Black-Box Learning: Supervised pattern classification approaches (e.g., SVMs trained on large fixation sets) achieve high numerical fixation prediction scores by incorporating high-level detectors (faces, objects), but often function as data-dependent black-boxes that lack mechanistic or neurobiological plausibility.

Coverage note — Specific algorithmic implementation details of individual external software toolkits (such as iNVT, STB, VOCUS) and granular summaries of specialized application subfields (e.g., thumbnailing, retargeting) were omitted in favor of the core computational models, theoretical unification, mathematical formulations, and evaluation methodologies.

References

  1. 1.K. Koch, J. McLean, R. Segev, M.A. Freed, M.J. Berry, V. Balasubramanian, and P. Sterling, ‘‘How Much the Eye Tells the Brain,’’ Current Biology, vol. 25, nos. 16-14, pp. 1428-34, 2006.
  2. 2.L. Itti, ‘‘Models of Bottom-Up and Top-Down Visual Attention,’’ PhD thesis, California Inst. of Technology, 2000.
  3. 3.D.J. Simons and D.T. Levin, ‘‘Failure to Detect Changes to Attended Objects,’’ Investigative Ophthalmology and Visual Science, vol. 38, no. 4, p. 3273, 1997.
  4. 4.R.A. Rensink, ‘‘How Much of a Scene Is Seen—The Role of Attention in Scene Perception,’’ Investigative Ophthalmology and Visual Science, vol. 38, p. 707, 1997.
  5. 5.D.J. Simons and C.F. Chabris, ‘‘Gorillas in Our Midst: Sustained Inattentional Blindness for Dynamic Events,’’ Perception, vol. 28, no. 9, pp. 1059-1074, 1999.
  6. 6.J.E. Raymond, K.L. Shapiro, and K.M. Arnell, ‘‘Temporary Suppression of Visual Processing in an RSVP Task: An Attentional Blink?’’ J. Experimental Psychology, vol. 18, no 3, pp. 849-60, 1992.
  7. 7.S. Treue and J.H.R. Maunsell, ‘‘Attentional Modulation of Visual Motion Processing in Cortical Areas MT and MST,’’ Nature, vol. 382, pp. 539-541, 1996.
  8. 8.S. Frintrop, E. Rome, and H.I. Christensen, ‘‘Computational Visual Attention Systems and Their Cognitive Foundations: A Survey,’’ ACM Trans. Applied Perception, vol. 7, no. 1, Article 6, 2010.
  9. 9.A. Rothenstein and J. Tsotsos, ‘‘Attention Links Sensing to Recognition,’’ J. Image and Vision Computing, vol. 26, pp. 114-126, 2006.
  10. 10.R. Desimone and J. Duncan, ‘‘Neural Mechanisms of Selective Visual Attention,’’ Ann. Rev. Neuroscience, vol. 18, pp. 193-222, 1995.
  11. 11.S.J. Luck, L. Chelazzi, S.A. Hillyard, and R. Desimone, ‘‘Neural Mechanisms of Spatial Selective Attention in Areas V1, V2, and V4 of Macaque Visual Cortex,’’ J. Neurophysiology, vol. 77, pp. 24-42, 1997.
  12. 12.C. Bundesen and T. Habekost, ‘‘Attention,’’ Handbook of Cognition, K. Lamberts and R. Goldstone, eds., 2005.
  13. 13.V. Navalpakkam, C. Koch, A. Rangel, and P. Perona, ‘‘Optimal Reward Harvesting in Complex Perceptual Environments,’’ Proc. Nat’l Academy of Sciences USA, vol. 107, no. 11, pp. 5232-5237, 2010.
  14. 14.L. Itti, C. Koch, and E. Niebur, ‘‘A Model of Saliency-Based Visual Attention for Rapid Scene Analysis,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 20, no. 11, pp. 1254-1259, Nov. 1998.
  15. 15.J.K. Tsotsos, S.M. Culhane, W.Y.K. Wai, Y. Lai, N. Davis, and F. Nuflo, ‘‘Modeling Visual Attention via Selective Tuning,’’ Artificial Intelligence, vol. 78, nos. 1-2, pp. 507-545, 1995.
  16. 16.R. Milanese, ‘‘Detecting Salient Regions in an Image: From Biological Evidence to Computer Implementation,’’ PhD thesis, Univ. Geneva, 1993.
  17. 17.S. Baluja and D. Pomerleau, ‘‘Using a Saliency Map for Active Spatial Selective Attention: Implementation & Initial Results,’’ Proc. Advances in Neural Information Processing Systems, pp. 451- 458, 1994.
  18. 18.C. Koch and S. Ullman, ‘‘Shifts in Selective Visual Attention: Towards the Underlying Neural Circuitry,’’ Human Neurobiology, vol. 4, no. 4, pp. 219-227, 1985.
  19. 19.K. Rayner, ‘‘Eye Movements in Reading and Information Proces- sing: 20 Years of Research,’’ Psychological Bull., vol. 134, pp. 372- 422, 1998.
  20. 20.J. Najemnik and W.S. Geisler, ‘‘Optimal Eye Movement Strategies in Visual Search,’’ Nature, vol. 434, pp. 387-391, 2005.
  21. 21.L.W. Renninger, J.M. Coughlan, P. Verghese, and J. Malik, ‘‘An Information Maximization Model of Eye Movements,’’ Advances in Neural Information Processing Systems, vol. 17, pp. 1121-1128, 2005.
  22. 22.U. Rutishauser and C. Koch, ‘‘Probabilistic Modeling of Eye Movement Data during Conjunction Search via Feature-Based Attention,’’ J. Vision, vol. 7, no. 6, pp. 1-20, 2007.
  23. 23.R. Rao, G. Zelinsky, M. Hayhoe, and D. Ballard, ‘‘Eye Movements in Iconic Visual Search,’’ Vision Research, vol. 42, pp. 1447-1463, 2002.
  24. 24.A.T. Duchowski, ‘‘A Breadth-First Survey of Eye-Tracking Applications,’’ Behavior Research Methods Instruments Computers J. Psychonomic Soc. Inc., vol. 34, pp. 455-470, 2002.
  25. 25.G.E. Legge, T.S. Klitz, and B. Tjan, ‘‘Mr. Chips: An Ideal-Observer Model of Reading,’’ Psychological Rev., vol. 104, pp. 524-553, 1997.
  26. 26.R.D. Rimey and C.M. Brown, ‘‘Controlling Eye Movements with Hidden Markov Models,’’ Int’l J. Computer Vision, vol. 7, no. 1, pp. 47-65, 1991.
  27. 27.S. Treue, ‘‘Neural Correlates of Attention in Primate Visual Cortex,’’ Trends in Neurosciences, vol. 24, no. 5, pp. 295-300, 2001.
  28. 28.S. Kastner and L.G. Ungerleider, ‘‘Mechanisms of Visual Attention in the Human Cortex,’’ Ann. Rev. Neurosciences, vol. 23, pp. 315- 341, 2000.
  29. 29.E.T. Rolls and G. Deco, ‘‘Attention in Natural Scenes: Neurophy- siological and Computational Bases,’’ Neural Networks, vol. 19, no. 9, pp. 1383-1394, 2006.
  30. 30.G.A. Carpenter and S. Grossberg, ‘‘A Massively Parallel Architec- ture for a Self-Organizing Neural Pattern Recognition Machine,’’ J. Computer Vision, Graphics, and Image Processing, vol. 37, no. 1, pp. 54-115, 1987.
  31. 31.N. Ouerhani and H. Hügli, ‘‘Real-Time Visual Attention on a Massively Parallel SIMD Architecture,’’ Real-Time Imaging, vol. 9, no. 3, pp. 189-196, 2003.
  32. 32.Q. Ma, L. Zhang, and B. Wang, ‘‘New Strategy for Image and Video Quality Assessment,’’ J. Electronic Imaging, vol. 19, pp. 1-14, 2010.
  33. 33.Y. Ma, X. Hua, L. Lu, and H. Zhang, ‘‘A Generic Framework of User Attention Model and Its Application in Video Summariza- tion,’’ IEEE Trans. Multimedia, vol. 7, no. 5, pp. 907-919, Oct. 2005.
  34. 34.A. Ninassi, O. Le Meur, P. Le Callet, and D. Barba, ‘‘Does Where You Gaze on an Image Affect Your Perception of Quality? Applying Visual Attention to Image Quality Metric,’’ Proc. IEEE Int’l Conf. Image Processing, vol. 2, pp. 169-172, 2007.
  35. 35.D. Walther and C. Koch, ‘‘Modeling Attention to Salient Proto- Objects,’’ Neural Networks, vol. 19, no. 9, pp. 1395-1407, 2006.
  36. 36.C. Siagian and L. Itti, ‘‘Biologically Inspired Mobile Robot Vision Localization,’’ IEEE Trans. Robotics, vol. 25, no. 4, pp. 861-873, Aug. 2009.
  37. 37.S. Frintrop and P. Jensfelt, ‘‘Attentional Landmarks and Active Gaze Control for Visual SLAM,’’ IEEE Trans. Robotics, vol. 24, no. 5, pp. 1054-1065, Oct. 2008.
  38. 38.D. DeCarlo and A. Santella, ‘‘Stylization and Abstraction of Photographs,’’ ACM Trans. Graphics, vol. 21, no. 3, pp. 769-776, 2002.
  39. 39.L. Itti, ‘‘Automatic Foveation for Video Compression Using a Neurobiological Model of Visual Attention,’’ IEEE Trans. Image Processing, vol. 13, no. 10, pp. 1304-1318, Oct. 2004.
  40. 40.L. Marchesotti, C. Cifarelli, and G. Csurka, ‘‘A Framework for Visual Saliency Detection with Applications to Image Thumbnail- ing,’’ Proc. 12th IEEE Int’l Conf. Computer Vision, 2009.
  41. 41.O. Le Meur, P. Le Callet, D. Barba, and D. Thoreau, ‘‘A Coherent Computational Approach to Model Bottom-Up Visual Attention,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 28, no. 5, pp. 802-817, May 2006.
  42. 42.G. Fritz, C. Seifert, L. Paletta, and H. Bischof, ‘‘Attentive Object Detection Using an Information Theoretic Saliency Measure,’’ Proc. Second Int’l Conf. Attention and Performance in Computational Vision, pp. 29-41, 2005.
  43. 43.T. Liu, J. Sun, N.N Zheng, and H.Y Shum, ‘‘Learning to Detect a Salient Object,’’ Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2007.
  44. 44.V. Setlur, R. Raskar, S. Takagi, M. Gleicher, and B. Gooch, ‘‘Automatic Image Retargeting, In Mobile and Ubiquitous Multi- media (MUM),’’ Proc. Fourth Int’l Conf. Mobile and Ubiquitous Multimedia, 2005.
  45. 45.C. Chamaret and O. Le Meur, ‘‘Attention-Based Video Reframing: Validation Using Eye-Tracking,’’ Proc. 19th Int’l Conf. Pattern Recognition, 2008.
  46. 46.S. Goferman, L. Zelnik-Manor, and A. Tal, ‘‘Context-Aware Saliency Detection,’’ Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2010.
  47. 47.N. Sadaka and L.J. Karam, ‘‘Efficient Perceptual Attentive Super- Resolution,’’ Proc. 16th IEEE Int’l Conf. Image Processing, 2009.
  48. 48.H. Liu, S. Jiang, Q. Huang, and C. Xu, ‘‘A Generic Virtual Content Insertion System Based on Visual Attention Analysis,’’ Proc. ACM Int’l Conf. Multimedia, pp. 379-388, 2008.
  49. 49.S. Marat, M. Guironnet, and D. Pellerin, ‘‘Video Summarization Using a Visual Attention Model,’’ Proc. 15th European Signal Processing Conf., 2007.
  50. 50.S. Frintrop, VOCUS: A Visual Attention System for Object Detection and Goal-Directed Search. Springer, 2006.
  51. 51.V. Navalpakkam and L. Itti, ‘‘An Integrated Model of Top-Down and Bottom-Up Attention for Optimizing Detection Speed,’’ Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2006.
  52. 52.A. Salah, E. Alpaydin, and L. Akrun, ‘‘A Selective Attention-Based Method for Visual Pattern Recognition with Application to Handwritten Digit Recognition and Face Recognition,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 24, no. 3, pp. 420-425, Mar. 2002.
  53. 53.S. Frintrop, ‘‘General Object Tracking with a Component-Based Target Descriptor,’’ Proc. IEEE Int’l Conf. Robotics and Automation, pp. 4531-4536, 2010.
  54. 54.M.S. El-Nasr, T. Vasilakos, C. Rao, and J. Zupko, ‘‘Dynamic Intelligent Lighting for Directing Visual Attention in Interactive 3D Scenes,’’ IEEE Trans. Computational Intelligence and AI in Games, vol. 1, no. 2, pp. 145-153, June 2009.
  55. 55.G. Boccignone, ‘‘Nonparametric Bayesian Attentive Video Analy- sis,’’ Proc. 19th Int’l Conf. Pattern Recognition, 2008.
  56. 56.G. Boccignone, A. Chianese, V. Moscato, and A. Picariello, ‘‘Foveated Shot Detection for Video Segmentation,’’ IEEE Trans. Circuits and Systems for Video Technology, vol. 15, no. 3, pp. 365-377, Mar. 2005.
  57. 57.B. Mertsching, M. Bollmann, R. Hoischen, and S. Schmalz, ‘‘The Neural Active Vision System,’’ Handbook of Computer Vision and Applications, Academic Press, 1999.
  58. 58.A. Dankers, N. Barnes, and A. Zelinsky, ‘‘A Reactive Vision System: Active-Dynamic Saliency,’’ Proc. Int’l Conf. Vision Systems, 2007.
  59. 59.N. Ouerhani, A. Bur, and H. Hügli, ‘‘Visual Attention-Based Robot Self-Localization,’’ Proc. European Conf. Mobile Robotics, pp. 803- 813, 2005.
  60. 60.S. Baluja and D. Pomerleau, ‘‘Expectation-Based Selective Atten- tion for Visual Monitoring and Control of a Robot Vehicle,’’ Robotics and Autonomous Systems, vol. 22, nos. 3/4, pp. 329-344, 1997.
  61. 61.C. Scheier and S. Egner, ‘‘Visual Attention in a Mobile Robot,’’ Proc. Int’l Symp. Industrial Electronics, pp. 48-53, 1997.
  62. 62.C. Breazeal, ‘‘A Context-Dependent Attention System for a Social Robot,’’ Proc. 16th Int’l Joint Conf. Artificial Intelligence, pp. 1146- 1151, 1999.
  63. 63.G. Heidemann, R. Rae, H. Bekel, I. Bax, and H. Ritter, ‘‘Integrating Context-Free and Context-Dependent Attentional Mechanisms for Gestural Object Reference,’’ Machine Vision Application, vol. 16, no. 1, pp. 64-73, 2004.
  64. 64.G. Heidemann, ‘‘Focus-of-Attention from Local Color Symme- tries,’’ IEEE Trans Pattern Analysis and Machine Intelligence, vol. 26, no. 7, pp. 817-830, July 2004.
  65. 65.A. Belardinelli, ‘‘Salience Features Selection: Deriving a Model from Human Evidence,’’ PhD thesis, 2008.
  66. 66.Y. Nagai, ‘‘From Bottom-up Visual Attention to Robot Action Learning,’’ Proc. Eighth IEEE Int’l Conf. Development and Learning, 2009.
  67. 67.C. Muhl, Y. Nagai, and G. Sagerer, ‘‘On Constructing a Communicative Space in HRI,’’ Proc. 30th German Conf. Artificial Intelligence, 2007.
  68. 68.T. Liu, S.D. Slotnick, J.T. Serences, and S. Yantis, ‘‘Cortical Mechanisms of Feature-Based Intentional Control,’’ Cerebral Cortex, vol. 13, no. 12, pp. 1334-1343, 2003.
  69. 69.B.W. Hong and M. Brady, ‘‘A Topographic Representation for Mammogram Segmentation,’’ Proc. Medical Image Computing and Computer Assisted Intervention, pp. 730-737, 2003.
  70. 70.N. Parikh, L. Itti, and J. Weiland, ‘‘Saliency-Based Image Processing for Retinal Prostheses,’’ J. Neural Eng., vol 7, no 1, pp. 1-10, 2010.
  71. 71.O.R. Joubert, D. Fize, G.A. Rousselet, and M. Fabre-Thorpe, ‘‘Early Interference of Context Congruence on Object Processing in Rapid Visual Categorization of Natural Scenes,’’ J. Vision, vol. 8, no. 13, pp. 1-18, 2008.
  72. 72.H. Li and K.N. Ngan, ‘‘Saliency Model-Based Face Segmentation and Tracking in Head-and-Shoulder Video Sequences,’’ J. Vision Comm. and Image Representation, vol. 19, pp. 320-333, 2008.
  73. 73.N. Courty and E. Marchand, ‘‘Visual Perception Based on Salient Features,’’ Proc. Int’l Conf. Intelligent Robots and Systems, 2003.
  74. 74.F. Shic and B. Scassellati, ‘‘A Behavioral Analysis of Computa- tional Models of Visual Attention,’’ Int’l J. Computer Vision, vol. 73, pp. 159-177, 2007.
  75. 75.H.C. Nothdurft, ‘‘Salience of Feature Contrast,’’ Neurobiology of Attention, L. Itti, G. Rees, and J. K. Tsotsos, eds., Academic Press, 2005.
  76. 76.M. Corbetta and G.L. Shulman, ‘‘Control of Goal-Directed and Stimulus-Driven Attention in the Brain,’’ Natural Rev., vol. 3, no. 3, pp. 201-215, 2002.
  77. 77.L. Itti and C. Koch, ‘‘Computational Modeling of Visual Atten- tion,’’ Natural Rev. Neuroscience, vol. 2, no. 3, pp. 194-203, 2001.
  78. 78.H.E. Egeth and S. Yantis, ‘‘Visual Attention: Control, Representa- tion, and Time Course,’’ Ann. Rev. Psychologogy, vol. 48, pp. 269- 297, 1997.
  79. 79.A.L. Yarbus, Eye-Movements and Vision. Plenum Press, 1967.
  80. 80.V. Navalpakkam and L. Itti, ‘‘Modeling the Influence of Task on Attention,’’ Vision Research, vol. 45, no. 2, pp. 205-231, 2005.
  81. 81.A.M. Treisman and G. Gelade, ‘‘A Feature Integration Theory of Attention,’’ Cognitive Psychology, vol. 12, pp. 97-136, 1980.
  82. 82.J.M. Wolfe, ‘‘Guided Search 4.0: Current Progress with a Model of Visual Search,’’ Integrated Models of Cognitive Systems, W.D. Gray, ed., Oxford Univ. Press, 2007.
  83. 83.G.J. Zelinsky, ‘‘A Theory of Eye Movements during Target Acquisition,’’ Psychological Rev., vol. 115, no. 4, pp. 787-835, 2008.
  84. 84.W. Einhauser, M. Spain, and P. Perona, ‘‘Objects Predict Fixations Better Than Early Saliency,’’ J. Vision, vol. 14, pp. 1-26, 2008.
  85. 85.M. Pomplun, ‘‘Saccadic Selectivity in Complex Visual Search Displays,’’ Vision Research, vol. 46, pp. 1886-1900, 2006.
  86. 86.A. Hwang and M. Pomplun, ‘‘A Model of Top-Down Control of Attention during Visual Search in Real-World Scenes,’’ J. Vision, vol. 8, no. 6, Article 681, 2008.
  87. 87.K. Ehinger, B. Hidalgo-Sotelo, A. Torralba, and A. Oliva, ‘‘Modeling Search for People in 900 Scenes: A Combined Source Model of Eye Guidance,’’ Visual Cognition, vol. 17, pp. 945-978, 2009.
  88. 88.A. Borji, M.N. Ahmadabadi, B.N. Araabi, and M. Hamidi, ‘‘Online Learning of Task-Driven Object-Based Visual Attention Control,’’ J. Image and Vision Computing, vol. 28, pp. 1130-1145, 2010.
  89. 89.A. Borji, M.N. Ahmadabadi, and B.N. Araabi, ‘‘Cost-Sensitive Learning of Top-Down Modulation for Attentional Control,’’ Machine Vision and Applications, vol. 22, pp. 61-76, 2011.
  90. 90.L. Elazary and L. Itti, ‘‘A Bayesian Model for Efficient Visual Search and Recognition,’’ Vision Research, vol. 50, pp. 1338-1352, 2010.
  91. 91.M.M. Chun and Y. Jiang, ‘‘Contextual Cueing: Implicit Learning and Memory of Visual Context Guides Spatial Attention,’’ Cognitive Psychology, vol. 36, pp. 28-71, 1998.
  92. 92.A. Torralba, ‘‘Modeling Global Scene Factors in Attention,’’ J. Optical Soc. Am., vol. 20, no. 7, pp. 1407-1418, 2003.
  93. 93.A. Oliva and A. Torralba, ‘‘Modeling the Shape of the Scene: A Holistic Representation of the Spatial Envelope,’’ Int’l J. Computer Vision, vol. 42, pp. 145-175, 2001.
  94. 94.L.W. Renninger and J. Malik, ‘‘When Is Scene Recognition Just Texture Recognition?’’ Vision Research, vol. 44, pp. 2301-2311, 2004.
  95. 95.C. Siagian and L. Itti, ‘‘Rapid Biologically-Inspired Scene Classification Using Features Shared with Visual Attention,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 29, no. 2, pp. 300-312, Feb. 2007.
  96. 96.M. Viswanathan, C. Siagian, and L. Itti, Vision Science Symp., 2007.
  97. 97.J. Triesch, D.H. Ballard, M.M. Hayhoe, and B.T. Sullivan, ‘‘What You See Is What You Need,’’ J. Vision, vol. 3, pp. 86-94, 2003.
  98. 98.M.I. Posner, ‘‘Orienting of Attention,’’ Quarterly J. Experimental Psychology, vol. 32, pp. 3-25, 1980.
  99. 99.M. Hayhoe and D. Ballard, ‘‘Eye Movements in Natural Behavior,’’ Trends in Cognitive Sciences, vol. 9, pp. 188-194, 2005.
  100. 100.M.S. Mirian, M.N. Ahmadabadi, B.N. Araabi, R.R. Siegwart, ‘‘Learning Active Fusion of Multiple Experts’ Decisions: An Attention-Based Approach,’’ Neural Computation, 2011.
  101. 101.R.J. Peters and L. Itti, ‘‘Beyond Bottom-up: Incorporating Task- dependent Influences into a Computational Model of Spatial Attention,’’ Proc. IEEE Conf. Computer Vision and Pattern Recogni- tion, 2007.
  102. 102.D. Pang, A. Kimura, T. Takeuchi, J. Yamato, and K. Kashino, ‘‘A Stochastic Model of Selective Visual Attention with a Dynamic Bayesian Network,’’ Proc. IEEE Int’l Conf. Multimedia and Expo., 2008.
  103. 103.Y. Zhai and M. Shah, ‘‘Visual Attention Detection in Video Sequences Using Spatiotemporal Cues,’’ Proc. ACM Int’l Conf. Multimedia, 2006.
  104. 104.S. Marat, T. Ho-Phuoc, L. Granjon, N. Guyader, D. Pellerin, and A. Guérin-Dugué, ‘‘Modeling Spatio-Temporal Saliency to Predict Gaze Direction for Short Videos,’’ Int’l J. Computer Vision, vol. 82, pp. 231-243, 2009.
  105. 105.V. Mahadevan and N. Vasconcelos, ‘‘Spatiotemporal Saliency in Dynamic Scenes,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 32, no. 1, pp. 171-177, Jan. 2010.
  106. 106.V. Mahadevan and N. Vasconcelos, ‘‘Saliency Based Discriminant Tracking,’’ Proc. IEEE Conf. Computer Vision and Pattern Recogni- tion, 2009.
  107. 107.N. Jacobson, Y-L. Lee, V. Mahadevan, N. Vasconcelos, and T.Q. Nguyen, ‘‘A Novel Approach to FRUC Using Discriminant Saliency and Frame Segmentation,’’ IEEE Trans. Image Processing, vol. 19, no. 11, pp. 2924-2934, Nov. 2010.
  108. 108.H.J. Seo and P. Milanfar, ‘‘Static and Space-Time Visual Saliency Detection by Self-Resemblance,’’ J. Vision, vol. 9, no. 12, pp. 1-27, 2009.
  109. 109.N. Sprague and D.H. Ballard, ‘‘Eye Movements for Reward Maximization,’’ Proc. Advances in Neural Information Processing, 2003.
  110. 110.http://tcts.fpms.ac.be / mousetrack/ 2012.
  111. 111.J. Bisley and M. Goldberg, ‘‘Neuronal Activity in the Lateral Intraparietal Area and Spatial Attention,’’ Science, vol. 299, pp. 81- 86, 2003.
  112. 112.J. Duncan, ‘‘Selective Attention and the Organization of Visual Information,’’ J. Experimental Psychology, vol. 113, pp. 501-517, 1984.
  113. 113.B.J. Scholl, ‘‘Objects and Attention: The State of the Art,’’ Cognition, vol. 80, pp. 1-46, 2001.
  114. 114.Z.W. Pylyshyn and R.W. Storm, ‘‘Tracking Multiple Independent Targets: Evidence for a Parallel Tracking Mechanism,’’ Spatial Vision, vol. 3, pp. 179-197, 1988.
  115. 115.E. Awh and H. Pashler, ‘‘Evidence for Split Attentional Foci,’’ J. Experimental Psychology Human Perception and Performance, vol. 26, pp. 834-846, 2000.
  116. 116.B.C. Russell, A. Torralba, K.P. Murphy, and W.T. Freeman, ‘‘LabelMe: A Database and Web-Based Tool for Image Annota- tion,’’ Int’l J. Computer Vision, vol. 77, nos. 1-3, pp. 157-173, 2008.
  117. 117.Y. Sun and R. Fisher, ‘‘Object-Based Visual Attention for Computer Vision,’’ Artificial Intelligence, vol. 146, no. 1, pp. 77- 123, 2003.
  118. 118.J.M. Wolfe and T.S. Horowitz, ‘‘What Attributes Guide the Deployment of Visual Attention and How Do They Do It?’’ Natural Rev. Neuroscience, vol. 5, pp. 1-7, 2004.
  119. 119.L. Itti, N. Dhavale, and F. Pighin, ‘‘Realistic Avatar Eye and Head Animation Using a Neurobiological Model of Visual Attention,’’ Proc. SPIE, vol. 5200, pp. 64-78, 2003.
  120. 120.R. Rae, ‘‘Gestikbasierte Mensch-Maschine-Kommunikation auf der Grundlage Visueller Aufmerksamkeit und Adaptivität,’’ PhD thesis, Universität Bielefeld, 2000.
  121. 121.J. Harel, C. Koch, and P. Perona, ‘‘Graph-Based Visual Saliency,’’ Neural Information Processing Systems, vol. 19, pp. 545-552, 2006.
  122. 122.O. Boiman and M. Irani, ‘‘Detecting Irregularities in Images and in Video,’’ Proc. IEEE Int’l Conf. Computer Vision, 2005.
  123. 123.B.W. Tatler, ‘‘The Central Fixation Bias in Scene Viewing: Selecting an Optimal Viewing Position Independently of Motor Bases and Image Feature Distributions,’’ J. Vision, vol. 14, pp. 1-17, 2007.
  124. 124.R. Milanese, ‘‘Detecting Salient Regions in an Image: From Biological Evidence to Computer Implementation,’’ PhD thesis, Univ. Geneva, 1993.
  125. 125.F.H. Hamker, ‘‘The Emergence of Attention by Poulation-based Inference and Its Role in Distributed Processing and Cognitive Control of Vision,’’ J. Computer Vision Image Understanding, vol. 100, nos. 1/2, pp. 64-106, 2005.
  126. 126.S. Vijayakumar, J. Conradt, T. Shibata, and S. Schaal, ‘‘Overt Visual Attention For a Humanoid Robot,’’ Proc. IEEE/RSJ Int’l Conf. Intelligent Robots and Systems, 2001.
  127. 127.C.M. Privitera and L.W. Stark, ‘‘Algorithms for Defining Visual Regions-of-Interest: Comparison with Eye Fixations,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 22, no. 9, pp. 970-982, Sept. 2000.
  128. 128.K. Lee, H. Buxton, and J. Feng, ‘‘Selective Attention for Cue- guided Search Using a Spiking Neural Network,’’ Proc. Int’l Workshop Attention and Performance in Computer Vision, p. 5562, 2003.
  129. 129.T. Kadir and M. Brady, ‘‘Saliency, Scale and Image Description,’’ Int’l J. Computer Vision, vol. 45, no. 2, pp. 83-105, 2001.
  130. 130.A. Maki, P. Nordlund, and J.O. Eklundh, ‘‘Attentional Scene Segmentation: Integrating Depth and Motion,’’ Computer Vision and Image Understanding, vol. 78, no. 3, pp. 351-373, 2000.
  131. 131.D. Parkhurst, K. Law, and E. Niebur, ‘‘Modeling the Role of Salience in the Allocation of Overt Visual Attention,’’ Vision Research, vol. 42, nos. 1, pp. 107-123, 2002.
  132. 132.T.S. Horowitz and J.M. Wolfe, ‘‘Visual Search Has No Memory,’’ Nature, vol. 394, pp. 575-577, 1998.
  133. 133.J. Li, Y. Tian, T. Huang, and W. Gao, ‘‘Probabilistic Multi-Task Learning for Visual Saliency Estimation in Video,’’ Int’l J. Computer Vision, vol. 90, pp. 150-165, 2010.
  134. 134.R. Peters, A. Iyer, L. Itti, and C. Koch, ‘‘Components of Bottom-Up Gaze Allocation in Natural Images,’’ Vision Research, vol. 45, pp. 2397-2416, 2005.
  135. 135.M. Land and M. Hayhoe, ‘‘In What Ways Do Eye Movements Contribute to Everyday Activities?’’ Vision Research, vol. 41, pp. 3559-3565, 2001.
  136. 136.G. Kootstra, A. Nederveen, and B. de Boer, ‘‘Paying Attention to Symmetry,’’ Proc. British Machine Vision Conf., pp. 1115-1125, 2008.
  137. 137.D. Reisfeld, H. Wolfson, and Y. Yeshurun, ‘‘Context-Free Atten- tional Operators: The Generalized Symmetry Transform,’’ Int’l J. Computer Vision, vol. 14, no. 2, pp. 119-130, 1995.
  138. 138.O. Le Meur, P. Le Callet, and D. Barba, ‘‘Predicting Visual Fixations on Video Based on Low-Level Visual Features,’’ Vision Research, vol. 47/19, pp. 2483-2498, 2007.
  139. 139.D.D. Salvucci, ‘‘An Integrated Model of Eye Movements and Visual Encoding,’’ Cognitive Systems Research, vol. 1, pp. 201-220, 2001.
  140. 140.A. Oliva, A. Torralba, M.S. Castelhano, and J.M. Henderson, ‘‘Top- Down Control of Visual Attention in Object Detection,’’ Proc. Int’l Conf. Image Processing, pp. 253-256, 2003.
  141. 141.L. Zhang, M.H. Tong, T.K. Marks, H. Shan, and G.W. Cottrell, ‘‘SUN: A Bayesian Framework for Saliency Using Natural Statistics,’’ J. Vision, vol. 8, no. 32, pp. 1-20, 2008.
  142. 142.L. Zhang, M.H. Tong, and G.W. Cottrell, ‘‘SUNDAy: Saliency Using Natural Statistics for Dynamic Analysis of Scenes,’’ Proc. 31st Ann. Cognitive Science Soc. Conf., 2009.
  143. 143.N.D.B. Bruce and J.K. Tsotsos, ‘‘Spatiotemporal Saliency: Towards a Hierarchical Representation of Visual Saliency,’’ Proc. Int’l Workshop Attention in Cognitive Systems, 2008.
  144. 144.N.D.B. Bruce and J.K. Tsotsos, ‘‘Saliency Based on Information Maximization,’’ Proc. Advances in Neural Information Processing Systems, 2005.
  145. 145.L. Itti and P. Baldi, ‘‘Bayesian Surprise Attracts Human Atten- tion,’’ Proc. Advances in Neural Information Processing Systems, 2005.
  146. 146.D. Gao and N. Vasconcelos, ‘‘Discriminant Saliency for Visual Recognition from Cluttered Scenes,’’ Proc. Advances in Neural Information Processing Systems, 2004.
  147. 147.D. Gao, S. Han, and N. Vasconcelos, ‘‘Discriminant Saliency, the Detection of Suspicious Coincidences, and Applications to Visual Recognition,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 31, no. 6, pp. 989-1005, June 2009.
  148. 148.E. Gu, J. Wang, and N.I. Badler, ‘‘Generating Sequence of Eye Fixations Using Decision-Theoretic Attention Model,’’ Proc. Work- shop Attention and Performance in Computational Vision, pp. 277-29, 2007.
  149. 149.T.S. Lee and S. Yu, ‘‘An Information-Theoretic Framework for Understanding Saccadic Behaviors,’’ Proc. Advanced in Neural Processing Systems, 2000.
  150. 150.X. Hou and L. Zhang, ‘‘Saliency Detection: A Spectral Residual Approach,’’ Proc. IEEE Conf. Computer Vision and Pattern Recogni- tion, 2007.
  151. 151.X. Hou and L. Zhang, ‘‘Dynamic Visual Attention: Searching for Coding Length Increments,’’ Proc. Advances in Neural Information Processing Systems, pp. 681-688, 2008.
  152. 152.M. Mancas, ‘‘Computational Attention: Modelisation and Appli- cation to Audio and Image Processing,’’ PhD thesis, 2007.
  153. 153.T. Avraham and M. Lindenbaum, ‘‘Esaliency (Extended Saliency): Meaningful Attention Using Stochastic Image Modeling,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 32, no. 4, pp. 693-708, Apr. 2010.
  154. 154.S. Chikkerur, T. Serre, C. Tan, and T. Poggio, ‘‘What and Where: A Bayesian Inference Theory of Visual Attention,’’ Vision Research, vol. 55, pp. 2233-2247, 2010.
  155. 155.P. Verghese, ‘‘Visual Search and Attention: A Signal Detection Theory Approach,’’ Neuron, vol. 31, pp. 523-535, 2001.
  156. 156.C. Guo, Q. Ma, and L. Zhang, ‘‘Spatio-Temporal Saliency Detection Using Phase Spectrum of Quaternion Fourier Transform,’’ Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2008.
  157. 157.C. Guo and L. Zhang, ‘‘A Novel Multiresolution Spatiotemporal Saliency Detection Model and Its Applications in Image and Video Compression,’’ IEEE Trans. Image Processing, vol. 19, no. 1, pp. 185- 198, Jan. 2010.
  158. 158.R. Achanta, S.S. Hemami, F.J. Estrada, and S. Süsstrunk, ‘‘Frequency-Tuned Salient Region Detection,’’ Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2009.
  159. 159.P. Bian and L. Zhang, ‘‘Biological Plausibility of Spectral Domain Approach for Spatiotemporal Visual Saliency,’’ Proc. 15th Int’l Conf. Advances in Neuro-Information Processing, pp. 251-258, 2009.
  160. 160.A. Garcia-Diaz, X.R. Fdez-Vidal, X.M. Pardo, and R. Dosil, ‘‘Decorrelation and Distinctiveness Provide with Human-Like Saliency,’’ Proc. Advanced Concepts for Intelligent Vision Systems, pp. 343-354, 2009.
  161. 161.N.J. Butko and J.R. Movellan, ‘‘Optimal Scanning for Faster Object Detection,’’ Proc. IEEE Conf. Computer Vision and Pattern Recogni- tion, 2009.
  162. 162.S. Jodogne and J. Piater, ‘‘Closed-Loop Learning of Visual Control Policies,’’ J. Artificial Intelligence Research, vol. 28, pp. 349-391, 2007.
  163. 163.R. McCallum, ‘‘Reinforcement Learning with Selective Perception and Hidden State,’’ PhD thesis, 1996.
  164. 164.L. Paletta, G. Fritz, and C. Seifert, ‘‘Q-Learning of Sequential Attention for Visual Object Recognition from Informative Local Descriptors,’’ Proc. 22nd Int’l Conf. Machine Learning, pp. 649-656, 2005.
  165. 165.W. Kienzle, M.O. Franz, B. Schölkopf, and F.A. Wichmann, ‘‘Center-Surround Patterns Emerge as Optimal Predictors for Human Saccade Targets,’’ J. Vision, vol. 9, pp. 1-15, 2009.
  166. 166.T. Judd, K. Ehinger, F. Durand, and A. Torralba, ‘‘Learning to Predict Where Humans Look,’’ Proc. 12th IEEE Int’l Conf. Computer Vision, 2009.
  167. 167.M. Cerf, J. Harel, W. Einhäuser, and C. Koch, ‘‘Predicting Human Gaze Using Low-Level Saliency Combined with Face Detection,’’ Advances in Neural Information Processing Systems, vol. 20, pp. 241- 248, 2007.
  168. 168.O. Ramström and H.I. Christensen, ‘‘Visual Attention Using Game Theory,’’ Proc. Biologically Motivated Computer Vision Conf., pp. 462- 471, 2002.
  169. 169.P.L. Rosin, ‘‘A Simple Method for Detecting Salient Regions,’’ Pattern Recognition, vol. 42, no. 11, pp. 2363-2371, 2009.
  170. 170.Z. Li, ‘‘A Saliency Map in Primary Visual Cortex,’’ Trends in Cognitive Sciences, vol. 6, no. 1, pp. 9-16, 2002.
  171. 171.Y. Li, Y. Zhou, J. Yan, and J. Yang, ‘‘Visual Saliency Based on Conditional Entropy,’’ Proc. Ninth Asian Conf. Computer Vision, 2009.
  172. 172.S.W Ban, I. Lee, and M. Lee, ‘‘Dynamic Visual Selective Attention Model,’’ Neurocomputing, vol. 71, nos. 4-6, pp. 853-856, 2008.
  173. 173.M.T. López, M.A. Fernndez, A. Fernández-Caballero, J. Mira, and A.E. Delgado, ‘‘Dynamic Visual Attention Model in Image Sequences,’’ J. Image and Vision Computing, vol. 25, pp. 597-613, 2007.
  174. 174.U. Rajashekar, I. van der Linde, A.C. Bovik, and L.K. Cormack, ‘‘GAFFE: A Gaze-Attentive Fixation Finding Engine,’’ IEEE Trans. Image Processing, vol. 17, no. 4, pp. 564-573, Apr. 2008.
  175. 175.G. Boccignone and M. Ferraro, ‘‘Modeling Gaze Shift as a Constrained Random Walk,‘‘ Physica A, vol. 331, 2004.
  176. 176.M.C. potter, ‘‘Meaning in Visual Scenes,’’ Science, vol. 187, pp. 965- 966, 1975.
  177. 177.J.M. Henderson and A. Hollingworth, ‘‘High-Level Scene Percep- tion,’’ Ann. Rev. Psychology, vol. 50, pp. 243-271, 1999.
  178. 178.R.A. Rensink, ‘‘The Dynamic Representation of Scenes,’’ Visual Cognition, vol. 7, pp. 17-42, 2000.
  179. 179.J. Bailenson and N. Yee, ‘‘Digital Chameleons: Automatic Assimilation of Nonverbal Gestures in Immersive Virtual Envir- onments,’’ Psychological Science, vol. 16, pp. 814-819, 2005.
  180. 180.M. Sodhi, B. Reimer, J.L. Cohen, E. Vastenburg, R. Kaars, and S. Kirschenbaum, ‘‘On-Road Driver Eye Movement Tracking Using Head-Mounted Devices,’’ Proc. Symp. Eye Tracking Research and Applications, 2002.
  181. 181.J.H. Reynolds and D.J. Heeger, ‘‘The Normalization Model of Attention,’’ Neuron, vol. 61, no. 2, pp. 168-185, 2009.
  182. 182.S. Engmann, B.M. Hart, T. Sieren, S. Onat, P. König, and W. Einhäuser, ‘‘Saliency on a Natural Scene Background: Effects of Color and Luminance Contrast Add Linearly,’’ Attention, Percep- tion and Psychophysics, vol. 71, no. 6, pp. 1337-1352, 2009.
  183. 183.A. Reeves and G. Sperling, ‘‘Attention Gating in Short-Term Visual Memory,’’ Psychological Rev., vol. 93, no. 2, pp. 180-206, 1986.
  184. 184.L. Itti, ‘‘Quantifying the Contribution of Low-Level Saliency to Human Eye Movements in Dynamic Scenes,’’ Visual Cognition, vol. 12, no. 6, pp. 1093-1123, 2005.
  185. 185.D. Gao, V. Mahadevan, and N. Vasconcelos, ‘‘On the Plausibility of the Discriminant Center-Surround Hypothesis for Visual Saliency,’’ J. Vision, vol. 8, nos. 7-13, pp. 1-18, 2008.
  186. 186.J. Yan, J. Liu, Y. Li, and Y. Liu, ‘‘Visual Saliency via Sparsity Rank Decomposition,’’ Proc. IEEE 17th Int’l Conf. Image Processing, 2010.
  187. 187.http://www.its.caltech.edu/~xhou/ 2012.
  188. 188.J. Yuen, B.C. Russell, C. Liu, and A. Torralba, ‘‘LabelMe Video: Building a Video Database with Human Annotations,’’ Proc. IEEE Int’l Conf. Computer Vision, 2009.
  189. 189.R. Rosenholtz, Y. Li, and L. Nakano, ‘‘Measuring Visual Clutter,’’ J. Vision, vol. 7, no. 17, pp. 1-22, 2007.
  190. 190.R. Rosenholtz, A. Dorai, and R. Freeman, ‘‘Do Predictions of Visual Perception Aid Design?’’ ACM Trans. Applied Perception, vol. 8, no. 2, Article 12, 2011.
  191. 191.R. Rosenholtz, ‘‘A Simple Saliency Model Predicts a Number of Motion Popout Phenomena,’’ Vision Research, vol. 39, pp. 3157- 3163, 1999.
  192. 192.X. Hou, J. Harel, and C. Koch, ‘‘Image Signature: Highlighting Sparse Salient Regions,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 34, no. 1, pp. 194-201, Jan. 2012.
  193. 193.R. Rosenholtz, A.L. Nagy, and N.R. Bell, ‘‘The Effect of Back- ground Color on Asymmetries in Color Search,’’ J. Vision, vol. 4, no. 3, pp. 224-240, 2004.
  194. 194.http://alpern.mit.edu/saliency/ 2012.
  195. 195.D. Green and J. Swets, Signal Detection Theory and Psychophysics. John Wiley, 1966.
  196. 196.T. Jost, N. Ouerhani, R. von Wartburg, R. Mäuri, and H. Häugli, ‘‘Assessing the Contribution of Color in Visual Attention,’’ Computer Vision and Image Understanding, vol. 100, pp. 107-123, 2005.
  197. 197.U. Rajashekar, A.C. Bovik, and L.K. Cormack, ‘‘Visual Search in Noise: Revealing the Influence of Structural Cues by Gaze- Contingent Classification Image Analysis,’’ J. Vision, vol. 13, pp. 379-386, 2006.
  198. 198.S.A. Brandt and L.W. Stark, ‘‘Spontaneous Eye Movements during Visual Imagery Reflect the Content of the Visual Scene,’’ J. Cognitive Neuroscience, vol. 9, nos. 27-38, pp. 27-38, 1997.
  199. 199.A.D. Hwang, H.C. Wang, and M. Pomplun, ‘‘Semantic Guidance of Eye Movements in Real-World Scenes,’’ Vision Research, vol. 51, pp. 1192-1205, 2011.
  200. 200.N. Murray, M. Vanrell, X. Otazu, and C. Alejandro Parraga, ‘‘Saliency Estimation Using a Non-Parametric Low-Level Vision Model,’’ Proc. IEEE Computer Vision and Pattern Recognition, 2011.
  201. 201.W. Wang, C. Chen, Y. Wang, T. Jiang, F. Fang, and Y. Yao, ‘‘Simulating Human Saccadic Scanpaths on Natural Images,’’ Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2011.
  202. 202.R.L Canosa, ‘‘Real-World Vision: Selective Perception and Task,’’ ACM Trans. Applied Perception, vol. 6, no. 2, Article 11, 2009.
  203. 203.M.S. Peterson, A.F. Kramer, and D.E. Irwin, ‘‘Covert Shifts of Attention Precede Involuntary Eye Movements,’’ Perception and Psychophysics, vol. 66, pp. 398-405, 2004.
  204. 204.F. Baluch and L. Itti, ‘‘Mechanisms of Top-Down Attention,’’ Trends in Neuroscience, vol. 34, no. 4, pp. 210-24, 2011.
  205. 205.J. Hayes and A. Efros, ‘‘Scene Completion Using Millions of Photographs,’’ Proc. ACM Siggraph, 2007.
  206. 206.P. Felzenszwalb, R. Girshick, D. McAllester, and D. Ramanan, ‘‘Object Detection with Discriminatively Trained Part Based Models,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 32, no. 9, pp. 1627-1645, Apr. 2010.
  207. 207.A.K. Mishra and Y. Aloimonos, ‘‘Active Segmentation,’’ Int’l J. Humanoid Robotics, vol 6, pp 361-386, 2009.
  208. 208.B. Suh, H. Lingm, B.B. Bederson, and D.W. Jacobs, ‘‘Automatic Thumbnail Cropping and Its Effectiveness,’’ Proc. 16th Ann. ACM Symp. User Interface Software and Technology, pp. 95-104, 2003.
  209. 209.S. Mitri, S. Frintrop, K. Pervolz, H. Surmann, and A. Nuchter, ‘‘Robust Object Detection at Regions of Interest with an Applica- tion in Ball Recognition,’’ Proc. IEEE Int’l Conf. Robotics and Animation, pp. 126-131, Apr. 2005.
  210. 210.N. Ouerhani, R. von Wartburg, H. Hugli, and R.M. Muri, ‘‘Empirical Validation of Saliency-Based Model of Visual Atten- tion,’’ Electronic Letters Computer Vision and Image Analysis, vol. 3, no. 1, pp. 13-24, 2003.
  211. 211.L.W. Stark and Y. Choi, ‘‘Experimental Metaphysics: The Scanpath as an Epistemological Mechanism,’’ Visual Attention and Cognition, pp. 3-69, 1996.
  212. 212.P. Reinagel and A. Zador, ‘‘Natural Scenes at the Center of Gaze,’’ Network, vol. 10, pp. 341-50, 1999.
  213. 213.U. Engelke, H.J. Zepernick, and A. Maeder, ‘‘Visual Attention Modeling: Region-of-Interest Versus Fixation Patterns,’’ Proc. Picture Coding Symp., 2009.
  214. 214.M. Verma and P.W. McOwana, ‘‘Generating Customised Experi- mental Stimuli for Visual Search Using Genetic Algorithms Shows Evidence for a Continuum of Search Efficiency,’’ Vision Research, vol. 49, no. 3, pp. 374-382, 2009.
  215. 215.S. Han and N. Vasconcelos, ‘‘Biologically Plausible Saliency Mechanisms Improve Feedforward Object Recognition,’’ Vision Research, vol. 50, no. 22, pp. 2295-2307, 2010.
  216. 216.D. Ballard, M. Hayhoe, and J. Pelz, ‘‘Memory Representations in Natural Tasks,’’ J. Cognitive Neuroscience, vol. 7, no. 1, pp. 66-80, 1995.
  217. 217.R. Rao, ‘‘Bayesian Inference and Attentional Modulation in the Visual Cortex,’’ NeuroReport, vol. 16, no. 16, pp. 1843-1848, 2005.
  218. 218.A.Borji D.N. Sihite and L. Itti, ‘‘Computational Modeling of Top- Down Visual Attention in Interactive Environments,’’ Proc. British Machine Vision Conf., 2011.
  219. 219.E. Niebur and C. Koch, ‘‘Control of Selective Visual Attention: Modeling the Where Pathway,’’ Proc. Advances in Neural Informa- tion Processing Systems, pp. 802-808, 1995.
  220. 220.P. Viola and M.J. Jones, ‘‘Robust Real-Time Face Detection,’’ Int’l J. Computer Vision, vol. 57, no. 2, pp. 137-154, 2004.
  221. 221.W. Kienzle, B. Sch%lkopf, F.A. Wichmann, and M.O. Franz, ‘‘How to Find Interesting Locations in Video: A Spatiotemporal Interest Point Detector Learned from Human Eye Movements,’’ Proc. 29th DAGM Conf. Pattern Recognition, pp. 405-414, 2007.
  222. 222.J. Wang, J. Sun, L. Quan, X. Tang, and H.Y Shum, ‘‘Picture Collage,’’ Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2006.
  223. 223.D. Gao and N. Vasconcelos, ‘‘Decision-Theoretic Saliency: Computational Principles, Biological Plausibility, and Implica- tions for Neurophysiology and Psychophysics,’’ Neural Computa- tion, vol. 21, pp. 239-271, 2009.
  224. 224.M. Carrasco, ‘‘Visual Attention: The Past 25 Years,’’ Vision Research, vol. 51, pp. 1484-1525, 2011.

Citation

MLA
Borji, A., and L. Itti. “State-of-the-Art in Visual Attention Modeling”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 35, no. 1, 2013, pp. 185–207, https://doi.org/10.1109/TPAMI.2012.89.
APA
Borji, A., & Itti, L. (2013). State-of-the-Art in Visual Attention Modeling. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(1), 185–207. https://doi.org/10.1109/TPAMI.2012.89
Chicago
Borji, A., and L. Itti. 2013. “State-of-the-Art in Visual Attention Modeling”. IEEE Transactions on Pattern Analysis and Machine Intelligence 35 (1): 185–207. https://doi.org/10.1109/TPAMI.2012.89.
Harvard
Borji, A. and Itti, L. (2013) “State-of-the-Art in Visual Attention Modeling”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(1), pp. 185–207. Available at: https://doi.org/10.1109/TPAMI.2012.89.
Vancouver
1. Borji A, Itti L (2013) State-of-the-Art in Visual Attention Modeling. IEEE Transactions on Pattern Analysis and Machine Intelligence 35:185–207

BibTeX

@article{Borji_2013, title={State-of-the-Art in Visual Attention Modeling}, volume={35}, ISSN={2160-9292}, url={http://dx.doi.org/10.1109/TPAMI.2012.89}, DOI={10.1109/tpami.2012.89}, number={1}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Borji, Ali and Itti, Laurent}, year={2013}, month=Jan, pages={185–207} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF