Learning to predict where humans look

Tilke JuddKrista EhingerFredo DurandAntonio Torralba

article2009ICCV2,198 citations

Presents a large benchmark dataset of human eye fixations across 1003 images and trains a supervised model combining low-, mid-, and high-level visual features to significantly outperform traditional bottom-up saliency algorithms.

Abstract

For many applications in graphics, design, and human computer interaction, it is essential to understand where humans look in a scene. Where eye tracking devices are not a viable option, models of saliency can be used to predict fixation locations. Most saliency approaches are based on bottom-up computation that does not consider top-down image semantics and often does not match actual eye movements. To address this problem, we collected eye tracking data of 15 viewers on 1003 images and use this database as training and testing examples to learn a model of saliency based on low, middle and high-level image features. This large database of eye tracking data is publicly available with this paper.

Table of Contents

  • 1. Introduction
  • 2. Database of eye tracking data
  • 2.1. Data gathering protocol
  • 2.2. Analysis of dataset
  • 3. Learning a model of saliency
  • 3.1. Features used for machine learning
  • 3.2. Training
  • 3.3. Performance
  • 3.4. Applications
  • 4. Conclusion
  • References

Knowls

  1. Knowl 1 — MIT Eye Tracking Dataset and Collection Protocol

    experimental setup

    The benchmark eye-tracking dataset comprises 1,003 natural images collected from Flickr Creative Commons and LabelMe, consisting of 779 landscape-oriented images and 228 portrait-oriented images. The longest dimension of each image is 1,024 pixels, while the shorter dimension ranges from 405 to 1,024 pixels (with the majority at 768 pixels).

    Eye-tracking data was gathered from 15 human observers (males and females aged 18 to 35; 2 researchers and 13 naive participants) under free-viewing conditions. Observers viewed images on a 19-inch screen (resolution 1280×10241280 \times 1024) from a distance of approximately 2 feet in a darkened room with head stabilization provided by a chin rest. Images were presented at full resolution for 3.0 seconds each, separated by 1.0 second of a blank gray screen. Observers completed two sessions of 500 randomly ordered images approximately one week apart. Camera calibration was re-verified every 50 images. To encourage attention, observers took a 100-image recognition memory test at the conclusion of each session.

    The initial fixation of each scanpath was removed to eliminate trivial center-start bias. A continuous ground-truth saliency map for each image was generated by convolving a 2D Gaussian filter over the spatial fixation coordinates across all 15 viewers.

  2. Knowl 2 — Statistical Properties and Center Bias of Human Eye Fixations

    empirical result

    Analysis of human fixation distributions across 1,003 natural images reveals distinct spatial and semantic properties:

    • Center Bias: Fixations are heavily concentrated near the center of the image frame: 40%40\% of all fixations fall within the central 11%11\% of image area, and 70%70\% fall within the central 25%25\% of image area.
    • Inter-Observer Consistency: Using the saliency map of one observer as a binary classifier to predict fixations from the remaining 14 observers yields high consistency under an ROC metric: 60%60\% of ground-truth fixations fall within the top 5%5\% most salient regions of an individual's map, and 90%90\% fall within the top 20%20\% most salient regions.
    • Semantic Preferences: Ground-truth bounding box labeling indicates that 10%10\% of all fixations fall directly on human faces, and 11%11\% fall on text regions. Fixations are also consistently attracted to people, body parts (e.g., eyes and hands), animals, and cars.
    • Entropy Distribution: Image fixation consistency varies with scene complexity: low-entropy fixation maps correspond to images containing a single dominant central object, whereas high-entropy fixation maps correspond to multi-textured scenes without a single focal target.
  3. Knowl 3 — Multi-Level Feature Representation for Visual Saliency

    model/method

    To train a supervised saliency model, each image is resized to 200×200200 \times 200 pixels and represented by a 33-dimensional feature vector computed at every pixel location, spanning low-, mid-, and high-level image properties:

    1. Low-Level Features (29 channels):

      • Steerable Pyramid Energy: Local energy of steerable pyramid filters computed across 4 orientations and 3 spatial scales (13 features).
      • Subband Saliency: Subband pyramid contrast features from Torralba and Rosenholtz (1 feature).
      • Itti-Koch Channels: Intensity contrast, orientation contrast, and color contrast computed via the standard Itti-Koch saliency model (3 features).
      • Color Statistics: Pixel-wise raw RGB channel values (3 features), marginal RGB color probabilities (3 features), and joint color probabilities evaluated from 3D color histograms after median filtering across 6 distinct spatial scales (6 features).
    2. Mid-Level Features (1 channel):

      • Horizon Detector: A probabilistic estimate of the horizon line location trained from spatial envelope (Gist) features.
    3. High-Level Features (2 channels):

      • Face Detection: Continuous detection score output from a Viola-Jones face detector.
      • Person Detection: Output from a Felzenszwalb deformable part model person detector.
    4. Spatial Prior (1 channel):

      • Distance to Center: The Euclidean distance from the pixel coordinate to the image center.
  4. Knowl 4 — Supervised Saliency Learning and Sampling Protocol

    model/method

    The saliency prediction model is trained using a linear Support Vector Machine (SVM) on labeled pixel samples from the image dataset:

    • Dataset Partitioning: 903 images are used for training and 100 images are reserved for testing.
    • Sample Selection: From each image, 10 positive pixels are sampled uniformly at random from the top 20%20\% most salient locations of the ground-truth smoothed fixation map, and 10 negative pixels are sampled from the bottom 70%70\% salient locations. Samples in the intermediate 20%–70%20\%\text{--}70\% zone and within 10 pixels of the image border are excluded to avoid label ambiguity and border artifacts. This yields 18,06018{,}060 training samples (9,0309{,}030 positive, 9,0309{,}030 negative) and 2,0002{,}000 testing samples.
    • Normalization: Training features are standardized to zero mean and unit variance (zz-score normalization), and the learned normalization parameters are applied to test features.
    • Model Training: A linear SVM is trained using LIBLINEAR with a misclassification cost parameter of c=1.0c = 1.0.
    • Continuous Saliency Inference: For a test pixel with feature vector x∈R33x \in \mathbb{R}^{33}, the continuous saliency value is computed as: S(x)=wTx+bS(x) = w^T x + b where w∈R33w \in \mathbb{R}^{33} is the learned weight vector and b∈Rb \in \mathbb{R} is the bias offset.
  5. Knowl 5 — Saliency Prediction Performance Across Feature Subsets

    empirical result

    Saliency models are evaluated by thresholding the predicted continuous saliency map S(x)S(x) at the top n%n\% of image pixels (n∈{1,3,5,10,15,20,25,30}n \in \{1, 3, 5, 10, 15, 20, 25, 30\}) and measuring the percentage of true human fixations captured within that salient area:

    • Full Model Performance: The combined model containing all 33 features outperforms all single-feature models and baselines (including Itti-Koch, Torralba-Rosenholtz, and Cerf et al.), achieving 88%88\% of human-level performance. At a threshold of 20%20\% salient area, the full model captures 75%75\% of human fixations compared to 85%85\% for human-to-human prediction.
    • Performance Without Center Prior: The model using all features except the center prior captures 60%60\% of human fixations at the 20%20\% salient threshold. This matches the standalone performance of the center-distance feature alone (~60%60\%), despite the all-features-without-center model receiving no spatial location information.
    • Feature Ablation Impact: Measuring the performance increase gained by adding individual feature sets to the center-distance prior shows that steerable pyramid subbands and Torralba subband features yield the largest performance gain, followed by color features, horizon detection, object detectors (face/person), and Itti-Koch channels.
    • Object Detector Features: While face and person detectors achieve lower standalone ROC scores across full images due to the absence of targets in many scenes, their inclusion provides substantial performance boosts on target-containing images.
  6. Knowl 6 — Spatial and Semantic Subgroup Robustness of Saliency Predictors

    empirical result

    Evaluating models across spatial and semantic subsets reveals performance divergences between purely location-based priors and image-feature-based models:

    • Subgroup Partitioning: Test samples are partitioned into circular central regions (distance to center ≤0.42\le 0.42, representing the center 27.7%27.7\% of the image area) versus peripheral regions, and face regions versus non-face regions. Performance is reported as balanced accuracy (average of true positive rate and true negative rate).
    • Center Prior Fragility: The standalone center-prior model performs near chance (50%50\%) when evaluated strictly on inside-only or outside-only subsets. Its strong performance on the aggregate dataset relies entirely on the global sample imbalance (79%79\% of positive fixations lie in the central region, whereas 75%75\% of negative samples lie in the periphery).
    • Feature Model Robustness: The all-features-without-center model maintains balanced predictive accuracy across both central and peripheral subsets, demonstrating genuine feature-driven saliency detection rather than reliance on spatial bias.
    • Semantic Subsets: Models incorporating high-level face, person, and car detector features outperform low-level and center models on sample subsets containing faces and people.
  7. Knowl 7 — Saliency-Guided Non-Photorealistic Photograph Abstraction

    model/method

    The supervised saliency prediction model can be directly applied to automated non-photorealistic rendering and image stylization (e.g., the DeCarlo and Santella framework) without requiring hardware-based eye-tracking measurements:

    1. The input image is processed through the trained linear SVM model to produce a continuous spatial saliency map S(x)=wTx+bS(x) = w^T x + b.
    2. Pixels assigned high saliency values correspond to predicted human gaze targets and are rendered with high visual detail and edge fidelity.
    3. Pixels assigned low saliency values are simplified and abstracted into broad, uniform color regions with reduced geometric detail.

Coverage note — Preliminary exploratory observations on the histogram of region-of-interest (ROI) bounding box radii on 30 images were omitted as they were identified as informal pilot analyses for future work rather than core methodological results.

References

  1. 1.T. Avraham and M. Lindenbaum. Esaliency: Meaningful attention using stochastic image modeling. IEEE Transactions on Pattern Analysis and Machine Intelligence, 99(1), 2009.
  2. 2.N. D. B. Bruce and J. K. Tsotsos. Saliency, attention, and visual search: An information theoretic approach. Journal of Vision, 9(3):1–24, 3 2009.
  3. 3.M. Cerf, J. Harel, W. Einhauser, and C. Koch. Predicting human gaze using low-level saliency combined with face detection. In J. C. Platt, D. Koller, Y. Singer, and S. T. Roweis, editors, NIPS. MIT Press, 2007.
  4. 4.D. DeCarlo and A. Santella. Stylization and abstraction of photographs. ACM Transactions on Graphics, 21(3):769–776, July 2002.
  5. 5.K. Ehinger, B. Hidalgo-Sotelo, A. Torralba, and A. Oliva. Modeling search for people in 900 scenes: A combined source model of eye guidance. Visual Cognition, 2009.
  6. 6.P. Felzenszwalb, D. McAllester, and D. Ramanan. A discriminatively trained, multiscale, deformable part model. Computer Vision and Pattern Recognition, 2008. CVPR 2008. IEEE Conference on, pages 1–8, June 2008.
  7. 7.W. S. Geisler and J. S. Perry. A real-time foveated multiresolution system for low-bandwidth video communication. In in Proc. SPIE, pages 294–305, 1998.
  8. 8.X. Hou and L. Zhang. Saliency detection: A spectral residual approach. Computer Vision and Pattern Recognition, IEEE Computer Society Conference on, 0:1–8, 2007.
  9. 9.L. Itti and C. Koch. A saliency-based search mechanism for overt and covert shifts of visual attention. Vision Research, 40:1489–1506, 2000.
  10. 10.W. Kienzle, F. A. Wichmann, B. Schölkopf, and M. O. Franz. A nonparametric approach to bottom-up visual saliency. In B. Schölkopf, J. C. Platt, and T. Hoffman, editors, NIPS, pages 689–696. MIT Press, 2006.
  11. 11.O. L. Meur, P. L. Callet, and D. Barba. Predicting visual fixations on video based on low-level visual features. Vision Research, 47(19):2483 – 2498, 2007.
  12. 12.A. Oliva and A. Torralba. Modeling the shape of the scene: A holistic representation of the spatial envelope. International Journal of Computer Vision, 42:145–175, 2001.
  13. 13.R. Rosenholtz. A simple saliency model predicts a number of motion popout phenomena. Vision Research 39, 19:3157–3163, 1999.
  14. 14.M. Rubinstein, A. Shamir, and S. Avidan. Improved seam carving for video retargeting. ACM Transactions on Graphics (SIGGRAPH), 2008.
  15. 15.B. Russell, A. Torralba, K. Murphy, and W. Freeman. Labelme: a database and web-based tool for image annotation. MIT AI Lab Memo AIM-2005-025, MIT CSAIL, Sept. 2005.
  16. 16.A. Santella, M. Agrawala, D. DeCarlo, D. Salesin, and M. Cohen. Gaze-based interaction for semi-automatic photo cropping. In CHI ’06: Proceedings of the SIGCHI conference on Human Factors in computing systems, pages 771–780, New York, NY, USA, 2006. ACM.
  17. 17.E. P. Simoncelli and W. T. Freeman. The steerable pyramid: A flexible architecture for multi-scale derivative computation. pages 444–447, 1995.
  18. 18.S. Sonnenburg, G. Rätsch, C. Schäfer, and B. Schölkopf. Large scale multiple kernel learning. J. Mach. Learn. Res., 7:1531–1565, 2006.
  19. 19.B. W. Tatler. The central fixation bias in scene viewing: Selecting an optimal viewing position independently of motor biases and image feature distributions. J. Vis., 7(14):1–17, 11 2007.
  20. 20.B. M. Velichkovsky, M. Pomplun, J. Rieser, and H. J. Ritter. Attention and Communication: Eye-Movement-Based Research Paradigms. Visual Attention and Cognition. Elsevier Science B.V., Amsterdam, 1996.
  21. 21.P. Viola and M. Jones. Robust real-time object detection. In International Journal of Computer Vision, 2001.
  22. 22.Z. Wang, L. Lu, and A. C. Bovik. Foveation scalable video coding with automatic fixation selection. IEEE Trans. Image Processing, 12:243–254, 2003.
  23. 23.L. Zhang, M. H. Tong, T. K. Marks, H. Shan, and G. W. Cottrell. SUN: A Bayesian framework for saliency using natural statistics. J. Vis., 8(7):1–20, 12 2008.

Citation

MLA
Judd, T., et al. “Learning to Predict Where Humans Look”. 2009 IEEE 12th International Conference on Computer Vision, 2009, pp. 2106–13, https://doi.org/10.1109/ICCV.2009.5459462.
APA
Judd, T., Ehinger, K., Durand, F., & Torralba, A. (2009). Learning to predict where humans look. 2009 IEEE 12th International Conference on Computer Vision, 2106–2113. https://doi.org/10.1109/ICCV.2009.5459462
Chicago
Judd, T., K. Ehinger, F. Durand, and A. Torralba. 2009. “Learning to Predict Where Humans Look”. 2009 IEEE 12th International Conference on Computer Vision, 2106–13. https://doi.org/10.1109/ICCV.2009.5459462.
Harvard
Judd, T. et al. (2009) “Learning to predict where humans look”, 2009 IEEE 12th International Conference on Computer Vision. IEEE, pp. 2106–2113. Available at: https://doi.org/10.1109/ICCV.2009.5459462.
Vancouver
1. Judd T, Ehinger K, Durand F, Torralba A (2009) Learning to predict where humans look. In: 2009 IEEE 12th International Conference on Computer Vision. IEEE, pp 2106–2113

BibTeX

@inproceedings{Judd_2009, title={Learning to predict where humans look}, url={http://dx.doi.org/10.1109/ICCV.2009.5459462}, DOI={10.1109/iccv.2009.5459462}, booktitle={2009 IEEE 12th International Conference on Computer Vision}, publisher={IEEE}, author={Judd, Tilke and Ehinger, Krista and Durand, Fredo and Torralba, Antonio}, year={2009}, month=Sept, pages={2106–2113} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE