The CMU Pose, Illumination, and Expression Database

Terence SimSimon BakerMaan Bsat

article2003TPAMI1,923 citations

Presents the CMU Pose, Illumination, and Expression database, providing an extensive benchmark of over 40,000 systematically varied facial images across 68 subjects to advance face recognition and 3D modeling research.

Listen

Real-world facial recognition and computer vision systems often struggle when faced with variations in viewing angles, lighting conditions, and facial expressions. Developing reliable algorithms requires standardized benchmarks that capture these real-world factors systematically. The article describes the creation, organization, and distribution of the Carnegie Mellon University Pose, Illumination, and Expression database to provide researchers with a controlled, high-volume image dataset.

The authors set out to construct and document an extensive image database capturing dozens of individuals across a wide spectrum of poses, lighting configurations, and facial expressions. The objective was to supply a rigorous benchmark to evaluate and improve face detection, head pose estimation, and recognition algorithms.

To achieve this, the team collected over 40,000 uncompressed color images from 68 individuals over a three-month period in late 2000. They used a specialized synchronized facility equipped with 13 progressive-scan cameras and an electronically controlled system of 21 flash units. Each subject completed an approximately 10-minute session featuring systematically varied viewing angles, lighting conditions both with and without ambient room illumination, and four distinct expressions (neutral, smiling, blinking, and talking).

The key findings center on the characteristics and utility of the collected dataset. First, the setup successfully produced 43 distinct illumination conditions per subject, capturing natural ambient lighting alongside direct flash variations to replicate realistic environments. Second, the 13 fixed camera positions enabled synchronized multi-angle coverage ranging from full left profile to full right profile, supplemented by typical surveillance and vertical angles. Third, the inclusion of precise geometric measurements for cameras and flashes, background reference frames, and subject demographic attributes provides a fully calibrated environment for three-dimensional modeling and algorithmic stress testing.

These results provide a critical foundation for advancing face-processing technologies by exposing algorithm vulnerabilities to common real-world conditions, such as altered lighting or eye closure during alignment. Access to such varied data reduces development risk and enhances algorithmic performance across security, surveillance, and human-computer interaction applications.

The article recommends that researchers utilize the dataset to benchmark pose-invariant detectors, cross-pose recognition systems, and 3D facial reconstruction models. Interested organizations can obtain the 40-gigabyte database by mailing physical storage drives to the researchers. While the dataset offers controlled precision, users should note that the sample size is limited to 68 subjects captured within an indoor laboratory setting, and the uncompressed storage format requires dedicated data handling.

Sim et al (2003).pdf
  • Paper: Eigenfaces vs. Fisherfaces: Recognition Using Class Specific Linear Projection, Peter N. Belhumeur et al. (1996). Belhumeur et al. establish foundational subspace and linear projection methods for handling lighting and expression variations, establishing the core challenges that motivated the creation of systematic datasets like CMU PIE.
  • Paper: A Morphable Model For The Synthesis Of 3D Faces, Volker Blanz et al. (1999). Blanz and Vetter introduce the 3D morphable model framework for synthesizing and modeling facial variations in pose and illumination, providing key motivation for controlled multi-view facial capture.
  • Paper: Active Appearance Models, Timothy F. Cootes et al. (1998). Cootes et al. develop Active Appearance Models to jointly fit shape and texture variations, highlighting the need for comprehensive benchmarks capturing pose and expression.
  • Paper: Detecting Faces in Images: A Survey, Ming-Hsuan Yang et al. (2002). Yang et al. provide an extensive survey of face detection challenges across varying lighting, pose, and expression conditions that contextualizes the requirement for the CMU PIE database.
  • Paper: Automatic Analysis of Facial Expressions: The State of the Art, Maja Pantic et al. (2000). Pantic and Rothkrantz survey facial expression analysis and outline the operational limitations imposed by existing constrained laboratory datasets.
Cover for The CMU Pose, Illumination, and Expression Database

Abstract

In the Fall of 2000 we collected a database of over 40,000 facial images of 68 people. Using the CMU 3D Room we imaged each person across 13 different poses, under 43 different illumination conditions, and with 4 different expressions. We call this the CMU Pose, Illumination, and Expression (PIE) database. We describe the imaging hardware, the collection procedure, the organization of the images, several possible uses, and how to obtain the database.

Table of Contents

  • 1 Introduction
  • 2 Capture Apparatus
  • 2.1 Setup of the Cameras: Pose
  • 2.2 The Flash System: Illumination
  • 3 Database Contents
  • 3.1 Pose Variation
  • 3.2 Pose and Illumination Variation
  • 3.3 Pose and Expression Variation
  • 3.4 Meta-Data
  • 4 Potential Uses of the Database
  • 5 Obtaining the Database
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — CMU Pose, Illumination, and Expression (PIE) Database Specifications

    experimental setup

    The CMU Pose, Illumination, and Expression (PIE) database is a multi-modal facial dataset collected from 68 human subjects. It contains over 40,000 uncompressed color images totaling approximately 40 GB (approximately 600 MB per subject). The capture protocol yields over 600 retained images per subject across:

    • 13 distinct poses recorded simultaneously using calibrated cameras.
    • 43 illumination conditions spanning compound ambient and directed point-source flash configurations.
    • 4 facial expressions/states: neutral, smiling, blinking (eyes shut), and talking.
    • An additional neutral capture sequence without glasses for subjects who normally wear eyeglasses.
  2. Knowl 2 — Multi-Camera Geometric Configuration for Pose Acquisition

    experimental setup

    To acquire 13 simultaneous poses per subject without moving hardware or relying on sequential subject re-positioning, 13 synchronized Sony DXC 9000 cameras (3 CCD3\text{ CCD}, progressive scan) operate simultaneously with gain and gamma correction turned off:

    • Horizontal sweep (9 cameras): Positioned at approximate head height along an arc from full left profile to full right profile, with adjacent cameras separated by approximately 22.5∘22.5^\circ.
    • Elevation perspectives (2 cameras): Positioned directly above and directly below the central frontal camera.
    • Surveillance perspectives (2 cameras): Positioned in upper room corners to duplicate typical surveillance camera geometry.

    Precise 3D Cartesian coordinates (x,y,z)(x, y, z) of the head location and all 13 camera centers were surveyed with a Leica theodolite and are provided in the database metadata.

  3. Knowl 3 — Programmable Electronic Flash System for Illumination Variation

    experimental setup

    Illumination variation is generated by an array of 21 Minolta 220X electronic flashes triggered via their hot-shoes using an Advantech PCL-734 32-channel digital output board:

    • Each flash discharge lasts approximately 1 ms1\text{ ms} and occurs entirely within the camera's open shutter window (duration ≈16 ms\approx 16\text{ ms}).
    • By firing one flash sequentially per frame, 21 images under distinct illumination directions are captured in 2130≈0.7 s\frac{21}{30}\approx 0.7\text{ s}.
    • Image sets are acquired under two environmental lighting conditions: with ambient room lights on (yielding natural compound illumination) and with room lights off (yielding isolated point-source illumination).

    Across 21 flashes with room lights on, 21 flashes with room lights off, and 1 baseline capture under ambient room lights alone, each pose is recorded under 21×2+1=4321 \times 2 + 1 = 43 distinct illumination conditions. Exact 3D positions (x,y,z)(x, y, z) of all 21 flashes were surveyed using a Leica theodolite.

  4. Knowl 4 — Facial Expression and Dynamic Talking Capture Protocol

    experimental setup

    Facial appearance variation is recorded across four distinct expressive states:

    1. Neutral Expression: Static neutral face recorded across all 13 camera poses.
    2. Smiling Expression: Static smile recorded across all 13 camera poses.
    3. Blinking State: Static capture with eyes intentionally closed across all 13 camera poses, designed to evaluate whether pupil-dependent face alignment algorithms degrade when subjects blink.
    4. Talking Sequence: A 2-second continuous video sequence (60 frames at 30 fps) captured synchronously from 3 selected viewpoints: the frontal camera, the three-quarter profile camera, and the full profile camera.

    For subjects who wear eyeglasses, an additional set of 13 neutral-expression images is acquired without glasses to isolate underlying facial geometry from occlusion.

  5. Knowl 5 — Auxiliary Calibration and Attribute Metadata in CMU PIE

    data/table

    The CMU PIE database provides calibration data and demographic metadata alongside facial images:

    • Spatial Coordinates: 3D (x,y,z)(x, y, z) coordinates of the subject's head position, all 13 cameras, and all 21 flashes measured via theodolite survey.
    • Background Reference Frames: Baseline background images captured from each of the 13 cameras without a subject present, enabling background subtraction and face localization.
    • Color Calibration Images: Images of standard color calibration charts captured under controlled lighting to calibrate camera gain, bias, and color response across manually adjusted apertures.
    • Demographic Attributes: Subject metadata recording sex, age, presence/absence of eyeglasses, presence/absence of facial hair (mustache and beard), and recording session date.
  6. Knowl 6 — Image Format and Vertical Interval TimeCode Encoding

    experimental setup

    Facial images in the CMU PIE database are stored in uncompressed 24-bit color raw PPM (Portable Pixmap) format at a resolution of 640×486640 \times 486 pixels. The first 6 image rows (6×6406 \times 640 pixels) contain binary-encoded Vertical Interval TimeCode (VITC) synchronization metadata generated by hardware capture units in the recording room. These initial 6 rows can be cropped or discarded to recover a standard 640×480640 \times 480 pixel active face region.

  7. Knowl 7 — Computer Vision Benchmark Applications of CMU PIE

    model/method

    The CMU PIE database serves as a standardized evaluation benchmark for multiple computer vision tasks:

    • Pose-Invariant Face Recognition: Evaluating probe-to-gallery matching where probe and gallery images have differing viewpoints, including multi-view probe fusion.
    • Illumination-Invariant Face Recognition: Evaluating classification robustness across directed point sources, ambient illumination, and compound lighting.
    • Head Pose Estimation: Evaluating continuous and discrete head orientation algorithms against ground-truth surveyed camera vectors.
    • 3D Face Reconstruction: Generating 3D facial shape via multi-view stereo across synchronized camera angles or via photometric stereo across fixed-pose flash sequences.
    • Pose-Invariant Face Detection: Training and testing component-based or holistic face detectors across full profiles, frontal views, and surveillance perspectives.

Coverage note — Logistical distribution details (shipping physical IDE hard drives) and acknowledgments were deliberately omitted as they do not constitute scientific contributions.

References

  1. 1.V. Blanz, S. Romdhani, and T. Vetter. Face identification across different poses and illuminations with a 3D morphable model. In Proceedings of the 5th IEEE International Conference on Face and Gesture Recognition, 2002.
  2. 2.A.S. Georghiades, P.N. Belhumeur, and D.J. Kriegman. From few to many: Generative models for recognition under variable pose and illumination. In Proceedings of the 4th IEEE International Conference on Automatic Face and Gesture Recognition, 2000.
  3. 3.R. Gross, J. Shi, and J. Cohn. Quo vadis face recognition. In Proceedings of the 3rd Workshop on Empirical Evaluation Methods in Computer Vision, 2001.
  4. 4.R. Gross, I. Matthews, and S. Baker. Eigen light-fields and face recognition across pose. In Proceedings of the 5th IEEE International Conference on Face and Gesture Recognition, 2002.
  5. 5.B. Heisele, T. Serre, M. Pontil, and T. Poggio. Component-based face detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2001.
  6. 6.B. Heisele, T. Serre, M. Pontil, T. Vetter, and T. Poggio. Categorization by learning and combining object parts. In Neural Information Processing Systems, 2001.
  7. 7.T. Kanade, H. Saito, and S. Vedula. The 3D room: Digitizing time-varying 3D events by synchronized multiple video streams. Technical Report CMU-RI-TR-98-34, Carnegie Mellon University Robotics Institute, 1998.
  8. 8.P.J. Philips, H. Moon, P. Rauss, and S.A. Rizvi. The FERET evaluation methodology for face-recognition algorithms. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1997.
  9. 9.T. Sim and T. Kanade. Combining models and exemplars for face recognition: An illuminating example. In Proceedings of the Workshop on Models versus Exemplars in Computer Vision, 2001.
  10. 10.T. Sim, S. Baker, and M. Bsat. The CMU pose, illumination, and expression (PIE) database. In Proceedings of the 5th IEEE International Conference on Face and Gesture Recognition, 2002.

Citation

MLA
Terence Sim, et al. “The CMU Pose, Illumination, and Expression Database”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 25, no. 12, 2003, pp. 1615–18, https://doi.org/10.1109/TPAMI.2003.1251154.
APA
Terence Sim, Baker, S., & Bsat, M. (2003). The CMU pose, illumination, and expression database. IEEE Transactions on Pattern Analysis and Machine Intelligence, 25(12), 1615–1618. https://doi.org/10.1109/TPAMI.2003.1251154
Chicago
Terence Sim, S. Baker, and M. Bsat. 2003. “The CMU Pose, Illumination, and Expression Database”. IEEE Transactions on Pattern Analysis and Machine Intelligence 25 (12): 1615–18. https://doi.org/10.1109/TPAMI.2003.1251154.
Harvard
Terence Sim, Baker, S. and Bsat, M. (2003) “The CMU pose, illumination, and expression database”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 25(12), pp. 1615–1618. Available at: https://doi.org/10.1109/TPAMI.2003.1251154.
Vancouver
1. Terence Sim, Baker S, Bsat M (2003) The CMU pose, illumination, and expression database. IEEE Transactions on Pattern Analysis and Machine Intelligence 25:1615–1618

BibTeX

@article{Terence_Sim_2003, title={The CMU pose, illumination, and expression database}, volume={25}, ISSN={0162-8828}, url={http://dx.doi.org/10.1109/TPAMI.2003.1251154}, DOI={10.1109/tpami.2003.1251154}, number={12}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Terence Sim and Baker, S. and Bsat, M.}, year={2003}, month=Dec, pages={1615–1618} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF