Age Progression/Regression by Conditional Adversarial Autoencoder
Zhifei ZhangYang SongHairong Qi
Proposes a conditional adversarial autoencoder that performs realistic facial age progression and regression from a single unlabeled photograph without requiring paired training data while preserving individual identity.
Predicting how a human face changes over time—both simulating future appearance through age progression and reconstructing younger appearance through age regression—is a critical capability for missing person searches, law enforcement, identity verification, and digital media. However, conventional automated systems face severe operational bottlenecks. Most legacy approaches depend on paired training data that captures the exact same individual across multiple decades, require the input image to be explicitly labeled with the subject's true age, and struggle to generate realistic baby or elderly appearances without introducing artificial ghosting or blurring.
The main objective of the article is to demonstrate a generative deep learning framework called the Conditional Adversarial Autoencoder (CAAE), which performs simultaneous age progression and regression on unlabeled face images without requiring paired longitudinal datasets.
To evaluate this framework, the authors assembled a balanced dataset of 10,670 face images across ten age brackets ranging from infancy to old age (0 to 80 years old), sourcing portraits from established image databases and web search queries. The technical approach maps input images into a simplified feature space that separates permanent facial identity from age attributes. The system employs two neural network components—an encoder and an image generator—supported by two competing discriminator networks. One discriminator enforces a smooth, continuous distribution of personal features, while the other ensures output images appear realistic and accurately match the target age.
The evaluation yielded several key findings. First, blinded human perception surveys comparing CAAE outputs to ground-truth images separated by more than 20 years found that 48.38% of respondents judged the synthesized faces to be the same individual, compared to 29.58% who judged them different and 22.04% who were uncertain. Second, in head-to-head comparisons against previous state-of-the-art aging techniques, respondents favored the CAAE method in 52.77% of votes, compared to 28.99% preferring prior methods and 18.24% rating them equal. Third, qualitative tests showed that the architecture successfully simulates plausible baby faces from adult inputs and synthesizes realistic, detailed skin wrinkles for older targets. Finally, the system demonstrated high visual stability and feature preservation even when input portraits exhibited strong facial expressions, non-frontal poses, or partial facial occlusions.
These findings indicate that high-fidelity facial aging and rejuvenation do not require cost-prohibitive, multi-decade photo tracking of specific individuals. By decoupling personal identity from age attributes, organizations can process arbitrary input photos without knowing the subject's chronological age at the time of capture. This substantially reduces data acquisition costs, minimizes preparation time, and broadens the operational applicability of automated aging tools in forensic, security, and verification workflows.
Based on these results, project leaders should consider piloting the CAAE architecture for visual identity tracking and age-invariant verification workflows. Development teams should also explore adapting individual network components for related operational tasks, such as using the feature encoder for cross-age identity recognition and the image discriminator for automated age estimation. Prior to full-scale deployment in critical forensic environments, practitioners should validate the model on higher-resolution imagery and evaluate larger demographic datasets to account for potential data-crawling labeling noise in early life brackets.
- Paper: Adversarial Autoencoders, Alireza Makhzani et al. (2015). This paper establishes the adversarial autoencoder framework that uses adversarial training to match latent representations to prior distributions, providing the direct architectural foundation for the conditional adversarial autoencoder.
- Paper: Conditional Generative Adversarial Nets, Mehdi Mirza et al. (2014). This foundational work introduces conditioning mechanisms into generative adversarial networks, which the source extends to condition autoencoders on age labels.
- Paper: Autoencoding beyond pixels using a learned similarity metric, Anders Boesen Lindbo Larsen et al. (2015). This paper introduces the hybrid combination of autoencoders and generative adversarial networks to overcome pixel-space blurriness, directly motivating CAAE's dual adversarial setup on latent and image spaces.
- Paper: Generative Visual Manipulation on the Natural Image Manifold, Jun-Yan Zhu et al. (2016). This work demonstrates how to traverse and project onto a learned natural image manifold for attribute editing, directly inspiring CAAE's manifold-traversal approach to age progression.
- Paper: Coupled Generative Adversarial Networks, Ming-Yu Liu et al. (2016). This paper demonstrates learning joint distributions across unaligned visual domains without paired training data, setting the precedent for CAAE's unpaired age-group mapping.
- Paper: Deep Learning Face Attributes in the Wild, Ziwei Liu et al. (2015). This paper provides benchmark methodologies and datasets (CelebA) for learning unconstrained facial attributes like age, which form the empirical basis for training face aging models.
- Paper: Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks, Alec Radford et al. (2016). This paper introduces stable deep convolutional architectures for generative adversarial networks, which form the structural building blocks for CAAE's convolutional encoder and deconvolutional generator.
- Paper: Generative Adversarial Networks, Ian J. Goodfellow et al. (2014). This seminal text introduces the minimax formulation of Generative Adversarial Networks, which provides the underlying objective for both discriminator networks in CAAE.
- Paper: StarGAN v2: Diverse Image Synthesis for Multiple Domains, Yunjey Choi et al. (2019). This work advances multi-domain image-to-image translation for facial attributes by introducing continuous style codes and multi-modal diversity across distinct domains.
- Paper: A Style-Based Generator Architecture for Generative Adversarial Networks, Tero Karras et al. (2019). This work develops a style-based generator architecture that enables fine-grained, disentangled control over facial features and age attributes across different scales.
- Paper: StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery, Or Patashnik et al. (2021). This paper combines pretrained style spaces with text encoders to provide intuitive natural language control over facial attributes, generalizing beyond predefined categorical age conditioning.
- Paper: Analyzing and Improving the Image Quality of StyleGAN, Tero Karras et al. (2020). This paper refines generative face modeling by redesigning normalization and introducing path length regularization to eliminate visual artifacts and ensure smooth latent space traversals.
- Paper: cGANs with Projection Discriminator, Takeru Miyato et al. (2018). This work improves conditional adversarial training by introducing projection-based discriminators, enhancing stability and fidelity in attribute-conditioned synthesis.
- Paper: Contrastive Learning for Unpaired Image-to-Image Translation, Taesung Park et al. (2020). This paper proposes contrastive unpaired translation to maintain structural identity while transforming domain attributes, addressing content preservation challenges inherent in unpaired image synthesis.
- Paper: Efficient Geometry-aware 3D Generative Adversarial Networks, Eric R. Chan et al. (2022). This paper extends generative face modeling into 3D-aware synthesis, ensuring multi-view consistency while performing facial generation and editing.
