Age Progression/Regression by Conditional Adversarial Autoencoder

Zhifei ZhangYang SongHairong Qi

article2017CVPR1,350 citations

Proposes a conditional adversarial autoencoder that performs realistic facial age progression and regression from a single unlabeled photograph without requiring paired training data while preserving individual identity.

Listen

Predicting how a human face changes over time—both simulating future appearance through age progression and reconstructing younger appearance through age regression—is a critical capability for missing person searches, law enforcement, identity verification, and digital media. However, conventional automated systems face severe operational bottlenecks. Most legacy approaches depend on paired training data that captures the exact same individual across multiple decades, require the input image to be explicitly labeled with the subject's true age, and struggle to generate realistic baby or elderly appearances without introducing artificial ghosting or blurring.

The main objective of the article is to demonstrate a generative deep learning framework called the Conditional Adversarial Autoencoder (CAAE), which performs simultaneous age progression and regression on unlabeled face images without requiring paired longitudinal datasets.

To evaluate this framework, the authors assembled a balanced dataset of 10,670 face images across ten age brackets ranging from infancy to old age (0 to 80 years old), sourcing portraits from established image databases and web search queries. The technical approach maps input images into a simplified feature space that separates permanent facial identity from age attributes. The system employs two neural network components—an encoder and an image generator—supported by two competing discriminator networks. One discriminator enforces a smooth, continuous distribution of personal features, while the other ensures output images appear realistic and accurately match the target age.

The evaluation yielded several key findings. First, blinded human perception surveys comparing CAAE outputs to ground-truth images separated by more than 20 years found that 48.38% of respondents judged the synthesized faces to be the same individual, compared to 29.58% who judged them different and 22.04% who were uncertain. Second, in head-to-head comparisons against previous state-of-the-art aging techniques, respondents favored the CAAE method in 52.77% of votes, compared to 28.99% preferring prior methods and 18.24% rating them equal. Third, qualitative tests showed that the architecture successfully simulates plausible baby faces from adult inputs and synthesizes realistic, detailed skin wrinkles for older targets. Finally, the system demonstrated high visual stability and feature preservation even when input portraits exhibited strong facial expressions, non-frontal poses, or partial facial occlusions.

These findings indicate that high-fidelity facial aging and rejuvenation do not require cost-prohibitive, multi-decade photo tracking of specific individuals. By decoupling personal identity from age attributes, organizations can process arbitrary input photos without knowing the subject's chronological age at the time of capture. This substantially reduces data acquisition costs, minimizes preparation time, and broadens the operational applicability of automated aging tools in forensic, security, and verification workflows.

Based on these results, project leaders should consider piloting the CAAE architecture for visual identity tracking and age-invariant verification workflows. Development teams should also explore adapting individual network components for related operational tasks, such as using the feature encoder for cross-age identity recognition and the image discriminator for automated age estimation. Prior to full-scale deployment in critical forensic environments, practitioners should validate the model on higher-resolution imagery and evaluate larger demographic datasets to account for potential data-crawling labeling noise in early life brackets.

  • Paper: Adversarial Autoencoders, Alireza Makhzani et al. (2015). This paper establishes the adversarial autoencoder framework that uses adversarial training to match latent representations to prior distributions, providing the direct architectural foundation for the conditional adversarial autoencoder.
  • Paper: Conditional Generative Adversarial Nets, Mehdi Mirza et al. (2014). This foundational work introduces conditioning mechanisms into generative adversarial networks, which the source extends to condition autoencoders on age labels.
  • Paper: Autoencoding beyond pixels using a learned similarity metric, Anders Boesen Lindbo Larsen et al. (2015). This paper introduces the hybrid combination of autoencoders and generative adversarial networks to overcome pixel-space blurriness, directly motivating CAAE's dual adversarial setup on latent and image spaces.
  • Paper: Generative Visual Manipulation on the Natural Image Manifold, Jun-Yan Zhu et al. (2016). This work demonstrates how to traverse and project onto a learned natural image manifold for attribute editing, directly inspiring CAAE's manifold-traversal approach to age progression.
  • Paper: Coupled Generative Adversarial Networks, Ming-Yu Liu et al. (2016). This paper demonstrates learning joint distributions across unaligned visual domains without paired training data, setting the precedent for CAAE's unpaired age-group mapping.
  • Paper: Deep Learning Face Attributes in the Wild, Ziwei Liu et al. (2015). This paper provides benchmark methodologies and datasets (CelebA) for learning unconstrained facial attributes like age, which form the empirical basis for training face aging models.
  • Paper: Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks, Alec Radford et al. (2016). This paper introduces stable deep convolutional architectures for generative adversarial networks, which form the structural building blocks for CAAE's convolutional encoder and deconvolutional generator.
  • Paper: Generative Adversarial Networks, Ian J. Goodfellow et al. (2014). This seminal text introduces the minimax formulation of Generative Adversarial Networks, which provides the underlying objective for both discriminator networks in CAAE.
Cover for Age Progression/Regression by Conditional Adversarial Autoencoder

Abstract

"If I provide you a face image of mine (without telling you the actual age when I took the picture) and a large amount of face images that I crawled (containing labeled faces of different ages but not necessarily paired), can you show me what I would look like when I am 80 or what I was like when I was 5?" The answer is probably a "No." Most existing face aging works attempt to learn the transformation between age groups and thus would require the paired samples as well as the labeled query image. In this paper, we look at the problem from a generative modeling perspective such that no paired samples is required. In addition, given an unlabeled image, the generative model can directly produce the image with desired age attribute. We propose a conditional adversarial autoencoder (CAAE) that learns a face manifold, traversing on which smooth age progression and regression can be realized simultaneously. In CAAE, the face is first mapped to a latent vector through a convolutional encoder, and then the vector is projected to the face manifold conditional on age through a deconvolutional generator. The latent vector preserves personalized face features (i.e., personality) and the age condition controls progression vs. regression. Two adversarial networks are imposed on the encoder and generator, respectively, forcing to generate more photo-realistic faces. Experimental results demonstrate the appealing performance and flexibility of the proposed framework by comparing with the state-of-the-art and ground truth.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Age Progression and Regression
  • 2.2 Generative Adversarial Network
  • 3 Traversing on the Manifold
  • 4 Approach
  • 4.1 Conditional Adversarial Autoencoder
  • 4.2 Objective Function
  • 4.3 Discriminator on 𝒛\boldsymbol{z}
  • 4.4 Discriminator on Face Images
  • 4.5 Differences from Other Generative Networks
  • 5 Experimental Evaluation
  • 5.1 Data Collection
  • 5.2 Implementation of CAAE
  • 5.3 Qualitative and Quantitative Comparison
  • 5.4 Tolerance to Pose, Expression, and Occlusion
  • 6 Discussion and Future Works
  • References

Knowls

  1. Knowl 1 — Conditional Adversarial Autoencoder (CAAE) Framework for Bidirectional Face Aging

    model/method

    The Conditional Adversarial Autoencoder (CAAE) is a generative framework designed for simultaneous face age progression (aging) and age regression (rejuvenation) without requiring paired longitudinal training images of the same individual across time or labeled age tags for query images at test time.

    CAAE operates on the premise that face images lie on a continuous low-dimensional manifold embedded in high-dimensional image space. The system consists of two primary operational modules:

    1. Convolutional Encoder (EE): Maps an input RGB face image x∈R128×128×3x \in \mathbb{R}^{128 \times 128 \times 3} to a latent personality vector z=E(x)∈Rnz = E(x) \in \mathbb{R}^n (with n=50n = 50). This vector zz captures high-level, age-invariant identity features.
    2. Deconvolutional Generator (GG): Synthesizes an output face image x^=G(z,l)\hat{x} = G(z, l) by conditioning the latent identity vector zz on a target age label vector ll.

    By uncoupling identity representation zz from age representation ll in the latent space, bidirectional age progression and rejuvenation are achieved by holding zz constant while traversing or specifying different target age vectors ll. During inference, only EE and GG are required.

  2. Knowl 2 — Composite Min-Max Objective Function for CAAE

    equation

    The CAAE network is trained using a composite min-max optimization objective that unifies image reconstruction, total variation regularization, latent distribution matching, and conditional adversarial image realism:

    min⁡E,Gmax⁡Dz,DimgλL(x,G(E(x),l))+γTV(G(E(x),l))+Ez∗∼p(z)[log⁡Dz(z∗)]+Ex∼pdata(x)[log⁡(1−Dz(E(x)))]+Ex,l∼pdata(x,l)[log⁡Dimg(x,l)]+Ex,l∼pdata(x,l)[log⁡(1−Dimg(G(E(x),l)))]\min_{E, G} \max_{D_z, D_{img}} \lambda \mathcal{L}\left(x, G(E(x), l)\right) + \gamma \text{TV}\left(G(E(x), l)\right) + \mathbb{E}_{z^* \sim p(z)}\left[\log D_z(z^*)\right] + \mathbb{E}_{x \sim p_{data}(x)}\left[\log\left(1 - D_z(E(x))\right)\right] + \mathbb{E}_{x, l \sim p_{data}(x, l)}\left[\log D_{img}(x, l)\right] + \mathbb{E}_{x, l \sim p_{data}(x, l)}\left[\log\left(1 - D_{img}(G(E(x), l))\right)\right]

    where:

    • x∈R128×128×3x \in \mathbb{R}^{128 \times 128 \times 3} is an input face image from the empirical data distribution pdata(x)p_{data}(x).
    • ll is the one-hot target age label vector.
    • EE is the convolutional encoder, and GG is the deconvolutional generator.
    • L(⋅,⋅)\mathcal{L}(\cdot, \cdot) denotes the L2L_2 pixel reconstruction loss: L(x,x^)=∥x−x^∥22\mathcal{L}(x, \hat{x}) = \|x - \hat{x}\|_2^2.
    • TV(⋅)\text{TV}(\cdot) denotes the total variation loss applied to the generated image to suppress ghosting artifacts.
    • p(z)p(z) is a uniform prior distribution over the latent vector space [−1,1]n[-1, 1]^n, from which z∗z^* is sampled.
    • DzD_z is the latent discriminator enforcing E(x)E(x) to conform to p(z)p(z).
    • DimgD_{img} is the conditional image discriminator distinguishing real from generated face-age pairs.
    • λ\lambda and γ\gamma are scalar balancing hyperparameters, set to λ=100\lambda = 100 and γ=10\gamma = 10.
  3. Knowl 3 — Latent Space Regularization via Adversarial Discriminator $D_z$

    model/method

    To ensure smooth traversal and interpolation in the latent identity space, an adversarial discriminator DzD_z is applied to the encoder's output z=E(x)z = E(x).

    Without DzD_z, mapping training images into the latent space leaves sparse unpopulated gaps ("holes"). Traversal or interpolation across these gaps causes the generator GG to produce distorted, non-face artifacts because intermediate latent points do not correspond to the learned face manifold.

    DzD_z is trained to distinguish between vectors z∗z^* drawn directly from a continuous uniform prior distribution p(z)=U(−1,1)p(z) = U(-1, 1) and generated latent vectors E(x)E(x). Simultaneously, EE is optimized to fool DzD_z. This adversarial competition forces the distribution of E(x)E(x) to evenly populate the latent space without holes, guaranteeing continuous identity-preserving morphing and consistent generation across arbitrary latent coordinates.

  4. Knowl 4 — Conditional Image Discriminator $D_{img}$ Architecture and Detail Enhancement

    model/method

    Relying strictly on an L2L_2 reconstruction loss during autoencoder training causes generated faces to appear blurry on unsampled test inputs because pixel-wise losses minimize average error across training instances. To generate photo-realistic textures conditional on age, CAAE incorporates a conditional image discriminator DimgD_{img}.

    DimgD_{img} receives a face image alongside its corresponding age label ll. To condition the discriminator:

    1. The 10-dimensional one-hot age vector ll is reshaped and spatially tiled/resized to 64×64×1064 \times 64 \times 10.
    2. This representation is concatenated along the channel dimension with the activation map from the first convolutional layer of DimgD_{img} (which operates on a 128×128×3128 \times 128 \times 3 input downsampled to 64×64×1664 \times 64 \times 16), producing a combined feature tensor of dimension 64×64×(10+16)64 \times 64 \times (10 + 16).

    This structure enables DimgD_{img} to evaluate both image fidelity and age-attribute alignment, forcing generator GG to synthesize realistic high-frequency textures (e.g., skin wrinkles for older target groups).

  5. Knowl 5 — Network Architecture and Training Setup of CAAE

    experimental setup

    The CAAE implementation uses the following architectural and optimization parameters:

    • Convolutional/Deconvolutional Structures: All convolutional and deconvolutional layers employ 5×55 \times 5 filters. Spatial downsampling in encoder EE and discriminator DimgD_{img} is performed using strided convolutions (stride 2) rather than max-pooling to maintain differentiability.
    • Normalization Rules: Batch normalization is omitted in EE and GG because it was found to blur personal identity features and cause output faces to drift away from query identities during testing. Conversely, batch normalization is applied within DimgD_{img} to stabilize adversarial training.
    • Activation Functions: Intermediate layers in E,G,Dz,E, G, D_z, and DimgD_{img} use ReLU activations. The output of EE (z∈R50z \in \mathbb{R}^{50}) and the output of GG (image RGB channels) use hyperbolic tangent (tanh⁡\tanh) activations, constraining values to [−1,1][-1, 1].
    • Condition Representation: Age labels are represented as 10-dimensional vectors corresponding to 10 age brackets. To match tanh⁡\tanh scaling, active categories are encoded as 11 and inactive categories as −1-1.
    • Optimization: Networks are updated sequentially using the Adam optimizer with learning rate α=0.0002\alpha = 0.0002, momentum β1=0.5\beta_1 = 0.5, mini-batch size of 100, λ=100\lambda = 100, and γ=10\gamma = 10, converging to plausible generation within roughly 50 epochs.
  6. Knowl 6 — Unpaired Age-Balanced Dataset Construction for Face Aging

    experimental setup

    To train CAAE without paired longitudinal samples across long time spans, an age-balanced dataset of 10,670 face images was assembled across 10 age categories: 0–5, 6–10, 11–15, 16–20, 21–30, 31–40, 41–50, 51–60, 61–70, and 71–80 years old.

    The dataset aggregates:

    1. Approximately 3,000 images sampled from the MORPH dataset (which contains 55,000 images of 13,000 subjects aged 16–77) and CACD (13,446 images of 2,000 subjects).
    2. 7,670 face images scraped from search engines (Bing and Google) using keyword queries (e.g., "baby", "teenager", "15 years old") to populate extreme age ranges (infants/children and older adults) absent from standard benchmarks.

    Faces were cropped, rotated, and aligned to 128×128×3128 \times 128 \times 3 resolution using a 68-landmark facial detector, ensuring uniform distribution across both age brackets and gender.

  7. Knowl 7 — Quantitative Assessment of Identity Preservation Against Ground Truth

    empirical result

    To evaluate personality and identity preservation across large age gaps, CAAE was tested on the FGNET dataset (1,002 images across 82 individuals, ages 0–69).

    Generated aged faces were paired with actual ground-truth photographs of the same subjects where the ground-truth age gap exceeded 20 years, yielding 856 evaluation pairs. A user study with 63 participants was conducted, where evaluators were presented with an original input image XX, a CAAE-generated image AA, and the ground-truth photograph BB under the matching age bracket.

    Across 3,208 recorded votes:

    • 48.38% judged that the CAAE-generated image AA was the same person as the ground truth BB.
    • 29.58% judged that they were not the same person.
    • 22.04% reported being unsure.

    These results demonstrate that CAAE reliably preserves underlying personal facial identity across extensive age spans.

  8. Knowl 8 — User Preference Study Comparing CAAE with Prior Face Aging Methods

    empirical result

    CAAE was benchmarked against prior state-of-the-art face age progression techniques (including prototype-based, dictionary learning, and recurrent neural network eigenface models) and the Face Transformer (FT) rejuvenation method using 235 paired test images from 79 subjects.

    In a blind user survey across 47 participants and 1,508 comparative votes:

    • 52.77% of votes preferred the output of CAAE over prior methods.
    • 28.99% preferred the result of the prior methods.
    • 18.24% assessed the outputs as equal in quality or hard to distinguish.

    CAAE demonstrated higher fidelity in synthetic baby/infant face rendering (avoiding the simple surface-texture-removal artifacts of baseline rejuvenation) and synthesized richer aging textures (such as authentic wrinkles) without inducing ghosting artifacts.

  9. Knowl 9 — Robustness to Pose, Expression, and Facial Occlusion

    empirical result

    Unlike prototype- and dictionary-based face aging frameworks that necessitate strict facial frontalization and landmark normalization, CAAE demonstrates tolerance to input variations including:

    • Dramatic facial expressions (e.g., open-mouth smiles or strained expressions),
    • Non-frontal profile head poses, and
    • Partial facial occlusions (e.g., skin marks, painted patterns, or foreground objects).

    CAAE projects non-standard input variations into the latent identity representation and synthesizes age-progressed or rejuvenated faces directly without requiring pre-alignment normalization or artifact removal pipelines.

Coverage note — None was omitted; all key contributions, mathematical objectives, architectural details, dataset construction strategies, and empirical evaluation results are fully covered.

References

  1. 1.D. M. Burt and D. I. Perrett. Perception of age in adult caucasian male faces: Computer graphic manipulation of shape and colour information. Proceedings of the Royal Society of London B: Biological Sciences, 259(1355):137–143, 1995.
  2. 2.B.-C. Chen, C.-S. Chen, and W. H. Hsu. Cross-age reference coding for age-invariant face recognition and retrieval. In Proceedings of the European Conference on Computer Vision, 2014.
  3. 3.X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, and P. Abbeel. InfoGAN: Interpretable representation learning by information maximizing generative adversarial nets. In Advances in Neural Information Processing Systems, 2016.
  4. 4.E. L. Denton, S. Chintala, R. Fergus, et al. Deep generative image models using a laplacian pyramid of adversarial networks. In Advances in Neural Information Processing Systems, pages 1486–1494, 2015.
  5. 5.Dlib C++ Library. http://dlib.net/. [Online].
  6. 6.Face Transformer (FT) demo. http://cherry.dcs.aber.ac.uk/transformer/. [Online].
  7. 7.Y. Fu and N. Zheng. M-face: An appearance-based photorealistic model for multiple facial attributes rendering. IEEE Transactions on Circuits and Systems for Video Technology, 16(7):830–842, 2006.
  8. 8.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems, pages 2672–2680, 2014.
  9. 9.K. Gregor, I. Danihelka, A. Graves, D. J. Rezende, and D. Wierstra. Draw: A recurrent neural network for image generation. arXiv preprint arXiv:1502.04623, 2015.
  10. 10.D. J. Im, C. D. Kim, H. Jiang, and R. Memisevic. Generating images with recurrent adversarial networks. arXiv preprint arXiv:1602.05110, 2016.
  11. 11.I. Kemelmacher-Shlizerman, S. Suwajanakorn, and S. M. Seitz. Illumination-aware age progression. In IEEE Conference on Computer Vision and Pattern Recognition, pages 3334–3341. IEEE, 2014.
  12. 12.D. Kingma and J. Ba. ADAM: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  13. 13.D. P. Kingma and M. Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  14. 14.A. Lanitis, C. J. Taylor, and T. F. Cootes. Toward automatic simulation of aging effects on face images. IEEE Transactions on Pattern Analysis and Machine Intelligence, 24(4):442–455, 2002.
  15. 15.A. B. L. Larsen, S. K. Sønderby, and O. Winther. Autoencoding beyond pixels using a learned similarity metric. arXiv preprint arXiv:1512.09300, 2015.
  16. 16.G. Levi and T. Hassner. Age and gender classification using convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 34–42, 2015.
  17. 17.M. Y. Liu and O. Tuzel. Coupled generative adversarial networks. In Advances in neural information processing systems, 2016.
  18. 18.Z. Liu, Z. Zhang, and Y. Shan. Image-based surface detail transfer. IEEE Computer Graphics and Applications, 24(3):30–35, 2004.
  19. 19.A. Makhzani, J. Shlens, N. Jaitly, and I. Goodfellow. Adversarial autoencoders. In International Conference on Learning Representations, 2016.
  20. 20.M. Mirza and S. Osindero. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784, 2014.
  21. 21.A. Radford, L. Metz, and S. Chintala. Unsupervised representation learning with deep convolutional generative adversarial networks. In International Conference on Learning Representations, 2016.
  22. 22.N. Ramanathan and R. Chellappa. Modeling age progression in young faces. In IEEE Computer Society Conference on Computer Vision and Pattern Recognition, volume 1, pages 387–394. IEEE, 2006.
  23. 23.N. Ramanathan and R. Chellappa. Modeling shape and textural variations in aging faces. In IEEE International Conference on Automatic Face & Gesture Recognition, pages 1–8. IEEE, 2008.
  24. 24.X. Shu, J. Tang, H. Lai, L. Liu, and S. Yan. Personalized age progression with aging dictionary. In Proceedings of the IEEE International Conference on Computer Vision, pages 3970–3978, 2015.
  25. 25.J. Suo, X. Chen, S. Shan, W. Gao, and Q. Dai. A concatenational graph evolution aging model. IEEE Transactions on Pattern Analysis and Machine Intelligence, 34(11):2083–2096, 2012.
  26. 26.J. Suo, S.-C. Zhu, S. Shan, and X. Chen. A compositional and dynamic model for face aging. IEEE Transactions on Pattern Analysis and Machine Intelligence, 32(3):385–401, 2010.
  27. 27.Y. Tazoe, H. Gohara, A. Maejima, and S. Morishima. Facial aging simulator considering geometry and patch-tiled texture. In ACM SIGGRAPH 2012 Posters, page 90. ACM, 2012.
  28. 28.B. Tiddeman, M. Burt, and D. Perrett. Prototyping and transforming facial textures for perception research. IEEE Computer Graphics and Applications, 21(5):42–50, 2001.
  29. 29.W. Wang, Z. Cui, Y. Yan, J. Feng, S. Yan, X. Shu, and N. Sebe. Recurrent face aging. In IEEE Conference on Computer Vision and Pattern Recognition, pages 2378–2386. IEEE, 2016.
  30. 30.X. Yu and F. Porikli. Ultra-resolving face images by discriminative generative networks. In European Conference on Computer Vision, pages 318–333. Springer, 2016.
  31. 31.J. Zhao, M. Mathieu, and Y. LeCun. Energy-based generative adversarial network. arXiv preprint arXiv:1609.03126, 2016.

Citation

MLA
Zhang, Z., et al. “Age Progression/Regression by Conditional Adversarial Autoencoder”. arXiv, 2017, http://arxiv.org/abs/1702.08423v2.
APA
Zhang, Z., Song, Y., & Qi, H. (2017). Age Progression/Regression by Conditional Adversarial Autoencoder. arXiv. http://arxiv.org/abs/1702.08423v2
Chicago
Zhang, Z., Y. Song, and H. Qi. 2017. “Age Progression/Regression by Conditional Adversarial Autoencoder”. arXiv. http://arxiv.org/abs/1702.08423v2.
Harvard
Zhang, Z., Song, Y. and Qi, H. (2017) “Age Progression/Regression by Conditional Adversarial Autoencoder”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1702.08423v2.
Vancouver
1. Zhang Z, Song Y, Qi H (2017) Age Progression/Regression by Conditional Adversarial Autoencoder. arXiv

BibTeX

@article{zhang2017age,
  title = {Age Progression/Regression by Conditional Adversarial Autoencoder},
  author = {Zhang, Zhifei and Song, Yang and Qi, Hairong},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1702.08423v2},
  eprint = {1702.08423}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/