Additive Margin Softmax for Face Verification

Feng WangJian ChengWeiyang LiuHaijun Liu

article2018IEEE Signal Processing Letters1,390 citationsIEEE Signal Processing Letters Best Paper Award

Proposes an additive margin Softmax loss function with feature normalization that simplifies angular margin learning and improves deep face verification accuracy on standard benchmarks such as MegaFace and LFW.

Listen

Automated face verification is essential for identity authentication across finance, defense, and public safety. Standard deep learning models rely on classification loss functions to distinguish individuals, but standard methods often struggle to create compact feature groupings for the same person while maintaining clear separation between different people. Existing approaches that enforce angular separation boundaries are computationally complex, difficult to train, and require cumbersome parameter tuning schedules.

The article demonstrates that introducing a direct additive margin to the classification loss function, combined with feature and weight normalization, creates a simpler, more interpretable, and higher-performing model for deep face verification.

The authors evaluated this additive margin approach by training deep neural networks from scratch on a standard dataset of roughly 494,000 face images across 10,575 identities. They rigorously removed overlapping identities between the training and testing sets to ensure realistic open-set evaluation. The models were tested against leading alternative loss functions across benchmark datasets, including Labeled Faces in the Wild and MegaFace, using identical network architectures.

The evaluation revealed several key findings. First, the additive margin method consistently outperformed existing state-of-the-art approaches across major benchmarks. On the challenging MegaFace identification benchmark with one million distractors, the proposed method achieved a 72.47% Rank-1 accuracy and an 84.44% verification rate, surpassing the previous leading angular method which achieved 67.41% and 78.19%, respectively. Second, performance peaked within a stable margin parameter range between 0.35 and 0.40. Third, the formulation removed the need for complex training annealing schedules, allowing models to converge stably from scratch. Finally, the analysis showed that feature normalization naturally acts as an automated hard-sample mining mechanism, assigning larger gradients to low-quality images to boost robustness.

These findings indicate that identity verification systems can achieve substantially higher accuracy under large-scale, real-world conditions without increasing architectural complexity or training instability. The additive formulation reduces the implementation and tuning overhead required to deploy robust biometric models, lowering development risk and engineering costs.

Organizations developing or upgrading facial recognition pipelines should adopt additive margin loss functions and audit their training pipelines to ensure strict overlap removal between training and testing data. Practitioners should apply feature normalization primarily when deploying models to environments characterized by low-quality imagery. Future work should focus on automatically determining optimal margin parameters and designing sample-specific or class-specific margins to further refine accuracy.

Confidence in these findings is high due to the controlled, fair comparisons across identical network backbones and benchmark datasets. However, decision-makers should note that optimal margin and normalization settings vary depending on target image quality, and extreme gradient adjustments on low-norm features present potential training risks if not properly scaled.

Cover for Additive Margin Softmax for Face Verification

Abstract

In this paper, we propose a conceptually simple and geometrically interpretable objective function, i.e. additive margin Softmax (AM-Softmax), for deep face verification. In general, the face verification task can be viewed as a metric learning problem, so learning large-margin face features whose intra-class variation is small and inter-class difference is large is of great importance in order to achieve good performance. Recently, Large-margin Softmax and Angular Softmax have been proposed to incorporate the angular margin in a multiplicative manner. In this work, we introduce a novel additive angular margin for the Softmax loss, which is intuitively appealing and more interpretable than the existing works. We also emphasize and discuss the importance of feature normalization in the paper. Most importantly, our experiments on LFW BLUFR and MegaFace show that our additive margin softmax loss consistently performs better than the current state-of-the-art methods using the same network architecture and training dataset. Our code has also been made available at this https URL

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 Additive Margin Softmax
  • 3.1 Definition
  • 3.2 Discussion
  • 3.2.1 Geometric Interpretation
  • 3.2.2 Angular Margin or Cosine Margin
  • 3.2.3 Feature Normalization
  • 3.2.4 Feature Distribution Visualization
  • 4 Experiment
  • 4.1 Implementation Details
  • 4.2 Dataset Overlap Removal
  • 4.3 Effect of Hyper-parameter mm
  • 5 Conclusion and Future Work
  • References

Knowls

  1. Knowl 1 — Additive Margin Softmax (AM-Softmax) Loss Formulation

    model/method

    Let fi∈Rdf_i \in \mathbb{R}^d denote the deep feature representation (the input to the last fully connected classification layer) for the ii-th training sample, and let yi∈{1,…,c}y_i \in \{1, \dots, c\} be its ground-truth class label among cc classes. Let Wj∈RdW_j \in \mathbb{R}^d denote the weight vector corresponding to the jj-th class in the classification layer.

    Both the weight vectors and the feature vectors are normalized to unit ℓ2\ell_2 norm: ∥Wj∥2=1\|W_j\|_2 = 1 for all j∈{1,…,c}j \in \{1, \dots, c\}, and ∥fi∥2=1\|f_i\|_2 = 1. The cosine similarity between weight WjW_j and feature fif_i is cos⁡θj=WjTfi\cos\theta_j = W_j^T f_i.

    The Additive Margin Softmax (AM-Softmax) loss is defined as:

    LAMS=−1n∑i=1nlog⁡es(cos⁡θyi−m)es(cos⁡θyi−m)+∑j=1,j≠yicescos⁡θj=−1n∑i=1nlog⁡es(WyiTfi−m)es(WyiTfi−m)+∑j=1,j≠yicesWjTfiL_{\text{AMS}} = -\frac{1}{n} \sum_{i=1}^n \log \frac{e^{s(\cos\theta_{y_i} - m)}}{e^{s(\cos\theta_{y_i} - m)} + \sum_{j=1, j \ne y_i}^c e^{s \cos\theta_j}} = -\frac{1}{n} \sum_{i=1}^n \log \frac{e^{s(W_{y_i}^T f_i - m)}}{e^{s(W_{y_i}^T f_i - m)} + \sum_{j=1, j \ne y_i}^c e^{s W_j^T f_i}}

    where:

    • nn is the mini-batch size,
    • s>0s > 0 is a fixed global scale parameter scaling the cosine similarities (typically s=30s = 30),
    • m>0m > 0 is a fixed additive angular margin scalar subtracted directly from the target class cosine similarity (typically m∈[0.35,0.4]m \in [0.35, 0.4]).

    In implementation, letting x=cos⁡θyi=WyiTfix = \cos\theta_{y_i} = W_{y_i}^T f_i, the modified target similarity is Ψ(x)=x−m\Psi(x) = x - m. Because the derivative is constant (Ψ′(x)=1\Psi'(x) = 1), the gradient flows directly through backpropagation without requiring piecewise angular formulations or dynamic parameter annealing strategies.

  2. Knowl 2 — Geometric Decision Boundary of AM-Softmax

    theoretical result

    On the unit hypersphere manifold where ∥W1∥2=∥W2∥2=∥f∥2=1\|W_1\|_2 = \|W_2\|_2 = \|f\|_2 = 1, the conventional softmax loss sets the decision boundary between class 1 and class 2 at a vector P0P_0 satisfying W1TP0=W2TP0W_1^T P_0 = W_2^T P_0, or equivalently cos⁡(θW1,P0)=cos⁡(θW2,P0)\cos(\theta_{W_1, P_0}) = \cos(\theta_{W_2, P_0}).

    In AM-Softmax, the decision boundary for class 1 is shifted to a vector P1P_1 satisfying:

    W1TP1−m=W2TP1  ⟺  cos⁡(θW1,P1)−m=cos⁡(θW2,P1)W_1^T P_1 - m = W_2^T P_1 \iff \cos(\theta_{W_1, P_1}) - m = \cos(\theta_{W_2, P_1})

    which yields:

    m=(W1−W2)TP1=cos⁡(θW1,P1)−cos⁡(θW2,P1)m = (W_1 - W_2)^T P_1 = \cos(\theta_{W_1, P_1}) - \cos(\theta_{W_2, P_1})

    Assuming equal intra-class variance across classes, the decision boundary for class 2 lies at vector P2P_2 such that cos⁡(θW2,P1)=cos⁡(θW1,P2)\cos(\theta_{W_2, P_1}) = \cos(\theta_{W_1, P_2}). The additive margin mm therefore equals the difference between the cosine similarities of class 1 evaluated at the two boundary extremes of the margin region:

    m=cos⁡(θW1,P1)−cos⁡(θW1,P2)m = \cos(\theta_{W_1, P_1}) - \cos(\theta_{W_1, P_2})

    This creates a hard geometric angular separation band on the hypersphere, forcing the intra-class variance to shrink and expanding inter-class separation.

  3. Knowl 3 — Benchmark Verification and Identification Performance Across Loss Functions

    data/table

    The table below compares the performance of deep convolutional neural networks trained on the CASIA-WebFace dataset (with overlapping identities removed) across different classification loss functions on the LFW 6,000-pair benchmark, the LFW BLUFR protocol, and the MegaFace Set 1 benchmark (with 1 million distractors).

    Loss Function mm LFW 6,000 pairs LFW BLUFR VR@FAR=0.01% LFW BLUFR VR@FAR=0.1% LFW BLUFR DIR@FAR=1% MegaFace Rank1@1e6 MegaFace VR@FAR=1e-6
    Softmax - 97.08% 60.26% 78.26% 50.85% 45.26% 50.12%
    Softmax+75% dropout - 98.62% 77.64% 90.91% 63.72% 57.32% 65.58%
    Center Loss - 99.00% 83.30% 94.50% 65.46% 63.38% 75.68%
    NormFace - 98.98% 88.15% 96.16% 75.22% 65.03% 75.88%
    A-Softmax ∼1.5\sim 1.5 99.08% 91.26% 97.06% 81.93% 67.41% 78.19%
    AM-Softmax 0.25 99.13% 91.97% 97.13% 81.42% 70.81% 83.01%
    AM-Softmax 0.30 99.08% 93.18% 97.56% 84.02% 72.01% 83.29%
    AM-Softmax 0.35 98.98% 93.51% 97.69% 84.82% 72.47% 84.44%
    AM-Softmax 0.40 99.17% 93.60% 97.71% 84.51% 72.44% 83.50%
    AM-Softmax 0.45 99.03% 93.44% 97.60% 84.59% 72.22% 83.00%
    AM-Softmax 0.50 99.10% 92.33% 97.28% 83.38% 71.56% 82.49%
    AM-Softmax w/o FN 0.35 99.08% 93.86% 97.63% 87.58% 70.71% 82.66%
    AM-Softmax w/o FN 0.40 99.12% 94.48% 97.96% 87.31% 70.96% 83.11%

    Center Loss and NormFace utilized a modified ResNet-28, whereas all other entries used a modified ResNet-20. AM-Softmax with m=0.35m=0.35 achieves 72.47%72.47\% Rank-1 accuracy and 84.44%84.44\% verification rate on MegaFace, substantially outperforming SphereFace / A-Softmax (67.41%67.41\% and 78.19%78.19\%) and Center Loss (63.38%63.38\% and 75.68%75.68\%).

  4. Knowl 4 — Gradient Dynamics and Hard Sample Mining Property of Feature Normalization

    theoretical result

    Applying feature normalization to map an unnormalized feature vector xx to the unit sphere y=x∥x∥2=xαy = \frac{x}{\|x\|_2} = \frac{x}{\alpha} (where α=∥x∥2\alpha = \|x\|_2) yields a backpropagation gradient scale that is inversely proportional to the original feature norm:

    dydx=1α\frac{dy}{dx} = \frac{1}{\alpha}

    Consequently, samples with smaller ℓ2\ell_2 feature norms receive substantially larger gradient updates during backpropagation than samples with larger feature norms. Because feature magnitude is positively correlated with image quality (e.g., clear, well-aligned frontal images exhibit large norms, whereas blurry, degraded, or occluded faces exhibit small norms), feature normalization functions as an implicit hard-sample mining mechanism. This makes feature normalization particularly advantageous for challenging, low-quality image datasets like MegaFace, while training without feature normalization performs better on clean, high-quality benchmarks like LFW.

  5. Knowl 5 — Hyperparameter Sensitivity of Additive Margin $m$ and Feature Normalization Ablation

    empirical result

    Ablation over the additive margin parameter m∈[0.25,0.50]m \in [0.25, 0.50] (with scaling parameter s=30s=30 on modified ResNet-20) demonstrates:

    1. Optimal Margin Interval: Recognition performance increases sharply from m=0.25m=0.25 to m=0.35m=0.35 (e.g., MegaFace Rank-1 rises from 70.81%70.81\% to 72.47%72.47\%, and VR@FAR=10−610^{-6} increases from 83.01%83.01\% to 84.44%84.44\%). Peak performance is achieved within m∈[0.35,0.40]m \in [0.35, 0.40]. When m>0.40m > 0.40, performance gradually declines (at m=0.50m=0.50, MegaFace Rank-1 drops to 71.56%71.56\%).
    2. Feature Normalization Trade-Off: AM-Softmax with feature normalization (FN) outperforms the non-normalized variant (w/o FN) on unconstrained low-quality images (MegaFace Rank-1 of 72.47%72.47\% with FN vs. 70.71%70.71\% w/o FN at m=0.35m=0.35). Conversely, AM-Softmax without feature normalization yields superior verification metrics on high-quality images (LFW BLUFR DIR@FAR=1% of 87.58%87.58\% w/o FN vs. 84.82%84.82\% with FN at m=0.35m=0.35, and 87.31%87.31\% vs. 84.51%84.51\% at m=0.40m=0.40).
  6. Knowl 6 — Impact of Training-Test Identity Overlap Removal on MegaFace Evaluation

    empirical result

    Cross-dataset identity overlap between training and testing sets introduces substantial evaluation bias. An identity audit between CASIA-WebFace (494,414 images, 10,575 identities) and evaluation benchmarks revealed:

    • 17 overlapping identities with LFW,
    • 42 overlapping identities with MegaFace Set 1 (which comprises only 80 target identities, meaning over 50%50\% of target identities were present in the training set).

    Evaluating the identical modified ResNet-20 model trained with AM-Softmax (m=0.35,s=30m=0.35, s=30) before and after removing the 42 overlapping identities from CASIA-WebFace shows:

    Loss Function Overlap Removal? MegaFace Rank1 MegaFace VR
    AM-Softmax No 75.23% 87.06%
    AM-Softmax Yes 72.47% 84.44%

    Removing the overlapping identities decreases the Rank-1 rate by 2.76%2.76\% (from 75.23%75.23\% to 72.47%72.47\%) and the Verification Rate at FAR=10−610^{-6} by 2.62%2.62\% (from 87.06%87.06\% to 84.44%84.44\%), demonstrating that uncleaned training sets artificially inflate benchmark accuracy.

  7. Knowl 7 — Face Verification Training and Evaluation Pipeline

    experimental setup

    The standard experimental configuration for Additive Margin Softmax is defined as follows:

    • Dataset: CASIA-WebFace cleaned of identity overlaps with LFW (17 identities removed) and MegaFace Set 1 (42 identities removed), containing 494,414 training images across 10,575 identities.
    • Preprocessing: Faces and 5 facial landmarks are detected via MTCNN. Aligned face crops are resized to 112×96112 \times 96 pixels and normalized by subtracting 128 and dividing by 128. Data augmentation consists exclusively of random horizontal flipping.
    • Architecture: Modified ResNet with 20 layers adapted for face recognition (using modified ResNet-28 for Center Loss and NormFace baselines).
    • Optimization: Models are trained from scratch using Caffe with batch size 256, weight decay 5×10−45 \times 10^{-4}, and initial learning rate 0.1 divided by 10 at 16K, 24K, and 28K iterations, terminating at 30K iterations.
    • Testing Metric: Feature vectors from the first inner-product layer are extracted for both the original and horizontally mirrored test images and summed to produce the final representation. Pairwise similarity is evaluated using cosine similarity.
  8. Knowl 8 — Limitations and Open Problems of AM-Softmax

    limitation

    AM-Softmax presents the following specific methodological limitations:

    • Gradient Instability with Very Small Feature Norms: In normalized feature learning, the feature gradient norm scales inversely with the original feature norm (∝1/∥x∥2\propto 1/\|x\|_2). When feature norms are exceptionally small, gradient magnitudes can become excessively large, increasing the risk of numerical instability and gradient explosion during optimization.
    • Learnable Scale Factor Degeneracy: When the scale parameter ss is set as a learnable parameter optimized by backpropagation alongside the additive margin mm, ss fails to increase and optimization converges extremely slowly. Consequently, ss must be manually tuned and fixed as a hyperparameter (e.g., s=30s = 30).
    • Uniform Global Margin: The margin parameter mm is a static global hyperparameter applied uniformly across all classes and samples, rather than dynamically adapting to class-specific distributions or sample-specific difficulty levels.

Coverage note — The illustrative 3D hypersphere feature visualization on Fashion-MNIST was omitted as it serves as a qualitative demonstration of the geometric boundary and variance shrinkage properties captured in the loss formulation and geometric boundary knowls.

References

  1. 1.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016. 5
  2. 2.G. Hinton, O. Vinyals, and J. Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. 1
  3. 3.G. B. Huang, M. Ramesh, T. Berg, and E. Learned-Miller. Labeled faces in the wild: A database for studying face recognition in unconstrained environments. Technical report, Technical Report 07-49, University of Massachusetts, Amherst, 2007. 5
  4. 4.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. In Proceedings of the 22nd ACM international conference on Multimedia, pages 675–678. ACM, 2014. 4
  5. 5.I. Kemelmacher-Shlizerman, S. M. Seitz, D. Miller, and E. Brossard. The megaface benchmark: 1 million faces for recognition at scale. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4873–4882, 2016. 1, 5, 6
  6. 6.X. Liang, X. Wang, Z. Lei, S. Liao, and S. Z. Li. Soft-margin softmax for deep classification. 24th International Conference on Neural Information Processing, pages 413–421, 2017. 1, 3
  7. 7.S. Liao, Z. Lei, D. Yi, and S. Z. Li. A benchmark study of large-scale unconstrained face recognition. In IEEE International Joint Conference on Biometrics, pages 1–8. IEEE, 2014. 1, 5, 6
  8. 8.T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár. Focal loss for dense object detection. arXiv preprint arXiv:1708.02002, 2017. 4
  9. 9.W. Liu, Y. Wen, Z. Yu, M. Li, B. Raj, and L. Song. Sphereface: Deep hypersphere embedding for face recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2017. 1, 2, 3, 4, 5, 6
  10. 10.W. Liu, Y. Wen, Z. Yu, and M. Yang. Large-margin softmax loss for convolutional neural networks. In International Conference on Machine Learning, pages 507–516, 2016. 1, 2, 3, 6
  11. 11.W. Liu, Y.-M. Zhang, X. Li, Z. Yu, B. Dai, T. Zhao, and L. Song. Deep hyperspherical learning. In Advances in Neural Information Processing Systems, pages 3953–3963, 2017. 2, 4
  12. 12.Y. Liu, H. Li, and X. Wang. Rethinking feature discrimination and polymerization for large-scale recognition. arXiv preprint arXiv:1710.00870, 2017. 1, 2, 3, 5
  13. 13.O. M. Parkhi, A. Vedaldi, and A. Zisserman. Deep face recognition. In BMVC, volume 1, page 6, 2015. 1
  14. 14.G. Pereyra, G. Tucker, J. Chorowski, Ł. Kaiser, and G. Hinton. Regularizing neural networks by penalizing confident output distributions. arXiv preprint arXiv:1701.06548, 2017. 2
  15. 15.R. Ranjan, C. D. Castillo, and R. Chellappa. L2-constrained softmax loss for discriminative face verification. arXiv preprint arXiv:1703.09507, 2017. 1, 3, 4, 5
  16. 16.F. Schroff, D. Kalenichenko, and J. Philbin. Facenet: A unified embedding for face recognition and clustering. In IEEE Conference on Computer Vision and Pattern Recognition, pages 815–823, 2015. 1, 4
  17. 17.Y. Sun, Y. Chen, X. Wang, and X. Tang. Deep learning face representation by joint identification-verification. In Advances in neural information processing systems, pages 1988–1996, 2014. 1
  18. 18.Y. Taigman, M. Yang, M. Ranzato, and L. Wolf. Deepface: Closing the gap to human-level performance in face verification. In IEEE Conference on Computer Vision and Pattern Recognition, pages 1701–1708, 2014. 1
  19. 19.F. Wang, X. Xiang, J. Cheng, and A. L. Yuille. Normface: L2 hypersphere embedding for face verification. In Proceedings of the 25th ACM international conference on Multimedia. ACM, 2017. 1, 2, 3, 5
  20. 20.Y. Wen, K. Zhang, Z. Li, and Y. Qiao. A discriminative feature learning approach for deep face recognition. In European Conference on Computer Vision, pages 499–515. Springer, 2016. 1, 5, 6
  21. 21.H. Xiao, K. Rasul, and R. Vollgraf. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017. 4
  22. 22.D. Yi, Z. Lei, S. Liao, and S. Z. Li. Learning face representation from scratch. arXiv preprint arXiv:1411.7923, 2014. 5
  23. 23.Y. Yuan, K. Yang, and C. Zhang. Feature incay for representation regularization. arXiv preprint arXiv:1705.10284, 2017.
  24. 24.K. Zhang, Z. Zhang, Z. Li, and Y. Qiao. Joint face detection and alignment using multitask cascaded convolutional networks. IEEE Signal Processing Letters, 23(10):1499–1503, 2016. 4

Citation

MLA
Wang, F., et al. “Additive Margin Softmax for Face Verification”. IEEE Signal Processing Letters, vol. 25, no. 7, 2018, pp. 926–30, https://doi.org/10.1109/LSP.2018.2822810.
APA
Wang, F., Cheng, J., Liu, W., & Liu, H. (2018). Additive Margin Softmax for Face Verification. IEEE Signal Processing Letters, 25(7), 926–930. https://doi.org/10.1109/LSP.2018.2822810
Chicago
Wang, F., J. Cheng, W. Liu, and H. Liu. 2018. “Additive Margin Softmax for Face Verification”. IEEE Signal Processing Letters 25 (7): 926–30. https://doi.org/10.1109/LSP.2018.2822810.
Harvard
Wang, F. et al. (2018) “Additive Margin Softmax for Face Verification”, IEEE Signal Processing Letters, 25(7), pp. 926–930. Available at: https://doi.org/10.1109/LSP.2018.2822810.
Vancouver
1. Wang F, Cheng J, Liu W, Liu H (2018) Additive Margin Softmax for Face Verification. IEEE Signal Processing Letters 25:926–930

BibTeX

@article{Wang_2018, title={Additive Margin Softmax for Face Verification}, volume={25}, ISSN={1558-2361}, url={http://dx.doi.org/10.1109/LSP.2018.2822810}, DOI={10.1109/lsp.2018.2822810}, number={7}, journal={IEEE Signal Processing Letters}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Wang, Feng and Cheng, Jian and Liu, Weiyang and Liu, Haijun}, year={2018}, month=July, pages={926–930} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-sa/4.0/