Built independently by an author, for readers. Read the story and support ChapterPal

keyword

siamese architecture

A Siamese architecture is an artificial neural network design consisting of two or more identical subnetworks that share the exact same configuration, parameters, and weights. Instead of classifying individual inputs into predefined categories, this architecture compares multiple inputs simultaneously by projecting them into a shared embedding space to determine their degree of similarity or semantic relationship. Each identical branch processes its respective input to generate a feature vector, and these resulting representations are evaluated using a distance metric, such as Euclidean distance or cosine similarity. Typically trained using discriminative objectives like contrastive or triplet loss to minimize distance for matching pairs while maximizing distance for non-matching pairs, Siamese architectures are widely employed in metric learning, one-shot learning, biometric verification, and text semantic matching.

3 items

Learning a similarity metric discriminatively, with application to face verification

Learning a similarity metric discriminatively, with application to face verification

Sumit Chopra, Raia Hadsell, Yann LeCun

OrganizationsNew York University

Why you should read this

Proposes a discriminative metric learning method using convolutional neural networks to map images into a semantic distance space, enabling effective face verification across extreme variations in lighting, pose, and occlusion when training data per subject is scarce.

We present a method for training a similarity metric from data. The method can be used for recognition or verification applications where the number of categories is very large and not known during training, and where the number of training samples for a single category is very small. The idea is to learn a function that maps input patterns into a target space such that the L1 norm in the target space approximates the “semantic” distance in the input space. The method is applied to a face verification task. The learning process minimizes a discriminative loss function that drives the similarity metric to be small for pairs of faces from the same person, and large for pairs from different persons. The mapping from raw to the target space is a convolutional network whose architecture is designed for robustness to geometric distortions. The system is tested on the Purdue/AR face database which has a very high degree of variability in the pose, lighting, expression, position, and artificial occlusions such as dark glasses and obscuring scarves.

Added

2026-09-10