keyword
siamese architecture
A Siamese architecture is an artificial neural network design consisting of two or more identical subnetworks that share the exact same configuration, parameters, and weights. Instead of classifying individual inputs into predefined categories, this architecture compares multiple inputs simultaneously by projecting them into a shared embedding space to determine their degree of similarity or semantic relationship. Each identical branch processes its respective input to generate a feature vector, and these resulting representations are evaluated using a distance metric, such as Euclidean distance or cosine similarity. Typically trained using discriminative objectives like contrastive or triplet loss to minimize distance for matching pairs while maximizing distance for non-matching pairs, Siamese architectures are widely employed in metric learning, one-shot learning, biometric verification, and text semantic matching.
3 items

LIFT: Learned Invariant Feature Transform
Kwang Moo Yi, Eduard Trulls, Vincent Lepetit, Pascal Fua
Why you should read this
Presents the first fully differentiable deep network architecture that unifies keypoint detection, orientation estimation, and descriptor generation into an end-to-end trainable feature extraction pipeline that outperforms classical methods across standard benchmarks.
We introduce a novel Deep Network architecture that implements the full feature point handling pipeline, that is, detection, orientation estimation, and feature description. While previous works have successfully tackled each one of these problems individually, we show how to learn to do all three in a unified manner while preserving end-to-end differentiability. We then demonstrate that our Deep pipeline outperforms state-of-the-art methods on a number of benchmark datasets, without the need of retraining.
Added
2026-09-25

Convolutional Neural Network Architectures for Matching Natural Language Sentences
Baotian Hu, Zhengdong Lu, Hang Li, Qingcai Chen
Why you should read this
Proposes general convolutional neural network architectures that capture both hierarchical sentence structure and multi-level semantic interactions without relying on language-specific prior knowledge, outperforming traditional baselines across diverse sentence-matching tasks.
Semantic matching is of central importance to many natural language tasks \cite{bordes2014semantic,RetrievalQA}. A successful matching algorithm needs to adequately model the internal structures of language objects and the interaction between them. As a step toward this goal, we propose convolutional neural network models for matching two sentences, by adapting the convolutional strategy in vision and speech. The proposed models not only nicely represent the hierarchical structures of sentences with their layer-by-layer composition and pooling, but also capture the rich matching patterns at different levels. Our models are rather generic, requiring no prior knowledge on language, and can hence be applied to matching tasks of different nature and in different languages. The empirical study on a variety of matching tasks demonstrates the efficacy of the proposed model on a variety of matching tasks and its superiority to competitor models.
Added
2026-09-25

Learning a similarity metric discriminatively, with application to face verification
Sumit Chopra, Raia Hadsell, Yann LeCun
Why you should read this
Proposes a discriminative metric learning method using convolutional neural networks to map images into a semantic distance space, enabling effective face verification across extreme variations in lighting, pose, and occlusion when training data per subject is scarce.
We present a method for training a similarity metric from data. The method can be used for recognition or verification applications where the number of categories is very large and not known during training, and where the number of training samples for a single category is very small. The idea is to learn a function that maps input patterns into a target space such that the L1 norm in the target space approximates the “semantic” distance in the input space. The method is applied to a face verification task. The learning process minimizes a discriminative loss function that drives the similarity metric to be small for pairs of faces from the same person, and large for pairs from different persons. The mapping from raw to the target space is a convolutional network whose architecture is designed for robustness to geometric distortions. The system is tested on the Purdue/AR face database which has a very high degree of variability in the pose, lighting, expression, position, and artificial occlusions such as dark glasses and obscuring scarves.
Added
2026-09-10
