Few-Shot Learning with Siamese Networks and Label Tuning
Thomas MüllerGuillermo Pérez-TorróMarc Franco-Salvador
Introduces a scalable Siamese network framework and label tuning method for few-shot text classification that achieves constant-time inference and allows a single text encoder to be shared across multiple tasks by adapting only label embeddings.
Organizations increasingly rely on automated text classification to process customer feedback, detect sentiment, and categorize content. However, building machine learning models usually demands thousands of human-labeled examples, creating severe cost and time bottlenecks. Zero-shot and few-shot learning methods address this challenge by classifying text with zero or very few labeled examples. The standard industry approach uses textual entailment cross-attention models, but these require processing every combination of input text and possible category label sequentially. As a result, operational costs and response latency scale steeply as the number of target labels grows, while deploying customized models for distinct tasks creates significant infrastructure overhead.
The article evaluates whether Siamese dual-encoder networks provide a faster, highly competitive alternative to traditional cross-attention systems for zero-shot and few-shot text classification. In addition, the article introduces "label tuning," an efficient adaptation technique that updates only category representations rather than modifying the underlying transformer backbone.
To establish these findings, the authors conducted comprehensive empirical evaluations across 16 diverse text classification benchmarks spanning topic categorization, sentiment analysis, emotion detection, acceptability, and subjectivity. The experiments covered English, German, and Spanish datasets across varying sample sizes, specifically 0, 8, 64, and 512 training examples per label. The authors systematically benchmarked cross-attention transformers against Siamese dual-encoders trained on natural language inference datasets, assessing both accuracy and runtime throughput.
The article reveals four primary findings. First, Siamese networks achieve classification accuracy comparable to cross-attention architectures across both zero-shot and few-shot settings; Siamese models perform slightly better on zero-shot tasks (averaging 47.6% Macro F1 in English compared to 47.2% for cross-attention), while cross-attention retains a slight advantage in few-shot scenarios. Second, Siamese networks demonstrate dramatic throughput advantages: their processing speed remains constant regardless of the number of target labels (roughly 18,000 to 26,000 tokens per second), whereas cross-attention throughput drops by roughly 50% each time the number of candidate labels doubles (falling to approximately 1,150 tokens per second at 10 labels). Third, using raw category names as prompts produces results statistically indistinguishable from carefully hand-crafted prompt templates, eliminating prompt-engineering overhead. Fourth, standalone label tuning introduces an average 5-point performance drop compared to full model fine-tuning, but combining label tuning with knowledge distillation using unlabeled text almost entirely recovers this gap in practical few-shot settings.
These results demonstrate that engineering teams can transition away from computationally expensive cross-attention classifiers toward Siamese dual encoders without sacrificing predictive accuracy. Label tuning allows organizations to serve dozens of distinct classification tasks using a single shared language model in memory. By storing only compact numerical label vectors (under 10 kilobytes per task) rather than dedicated multimegabyte model weights, companies can drastically reduce cloud hosting costs, lower latency, and simplify deployment pipelines.
For production deployments involving multiple categorization tasks or large label sets, engineering leaders should adopt Siamese dual-encoder architectures and implement distilled label tuning. Teams can safely bypass manual prompt engineering and use straightforward label names. Before deploying systems at scale, practitioners should account for two specific operational limitations: Siamese networks perform slightly worse on very long text passages (over 160 tokens) and complex sentence negations compared to cross-attention models, and both model types struggle to accurately identify neutral sentiment in informal social media text. When deploying in multilingual or long-form domains, teams should run targeted validation pilots on representative business data.
- Paper: Universal Sentence Encoder, Daniel Cer et al. (2018). Its transferable sentence-embedding approach, trained partly on natural-language-inference data, provides useful foundations for understanding the article’s Siamese encoders.
- Paper: Exploiting Cloze-Questions for Few-Shot Text Classification and Natural Language Inference, Timo Schick et al. (2020). PET’s few-shot classification and distillation pipeline supplies key context for the article’s comparison of low-data adaptation methods and its use of distillation.
- Paper: Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning, Haokun Liu et al. (2022). T-Few introduces parameter-efficient, task-specific adaptation, preparing readers to understand why the article tests updating compact label representations instead of the full model.
No sufficiently relevant recommendations were found.
