Few-Shot Learning with Siamese Networks and Label Tuning

Thomas MüllerGuillermo Pérez-TorróMarc Franco-Salvador

article2022ACL52 citations

Introduces a scalable Siamese network framework and label tuning method for few-shot text classification that achieves constant-time inference and allows a single text encoder to be shared across multiple tasks by adapting only label embeddings.

Listen

Organizations increasingly rely on automated text classification to process customer feedback, detect sentiment, and categorize content. However, building machine learning models usually demands thousands of human-labeled examples, creating severe cost and time bottlenecks. Zero-shot and few-shot learning methods address this challenge by classifying text with zero or very few labeled examples. The standard industry approach uses textual entailment cross-attention models, but these require processing every combination of input text and possible category label sequentially. As a result, operational costs and response latency scale steeply as the number of target labels grows, while deploying customized models for distinct tasks creates significant infrastructure overhead.

The article evaluates whether Siamese dual-encoder networks provide a faster, highly competitive alternative to traditional cross-attention systems for zero-shot and few-shot text classification. In addition, the article introduces "label tuning," an efficient adaptation technique that updates only category representations rather than modifying the underlying transformer backbone.

To establish these findings, the authors conducted comprehensive empirical evaluations across 16 diverse text classification benchmarks spanning topic categorization, sentiment analysis, emotion detection, acceptability, and subjectivity. The experiments covered English, German, and Spanish datasets across varying sample sizes, specifically 0, 8, 64, and 512 training examples per label. The authors systematically benchmarked cross-attention transformers against Siamese dual-encoders trained on natural language inference datasets, assessing both accuracy and runtime throughput.

The article reveals four primary findings. First, Siamese networks achieve classification accuracy comparable to cross-attention architectures across both zero-shot and few-shot settings; Siamese models perform slightly better on zero-shot tasks (averaging 47.6% Macro F1 in English compared to 47.2% for cross-attention), while cross-attention retains a slight advantage in few-shot scenarios. Second, Siamese networks demonstrate dramatic throughput advantages: their processing speed remains constant regardless of the number of target labels (roughly 18,000 to 26,000 tokens per second), whereas cross-attention throughput drops by roughly 50% each time the number of candidate labels doubles (falling to approximately 1,150 tokens per second at 10 labels). Third, using raw category names as prompts produces results statistically indistinguishable from carefully hand-crafted prompt templates, eliminating prompt-engineering overhead. Fourth, standalone label tuning introduces an average 5-point performance drop compared to full model fine-tuning, but combining label tuning with knowledge distillation using unlabeled text almost entirely recovers this gap in practical few-shot settings.

These results demonstrate that engineering teams can transition away from computationally expensive cross-attention classifiers toward Siamese dual encoders without sacrificing predictive accuracy. Label tuning allows organizations to serve dozens of distinct classification tasks using a single shared language model in memory. By storing only compact numerical label vectors (under 10 kilobytes per task) rather than dedicated multimegabyte model weights, companies can drastically reduce cloud hosting costs, lower latency, and simplify deployment pipelines.

For production deployments involving multiple categorization tasks or large label sets, engineering leaders should adopt Siamese dual-encoder architectures and implement distilled label tuning. Teams can safely bypass manual prompt engineering and use straightforward label names. Before deploying systems at scale, practitioners should account for two specific operational limitations: Siamese networks perform slightly worse on very long text passages (over 160 tokens) and complex sentence negations compared to cross-attention models, and both model types struggle to accurately identify neutral sentiment in informal social media text. When deploying in multilingual or long-form domains, teams should run targeted validation pilots on representative business data.

No sufficiently relevant recommendations were found.

Cover for Few-Shot Learning with Siamese Networks and Label Tuning

Abstract

We study the problem of building text classifiers with little or no training data, commonly known as zero and few-shot text classification. In recent years, an approach based on neural textual entailment models has been found to give strong results on a diverse range of tasks. In this work, we show that with proper pre-training, Siamese Networks that embed texts and labels offer a competitive alternative. These models allow for a large reduction in inference cost: constant in the number of labels rather than linear. Furthermore, we introduce label tuning, a simple and computationally efficient approach that allows to adapt the models in a few-shot setup by only changing the label embeddings. While giving lower performance than model fine-tuning, this approach has the architectural advantage that a single encoder can be shared by many different tasks.

Table of Contents

  • 1 Introduction
  • 2 Methodology
  • 2.1 Fine-Tuning
  • 2.2 Label Tuning
  • 2.3 Knowledge Distillation
  • 3 Related Work
  • 4 Experimental Setup
  • 4.1 Models
  • 4.2 Datasets
  • 4.3 Hypotheses
  • 4.4 Fine-Tuning
  • 5 Results
  • 5.1 Siamese Network and Cross Attention
  • 5.2 Label Tuning
  • 6 Analysis
  • 7 Conclusion
  • Acknowledgements
  • References
  • A Unified Emotions
  • B Hypotheses
  • C Paraphrase datasets
  • D Hyperparameters
  • E Additional Analysis
  • F Computing Requirements
  • G NLI Training sets
  • H Multilingual Label Tuning Results

Knowls

  1. Knowl 1 — Siamese text classification with precomputed label embeddings

    model/method

    A Siamese Network (SN) represents an input text and each text-form label with the same transformer encoder, followed by mean pooling of contextual token embeddings. The resulting vectors are compared by dot product: for input xx and label text yy, the score is S(x,y)=f(x)⋅f(y)S(x,y)=f(x)\cdot f(y), where ff is the shared text encoder. Because input and label texts are encoded independently, class-label embeddings can be precomputed. At prediction time, the encoder processes each new input once, and its embedding is compared with the stored label embeddings. In contrast, a cross-attention (CA) classifier processes input-label pairs, requiring model execution for each candidate label. The SN therefore reduces inference-time encoder calls as the number of labels grows, while retaining a common encoder that can be shared across tasks.

  2. Knowl 2 — Batch-softmax training for Siamese Networks

    equation

    To fine-tune a Siamese Network on labeled examples, each training batch contains exactly one example for each class, so the other examples in the batch provide in-batch negative labels. For a batch of size BB, let xix_i be the input text and yiy_i its correct label text for batch position ii, and let ff be the current shared encoder. With dot-product score S(x,y)=f(x)⋅f(y)S(x,y)=f(x)\cdot f(y), the batch-softmax objective is

    J=−1B∑i=1B[S(xi,yi)−log⁡(∑j=1Bexp⁡(S(xi,yj)))].\mathcal{J}=-\frac{1}{B}\sum_{i=1}^{B}\left[S(x_i,y_i)-\log\left(\sum_{j=1}^{B}\exp(S(x_i,y_j))\right)\right].

    The encoder is updated to raise the score of each input's correct label relative to the other labels in the batch. Unlike a conventional classifier's fixed class-weight lookup, this method computes each label representation from its text during the forward pass.

  3. Knowl 3 — Label Tuning updates label vectors while freezing the encoder

    model/method

    Label Tuning (LT) adapts a pretrained Siamese Network to a task without changing the encoder parameters. Given NN labeled inputs, KK classes, and embedding dimension dd, the fixed encoder produces input matrix X∈RN×dX\in\mathbb{R}^{N\times d} and initial label matrix Y0∈RK×dY_0\in\mathbb{R}^{K\times d}. LT optimizes only the label matrix Y∈RK×dY\in\mathbb{R}^{K\times d}, using scores S=XY⊤∈RN×KS=XY^\top\in\mathbb{R}^{N\times K} and the multiclass cross-entropy loss

    LLT=−1N∑i=1N[Si,zi−log⁡(∑j=1Kexp⁡(Si,j))]+λ∥Y−Y0∥F,\mathcal{L}_{\mathrm{LT}}=-\frac{1}{N}\sum_{i=1}^{N}\left[S_{i,z_i}-\log\left(\sum_{j=1}^{K}\exp(S_{i,j})\right)\right]+\lambda\lVert Y-Y_0\rVert_F,

    where zi∈{1,…,K}z_i\in\{1,\ldots,K\} is the correct class index for input ii, Si,jS_{i,j} is its score for class jj, ∥⋅∥F\lVert\cdot\rVert_F is the Frobenius norm, and λ\lambda controls the penalty for moving the label vectors away from their initial values. The method also applies dropout by sampling a dd-dimensional binary mask, with each component zero at the selected dropout rate, and multiplying the mask component-wise into every row of YY at each gradient step. The paper selected learning rate, number of epochs, regularizer coefficient, and dropout rate using four-fold cross-validation on the few-shot training set; tested values were respectively {0.01,0.1}\{0.01,0.1\}, {1000,2000}\{1000,2000\}, {0.01,0.1}\{0.01,0.1\}, and {0.01,0.1}\{0.01,0.1\}. A tuned task requires storing its label vectors rather than a separate copy of the encoder.

  4. Knowl 4 — Distillation transfers a fine-tuned teacher to label vectors

    model/method

    To improve Label Tuning without giving up a frozen, shareable encoder, the paper uses a fine-tuned model as a teacher. The teacher produces label distributions for unlabeled examples; those distributions form a silver training set used to train label embeddings for a student whose encoder remains untuned. Experiments used up to 10,000 unlabeled examples drawn from the training data. This procedure adds teacher-training cost and requires unlabeled examples, but the resulting task adaptation still resides in the label embeddings rather than task-specific encoder weights.

  5. Knowl 5 — English zero-shot results show similar quality for Siamese and cross-attention models

    empirical result

    On 11 English text-classification datasets, the paper compared MPNET-based Siamese Networks (SN) and cross-attention models (CA) pretrained on SNLI and MNLI, using mean macro F1 across datasets. With hypothesis-pattern label texts, the zero-shot means were 47.6 for SN and 47.2 for CA; with identity hypotheses (the raw label names), they were 48.7 for SN and 47.6 for CA. The SN was significantly better on four datasets, significantly worse on four, and statistically on par on the remaining three. Thus, the mean scores were close despite mixed dataset-level outcomes. Using raw label names rather than hypothesis patterns produced comparable average performance, with the identity-hypothesis averages less than one point higher.

  6. Knowl 6 — English few-shot results keep the average gap between architectures small

    empirical result

    In English few-shot experiments, models were adapted with 8, 64, or 512 labeled examples per class. Mean macro F1 across the 11 datasets for the hypothesis-pattern MPNET models was, respectively, 62.8, 70.2, and 74.6 for cross-attention (CA), versus 61.3, 68.6, and 73.9 for Siamese Networks (SN). The authors report that CA's average advantage remained within the confidence interval. For 8 and 64 examples per class, CA significantly exceeded SN on three datasets and the models were comparable on the other eight; with 512 examples, they were comparable on six datasets, CA was better on three, and SN was better on two. Both architectures substantially outperformed a character n-gram SVM trained on the same number of examples.

  7. Knowl 7 — Multilingual experiments retain the zero-shot and few-shot pattern

    empirical result

    Across German, English, and Spanish classification datasets, the multilingual MPNET-derived models achieved mean macro F1 of 27.2 for cross-attention (CA) and 32.4 for Siamese Networks (SN) in zero-shot evaluation; corresponding identity-hypothesis means were 27.2 and 31.8. FastText averaged 26.0. In few-shot evaluation with 8, 64, and 512 examples per class, CA means were 45.7, 55.5, and 60.3, while SN means were 45.4, 54.2, and 60.3. The results therefore favor SN for zero-shot classification and show close few-shot averages, with CA's largest mean advantage at 64 examples per class. The paper reports significant CA–SN differences on only two to four datasets in the few-shot settings; both neural approaches outperformed the character SVM by a large margin.

  8. Knowl 8 — Label Tuning gains accuracy with data, and distillation recovers much of its gap

    empirical result

    For the English benchmark average, full fine-tuning (FT), Label Tuning (LT), and LT with distillation (LT-DIST) achieved macro F1 of 63.9, 57.8, and 63.3 with 8 examples per class; 69.9, 64.6, and 68.7 with 64 examples; and 74.8, 69.3, and 71.0 with 512 examples. BitFit, which updates a small subset of model parameters, averaged 61.9, 68.8, and 74.5 at those respective data sizes. Thus LT improved as labeled data increased but generally trailed FT and BitFit; distillation recovered most of the FT gap at 8 and 64 examples. In the multilingual benchmark, FT/LT/LT-DIST means were 50.6/46.5/50.0 at 8 examples, 57.9/54.7/56.9 at 64, and 62.6/58.4/59.9 at 512. For context, English zero-shot MPNET averaged 52.0, increasing to 54.0 with label refinement (LR); LR with 8 labeled examples averaged 55.8, while LT averaged 57.8. These are benchmark means, not guarantees for an individual dataset.

  9. Knowl 9 — Siamese inference is substantially faster as the label set grows

    empirical result

    Processing throughput was measured in thousand tokens per second on English test sets, grouped by label-set size of 2–3, 4–6, and 10 classes. Word-embedding models achieved 192.90, 195.82, and 208.40 thousand tokens per second; MPNET cross-attention models achieved 5.12, 2.22, and 1.15; and MPNET Siamese Networks achieved 26.08, 18.30, and 18.85. The CA throughput fell as the label set grew, whereas SN throughput did not show that scaling pattern, consistent with precomputing label vectors and encoding each input once. The measurements used an AMD Ryzen Threadripper 1950X CPU and an NVIDIA GeForce GTX 1080 Ti GPU.

  10. Knowl 10 — Evaluation protocol and pretrained model coverage

    experimental setup

    The evaluation covered diverse classification tasks—including topic, sentiment, emotion, review rating, acceptability, subjectivity, and question type—in English, German, and Spanish. Few-shot conditions used 8, 64, or 512 labeled examples per class, with five different sampled training sets for each experiment; macro F1 was the evaluation metric. Zero-shot standard deviations were estimated by bootstrapping, and significance was tested with Welch's t-test at p=0.05p=0.05. English MPNET comparisons used CA and SN models trained on SNLI and MNLI; the SN experiments also included a sentence model trained on paraphrase data. Multilingual comparisons used XLM-R-based models, with NLI training including SNLI, MNLI, adversarial NLI, and cross-lingual NLI; the latter covers 15 languages. CA fine-tuning used learning rate 10−510^{-5}, batch size 8, and 10 epochs. Siamese fine-tuning used learning rate 2×10−52\times10^{-5}, or 2×10−42\times10^{-4} in BitFit experiments.

Coverage note — Supplementary subgroup analyses by text length, hypothesis length, task type, language, label-set size, and negation are omitted because they do not change the paper's main architecture, adaptation, or benchmark conclusions.

References

  1. 1.Cecilia Ovesdotter Alm, Dan Roth, and Richard Sproat. 2005. Emotions from text: machine learning for text-based emotion prediction. In Proceedings of the conference on human language technology and empirical methods in natural language processing, pages 579–586. Association for Computational Linguistics.
  2. 2.Cecilia Ovesdotter Alm and Richard Sproat. 2005. Perceptions of emotions in expressive storytelling. In Ninth European Conference on Speech Communication and Technology.
  3. 3.Ebba Cecilia Ovesdotter Alm. 2008. Affect in* text and speech. University of Illinois at Urbana-Champaign.
  4. 4.Livio Baldini Soares, Nicholas FitzGerald, Jeffrey Ling, and Tom Kwiatkowski. 2019. Matching the blanks: Distributional similarity for relation learning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2895–2905, Florence, Italy. Association for Computational Linguistics.
  5. 5.Elad Ben-Zaken, Shauli Ravfogel, and Yoav Goldberg. 2021. Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. ArXiv, abs/2106.10199.
  6. 6.Steven Bird, Ewan Klein, and Edward Loper. 2009. Natural Language Processing with Python, 1st edition. O’Reilly Media, Inc.
  7. 7.Timo Block. 2019. Ten thousand german news articles dataset. https://tblock.github.io/10kGNAD/. Accessed: 2021-08-25.
  8. 8.Laura Ana Maria Bostan and Roman Klinger. 2018. An analysis of annotated corpora for emotion classification in text. In Proceedings of the 27th International Conference on Computational Linguistics, pages 2104–2119. Association for Computational Linguistics.
  9. 9.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  10. 10.Jane Bromley, Isabelle Guyon, Yann LeCun, Eduard Sackinger, and Roopak Shah. 1993. Signature verification using a ”siamese” time delay neural network. In Proceedings of the 6th International Conference on Neural Information Processing Systems, NIPS’93, page 737–744, San Francisco, CA, USA. Morgan Kaufmann Publishers Inc.
  11. 11.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  12. 12.Ming-Wei Chang, Lev-Arie Ratinov, D. Roth, and Vivek Srikumar. 2008. Importance of semantic representation: Dataless classification. In AAAI.
  13. 13.Zewei Chu, Karl Stratos, and Kevin Gimpel. 2021. Unsupervised label refinement improves dataless text classification. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4165–4178, Online. Association for Computational Linguistics.
  14. 14.Mark Cieliebak, Jan Milan Deriu, Dominic Egger, and Fatih Uzdilli. 2017. A Twitter corpus and benchmark resources for German sentiment analysis. In Proceedings of the Fifth International Workshop on Natural Language Processing for Social Media, pages 45–51, Valencia, Spain. Association for Computational Linguistics.
  15. 15.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. Xnli: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics.
  16. 16.Crowdflower. 2016. The emotion in text, published by crowdflower.
  17. 17.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2006. The pascal recognising textual entailment challenge. In Machine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tectual Entailment, pages 177–190, Berlin, Heidelberg. Springer Berlin Heidelberg.
  18. 18.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  19. 19.Evgeniy Gabrilovich and Shaul Markovitch. 2007. Computing semantic relatedness using wikipedia-based explicit semantic analysis. In Proceedings of the 20th International Joint Conference on Artifical Intelligence, IJCAI’07, page 1606–1611, San Francisco, CA, USA. Morgan Kaufmann Publishers Inc.
  20. 20.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830, Online. Association for Computational Linguistics.
  21. 21.Diman Ghazi, Diana Inkpen, and Stan Szpakowicz. 2015. Detecting emotion stimuli in emotion-bearing sentences. In CICLing (2), pages 152–165.
  22. 22.Edouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, and Tomas Mikolov. 2018. Learning word vectors for 157 languages. In Proceedings of the International Conference on Language Resources and Evaluation (LREC 2018).
  23. 23.Antonio Gulli. 2005. AG’s corpus of news articles. http://groups.di.unipi.it/~gulli/AG_corpus_of_news_articles.html. Accessed: 2021-07-08.
  24. 24.Kishaloy Halder, Alan Akbik, Josip Krapac, and Roland Vollgraf. 2020. Task-aware representation of sentences for generic text classification. In Proceedings of the 28th International Conference on Computational Linguistics, pages 3202–3213, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  25. 25.Karen Hambardzumyan, Hrant Khachatrian, and Jonathan May. 2021. WARP: Word-level Adversarial ReProgramming. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4921–4933, Online. Association for Computational Linguistics.
  26. 26.Marti A. Hearst, Susan T Dumais, Edgar Osuna, John Platt, and Bernhard Scholkopf. 1998. Support vector machines. IEEE Intelligent Systems and their applications, 13(4):18–28.
  27. 27.Matthew L. Henderson, Rami Al-Rfou, Brian Strope, Yun-Hsuan Sung, Laszlo Lukacs, Ruiqi Guo, Sanjiv Kumar, Balint Miklos, and Ray Kurzweil. 2017. Efficient natural language response suggestion for smart reply. CoRR, abs/1705.00652.
  28. 28.Geoffrey E. Hinton, Oriol Vinyals, and J. Dean. 2014. Distilling the knowledge in a neural network. In The NIPS 2014 Learning Semantics Workshop.
  29. 29.Phillip Keung, Yichao Lu, Gyorgy Szarvas, and Noah A. Smith. 2020. The multilingual Amazon reviews corpus. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4563–4568, Online. Association for Computational Linguistics.
  30. 30.Philipp Koehn. 2004. Statistical significance tests for machine translation evaluation. In Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing, pages 388–395, Barcelona, Spain. Association for Computational Linguistics.
  31. 31.Hugo Larochelle, Dumitru Erhan, and Yoshua Bengio. 2008. Zero-data learning of new tasks. In Proceedings of the 23rd National Conference on Artificial Intelligence - Volume 2, AAAI’08, page 646–651. AAAI Press.
  32. 32.Teven Le Scao and Alexander Rush. 2021. How many data points is a prompt worth? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2627–2636, Online. Association for Computational Linguistics.
  33. 33.Xin Li and Dan Roth. 2002. Learning question classifiers. In COLING 2002: The 19th International Conference on Computational Linguistics.
  34. 34.Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. Dailydialog: A manually labelled multi-turn dialogue dataset. arXiv preprint arXiv:1710.03957.
  35. 35.V. Liu, C. Banea, and R. Mihalcea. 2007. Grounded emotions. In International Conference on Affective Computing and Intelligent Interaction (ACII 2017), San Antonio, Texas.
  36. 36.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  37. 37.Robert L. Logan IV, Ivana Balazevic, Eric Wallace, Fabio Petroni, Sameer Singh, and Sebastian Riedel. 2021. Cutting down on prompts and parameters: Simple few-shot learning with language models. In Advances in Neural Information Processing Systems.
  38. 38.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 142–150, Portland, Oregon, USA. Association for Computational Linguistics.
  39. 39.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. In 1st International Conference on Learning Representations, ICLR 2013, Scottsdale, Arizona, USA, May 2-4, 2013, Workshop Track Proceedings.
  40. 40.Saif Mohammad. 2012. #Emotional Tweets. In *SEM 2012: The First Joint Conference on Lexical and Computational Semantics – Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation (SemEval 2012), pages 246–255, Montreal, Canada. Association for Computational Linguistics.
  41. 41.Saif M Mohammad, Parinaz Sobhani, and Svetlana Kiritchenko. 2017. Stance and sentiment in tweets. ACM Transactions on Internet Technology (TOIT), 17(3):1–23.
  42. 42.Saif M Mohammad, Xiaodan Zhu, Svetlana Kiritchenko, and Joel Martin. 2015. Sentiment, emotion, purpose, and style in electoral tweets. Information Processing & Management, 51(4):480–499.
  43. 43.Preslav Nakov, Alan Ritter, Sara Rosenthal, Veselin Stoyanov, and Fabrizio Sebastiani. 2016. SemEval-2016 task 4: Sentiment analysis in Twitter. In Proceedings of the 10th International Workshop on Semantic Evaluation, SemEval ’16, San Diego, California. Association for Computational Linguistics.
  44. 44.María Navas-Loro, Víctor Rodríguez-Doncel, Idafen Santana-Perez, and Alberto Sanchez. 2017. Spanish corpus for sentiment analysis towards brands. In Speech and Computer, pages 680–689, Cham. Springer International Publishing.
  45. 45.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020. Adversarial NLI: A new benchmark for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics.
  46. 46.Bo Pang and Lillian Lee. 2004. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. In Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics (ACL-04), pages 271–278, Barcelona, Spain.
  47. 47.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830.
  48. 48.Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021. True few-shot learning with language models. In Advances in Neural Information Processing Systems.
  49. 49.Fabio Petroni, Tim Rocktaschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2463–2473, Hong Kong, China. Association for Computational Linguistics.
  50. 50.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics.
  51. 51.K. Scherer and H. G. Wallbott. 1994. Evidence for universality and cultural variation of differential emotion response patterning. Journal of personality and social psychology, 66 2:310–28.
  52. 52.Timo Schick and Hinrich Schutze. 2021. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 255–269, Online. Association for Computational Linguistics.
  53. 53.Jake Snell, Kevin Swersky, and Richard Zemel. 2017. Prototypical networks for few-shot learning. Advances in neural information processing systems, 30.
  54. 54.Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2020. Mpnet: Masked and permuted pre-training for language understanding. In Advances in Neural Information Processing Systems, volume 33, pages 16857–16867. Curran Associates, Inc.
  55. 55.Enrica Troiano, Sebastian Pado, and Roman Klinger. 2019. Crowdsourcing and validating event-focused emotion corpora for German and English. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4005–4011, Florence, Italy. Association for Computational Linguistics.
  56. 56.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  57. 57.David Vilares and Carlos Gomez-Rodríguez. 2019. HEAD-QA: A healthcare dataset for complex reasoning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 960–966, Florence, Italy. Association for Computational Linguistics.
  58. 58.Sinong Wang, Han Fang, Madian Khabsa, Hanzi Mao, and Hao Ma. 2021. Entailment as few-shot learner. ArXiv, abs/2104.14690.
  59. 59.Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019. Neural network acceptability judgments. Transactions of the Association for Computational Linguistics, 7:625–641.
  60. 60.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122. Association for Computational Linguistics.
  61. 61.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  62. 62.Wenpeng Yin, Jamaal Hay, and Dan Roth. 2019. Benchmarking zero-shot text classification: Datasets, evaluation and entailment approach. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3914–3923, Hong Kong, China. Association for Computational Linguistics.
  63. 63.Wenpeng Yin, Nazneen Fatema Rajani, Dragomir Radev, Richard Socher, and Caiming Xiong. 2020. Universal natural language processing with limited annotations: Try few-shot textual entailment as a start. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8229–8239, Online. Association for Computational Linguistics.
  64. 64.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc.

Citation

MLA
Müller, T., et al. “Few-Shot Learning with Siamese Networks and Label Tuning”. arXiv, 2022, http://arxiv.org/abs/2203.14655v2.
APA
Müller, T., Pérez-Torró, G., & Franco-Salvador, M. (2022). Few-Shot Learning with Siamese Networks and Label Tuning. arXiv. http://arxiv.org/abs/2203.14655v2
Chicago
Müller, T., G. Pérez-Torró, and M. Franco-Salvador. 2022. “Few-Shot Learning with Siamese Networks and Label Tuning”. arXiv. http://arxiv.org/abs/2203.14655v2.
Harvard
Müller, T., Pérez-Torró, G. and Franco-Salvador, M. (2022) “Few-Shot Learning with Siamese Networks and Label Tuning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.14655v2.
Vancouver
1. Müller T, Pérez-Torró G, Franco-Salvador M (2022) Few-Shot Learning with Siamese Networks and Label Tuning. arXiv

BibTeX

@article{muller2022few,
  title = {Few-Shot Learning with Siamese Networks and Label Tuning},
  author = {Müller, Thomas and Pérez-Torró, Guillermo and Franco-Salvador, Marc},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.14655v2},
  eprint = {2203.14655}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/