A Closer Look at How Fine-tuning Changes BERT

Yichu ZhouVivek Srikumar

article2022ACL87 citations

Reveals how fine-tuning improves BERT representations by widening separation distances between distinct label clusters while preserving the model's underlying geometric structure across downstream tasks.

Listen

Context and problem: Pre-trained language models have become the standard foundation for modern natural language processing applications. Organizations routinely adapt these models to specific downstream tasks through a process called fine-tuning. While fine-tuning almost always improves performance on practical tasks, the underlying mechanics of why it succeeds and how it alters the internal data representations have remained an unexamined black box.

Objective: The article evaluates how fine-tuning alters the underlying geometric structure of language representations and examines whether the common assumption that fine-tuning always improves performance holds true across tasks and model sizes.

Approach: The authors conducted systematic experiments on the English BERT model family across five distinct language tasks, spanning syntax, semantic disambiguation, and text classification. To analyze the models, the study combined two probing methods: training supervised two-layer neural network classifiers on frozen representations to evaluate predictive accuracy, and applying a geometric probing technique called DIRECTPROBE. This geometric technique measures data complexity by tracking how data clusters group together, calculating Euclidean distances between different label groups, and assessing spatial similarity across training and testing splits as well as across individual model layers.

Key findings: The analysis revealed four core findings:

  1. Fine-tuning simplifies representation geometry: When initial data points cannot be separated by simple linear boundaries, fine-tuning groups points sharing the same label into fewer, consolidated clusters. When labels are already linearly separable, fine-tuning actively pushes distinct label clusters farther apart, creating wider separation margins that accommodate more robust decision boundaries.
  2. Fine-tuning is not universally beneficial: Fine-tuning consistently increases geometric divergence between training and test sets. In one observed case—a smaller BERT model fine-tuned on preposition semantic function prediction—this divergence was severe enough (similarity dropping to 0.44) that performance decreased from 86.26% to 85.08%.
  3. Cross-task fine-tuning exhibits transfer dynamics based on task alignment: Fine-tuning on a related task increases cluster distances and improves target performance, whereas fine-tuning on a conflicting task shrinks cluster distances (by 1.68 on average) and degrades target accuracy from 87.75% to 83.24%.
  4. Higher model layers preserve foundational structure: While higher model layers change substantially more than lower layers, they do not change arbitrarily; they maintain a spatial correlation greater than 0.5 with the original pre-trained space, preserving core structural relationships while making targeted adjustments.

Implications and interpretation: These findings explain why fine-tuning is widely effective: it improves model generalization by geometrically enlarging separation margins between task categories rather than simply memorizing training instances. However, the discovery that fine-tuning introduces training-test divergence introduces a potential risk of representation-level overfitting. For practitioners and decision-makers, this highlights that task adaptation is not purely additive; applying a model fine-tuned on an opposing or misaligned objective can inadvertently strip out valuable capabilities and degrade downstream performance.

Recommendations and next steps: Engineering teams deploying compact or lightweight models should pair them with non-linear classification heads rather than simple linear classifiers, as smaller models retain more complex, non-linear geometric structures. Additionally, organizations should monitor validation performance closely rather than assuming fine-tuning is strictly beneficial, particularly when adapting models across distinct or potentially conflicting tasks. Prior to establishing new operational guidelines, teams should conduct further exploratory work to determine whether tracking geometric divergence can serve as an early-warning metric to prevent fine-tuning failures.

Limitations and confidence: Confidence in the geometric behavior described is high across the evaluated BERT architectures and tasks. However, key limitations remain: the study is restricted to English-language BERT variants, excludes the largest model due to training instability, and focuses primarily on cluster-level distances rather than the internal distribution within clusters. Readers should exercise caution before generalizing these specific geometric thresholds to non-transformer architectures or non-English datasets without pilot validation.

Zhou et al (2022).pdf
Cover for A Closer Look at How Fine-tuning Changes BERT

Abstract

Given the prevalence of pre-trained contextualized representations in today's NLP, there have been many efforts to understand what information they contain, and why they seem to be universally successful. The most common approach to use these representations involves fine-tuning them for an end task. Yet, how fine-tuning changes the underlying embedding space is less studied. In this work, we study the English BERT family and use two probing techniques to analyze how fine-tuning changes the space. We hypothesize that fine-tuning affects classification performance by increasing the distances between examples associated with different labels. We confirm this hypothesis with carefully designed experiments on five different NLP tasks. Via these experiments, we also discover an exception to the prevailing wisdom that "fine-tuning always improves performance". Finally, by comparing the representations before and after fine-tuning, we discover that fine-tuning does not introduce arbitrary changes to representations; instead, it adjusts the representations to downstream tasks while largely preserving the original spatial structure of the data points.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries: Probing Methods
  • 2.1 Classifiers as Probes
  • 2.2 DIRECTPROBE: Probing the Geometric Structure
  • 3 Experimental Setup
  • 3.1 Representations
  • 3.2 Tasks
  • 3.3 Fine-tuning Setup
  • 4 Observations and Analysis
  • 4.1 Fine-tuned Performance
  • 4.2 Linearity of Representations
  • 4.3 Spatial Structure of Labels
  • 4.4 Cross-task Fine-tuning
  • 4.5 Layer Behavior
  • 5 Discussion
  • 6 Related Work
  • 7 Conclusions
  • Acknowledgments
  • References
  • A Fine-tuning Details
  • B Summary of Tasks
  • C Probing Performance
  • D Dynamics of Minimum Distances
  • E PCA Projections of the Movements
  • F Cluster Number Revision

Knowls

  1. Knowl 1 — Fine-tuning widens label-cluster gaps in linearly separable representations

    empirical result

    For BERTbase on part-of-speech tagging (POS), dependency-relation prediction (DEP), preposition supersense function (PS-fxn), and preposition supersense role (PS-role), the representations were linearly separable before fine-tuning. DirectProbe therefore identified one cluster per label. Across these tasks, fine-tuning increased each label cluster’s minimum Euclidean distance to the other label clusters overall, although a distance could dip during training; for example, the minimum distance for STUFF in PS-role temporarily decreased. In POS, the centroids of three initially nearby label clusters moved in different directions as fine-tuning progressed. The authors propose that wider gaps allow more classifiers to separate the labels and may explain improved generalization; this is an interpretation of the observed geometry, not a theoretical guarantee.

  2. Knowl 2 — Cross-task fine-tuning changes target-task distances in ways associated with probe accuracy

    empirical result

    The authors fine-tuned BERTbase separately on PS-fxn, the related task PS-role, or POS, then froze each resulting representation and trained a classifier probe for PS-fxn without further fine-tuning. For each PS-fxn label, they measured the change in its minimum Euclidean distance to another label cluster. Fine-tuning on PS-fxn increased all 40 measured minimum distances; PS-role increased distances for most labels; POS decreased them for all 40. The corresponding probe accuracies and mean distance changes were:

    Fine-tuning taskPS-fxn labels with distance increasedPS-fxn labels with distance decreasedMean distance changePS-fxn probe accuracy
    None———87.75
    PS-fxn4005.2989.58
    PS-role27131.0288.53
    POS040-1.6883.24

    The pattern is consistent with task-related fine-tuning helping when it enlarges target-label gaps, and harming when it shrinks them. The authors interpret this as evidence that cross-task fine-tuning can add or remove information useful to a target task, depending on the relationship between tasks.

  3. Knowl 3 — Fine-tuning often simplifies nonlinear label geometry

    empirical result

    DirectProbe represents a task’s labeled examples as non-overlapping convex clusters, each containing examples of only one label. If the number of clusters equals the number of labels, the representation is linearly separable; more clusters indicate that at least one label is split across multiple regions. For BERTtiny, fine-tuning reduced the cluster count on every task with initially nonlinear geometry, though it did not always make the representation linear. The probe accuracy values below are reported scores for the original and fine-tuned representations.

    TaskClusters, original → fine-tunedLinear, original → fine-tunedProbe accuracy, original → fine-tuned
    POS30 → 18No → No90.76 → 91.67
    DEP50 → 46No → Yes86.74 → 89.04
    PS-fxn42 → 40No → Yes74.14 → 74.40
    PS-role46 → 46Yes → Yes58.38 → 60.31
    TREC-5058 → 51No → No68.12 → 84.04

    The authors’ interpretation is that fine-tuning can gather same-label examples into fewer regions, simplifying a representation that was not linearly separable. For TREC-50, the number of clusters also fell for BERTmini, BERTsmall, and BERTbase, but remained 52 for BERTmedium; TREC-50 remained nonlinear for all five BERT sizes.

  4. Knowl 4 — Probe accuracy generally improves after fine-tuning, with one observed exception

    data/table

    The study compared frozen original and task-fine-tuned BERT representations using separately trained classifier probes on five NLP tasks. The table reports the probe accuracy scores in the paper; each entry is original → fine-tuned. Fine-tuning increased accuracy in 24 of the 25 model–task comparisons. The sole decrease was BERTsmall on PS-fxn, from 86.26 to 85.08.

    ModelPOSDEPPS-fxnPS-roleTREC-50
    BERTtiny90.76 → 91.6786.74 → 89.0474.14 → 74.4058.38 → 60.3168.12 → 84.04
    BERTmini93.81 → 94.9191.82 → 93.5582.45 → 84.2568.05 → 71.9074.12 → 88.36
    BERTsmall94.26 → 95.4392.93 → 94.4886.26 → 85.0874.22 → 74.5781.32 → 89.60
    BERTmedium94.40 → 95.5692.54 → 94.7686.56 → 88.4576.28 → 78.8680.68 → 89.80
    BERTbase93.39 → 95.6889.39 → 94.7687.75 → 89.5874.49 → 81.1485.24 → 90.36

    These are probing results, not scores from the fine-tuning task heads: probes were trained from scratch on each frozen representation. The authors caution that the single counterexample is insufficient to establish why fine-tuning hurt in that setting.

  5. Knowl 5 — Fine-tuning reduces training–test spatial similarity, including in the performance exception

    empirical result

    For each linearly separable task representation, the authors applied DirectProbe separately to the training and test examples and correlated the resulting vectors of pairwise label-cluster distances using Pearson correlation. This spatial-similarity score decreased after fine-tuning in every reported comparison, indicating that the training and test sets’ relative label geometry became less alike. For BERTsmall, the original → fine-tuned probe accuracies and training–test spatial similarities were:

    TaskProbe accuracy, original → fine-tunedTraining–test spatial similarity, original → fine-tuned
    POS94.26 → 95.430.96 → 0.72
    DEP92.93 → 94.480.93 → 0.78
    PS-fxn86.26 → 85.080.82 → 0.44
    PS-role74.22 → 74.570.84 → 0.54
    TREC-5081.32 → 89.60Not computed

    TREC-50 was not assigned a spatial-similarity score because its representations were nonlinear. The BERTsmall PS-fxn exception had the lowest fine-tuned training–test similarity in this comparison, 0.44. The authors conjecture that limiting training–test divergence might help preserve fine-tuning gains, but state that this relationship requires further study.

  6. Knowl 6 — DirectProbe operationalizes linearity, label-cluster distance, and spatial similarity

    model/method

    DirectProbe analyzes a representation for a specified labeling task by producing clusters whose members share a label and whose convex hulls do not overlap. If the number of clusters equals the number of labels, each label occupies one cluster and a linear multiclass classifier suffices; if there are more clusters than labels, at least one label is split across regions and a nonlinear classifier is needed. The study used this cluster count to track changes in label geometry.

    For two clusters, the authors trained a maximum-margin linear SVM and defined their Euclidean distance as twice the SVM margin. When there was exactly one cluster per label, they formed a distance vector from the distances for all unordered pairs of labels. With nn labels, this vector has n(n−1)/2n(n-1)/2 entries. They measured spatial similarity between two representations for the same task as the Pearson correlation between their distance vectors. This measure captures the relative arrangement of label clusters, rather than the geometry of points within a cluster.

  7. Knowl 7 — Experiments span five English BERT sizes and five syntax and semantics tasks

    experimental setup

    The experiments used uncased English BERT models and five tasks. For subword-tokenized words, the representation was the average of the subword embeddings. POS and dependency-relation examples came from the English Universal Dependencies Parallel Universal Dependencies Treebank; dependency examples represented a token pair by concatenating the two contextualized token vectors. Preposition supersense role and function were trained and evaluated on single-token prepositions from Streusle v4.2. TREC-50 used the [CLS] vector as the sentence representation.

    TaskTraining examplesTest examplesLabels
    POS16,8604,32317
    Dependency relation16,0544,12246
    Preposition supersense role4,28245747
    Preposition supersense function4,28245740
    TREC-505,45250050
    ModelLayersAttention headsHidden dimensionParameters
    BERTtiny221284.4M
    BERTmini4425611.3M
    BERTsmall4851229.1M
    BERTmedium8851241.7M
    BERTbase1212768110.1M

    Each model was fine-tuned separately on each task. Fine-tuning used Adam, batch size 32, a linear learning-rate schedule with learning rate 3×10−43\times10^{-4} and 10% of update steps for warmup, and one Titan GPU. Models were fine-tuned for 10 epochs, except BERTbase, which was fine-tuned for three. BERTlarge was excluded because preliminary fine-tuning runs were highly variable.

  8. Knowl 8 — Layerwise fine-tuning changes upper layers more, while retaining their relative label geometry

    empirical result

    For BERTbase, the authors compared each layer during POS, dependency-relation, preposition supersense role, and preposition supersense function fine-tuning with the corresponding original pretrained layer. They measured spatial similarity as the Pearson correlation between the two layers’ vectors of pairwise label-cluster distances. In the higher layers, this correlation remained above 0.5, so the changes did not amount to an arbitrary rearrangement: the relative positions of label clusters were largely preserved while the representations were adapted to the task. TREC-50 was excluded because its representations were nonlinear and did not provide the required distance vectors.

    A complementary analysis tracked each label cluster’s centroid displacement before and after fine-tuning. In lower layers, these displacements occupied a much smaller region and tended to point in similar directions, whereas higher-layer displacements covered a wider range. In the two-dimensional PCA projections reported for POS, layer 2 spanned approximately -1 to 3 on one axis and -3 to 3 on the other; layer 12 spanned approximately -12 to 13 and -12 to 8. These projected ranges illustrate the greater movement in the upper layer, not distances in the original embedding space.

  9. Knowl 9 — Classifier probes use independently trained two-layer networks on frozen representations

    experimental setup

    To compare how well original and fine-tuned representations support task prediction, the authors froze the representation and trained a two-layer neural-network probe from scratch. They selected probe hyperparameters by grid search: each of the two hidden-layer sizes came from {32, 64, 128, 256}, and the regularizer weight was searched from 10−710^{-7} to 10010^0. The hidden layers used ReLU activations, and the network was optimized with Adam for at most 1,000 iterations. Each selected probe was trained with five random initializations; the reported accuracy is the average and the accompanying standard deviation is across those runs. This probe-training phase was separate from fine-tuning, so both original and fine-tuned representations were evaluated using frozen embeddings and newly trained probes.

  10. Knowl 10 — The evidence is limited to English BERT experiments and does not establish a theory of generalization

    limitation

    The experiments cover English tasks and the BERT family; whether the findings extend to other languages or transformer architectures remains unverified. The geometric analysis focuses on distances and arrangements between label clusters and omits the structure of points within each cluster. The authors also do not provide a theoretical account linking their geometric observations to representation learnability or classifier generalization, and they do not explain why higher layers change more than lower layers. The proposed connection between training–test geometric divergence and fine-tuning performance is a conjecture rather than an established causal result.

Coverage note — The paper’s plots of per-label distance trajectories and centroid paths are summarized through their stated trends rather than reproduced label by label; the additional layerwise PCA plots for DEP and preposition supersense tasks are omitted because they illustrate the same movement pattern described in the layerwise result.

References

  1. 1.Armen Aghajanyan, Sonal Gupta, and Luke Zettlemoyer. 2021. Intrinsic dimensionality explains the effectiveness of language model fine-tuning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7319–7328, Online. Association for Computational Linguistics.
  2. 2.Guillaume Alain and Yoshua Bengio. 2017. Understanding intermediate layers using linear classifier probes. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Workshop Track Proceedings. OpenReview.net.
  3. 3.Yonatan Belinkov. 2021. Probing classifiers: Promises, shortcomings, and alternatives. CoRR, abs/2102.12452.
  4. 4.Chih-Chung Chang and Chih-Jen Lin. 2011. Libsvm: a library for support vector machines. ACM transactions on intelligent systems and technology (TIST), 2(3):1–27.
  5. 5.Boli Chen, Yao Fu, Guangwei Xu, Pengjun Xie, Chuanqi Tan, Mosha Chen, and Liping Jing. 2021. Probing BERT in hyperbolic spaces. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  6. 6.Sanyuan Chen, Yutai Hou, Yiming Cui, Wanxiang Che, Ting Liu, and Xiangzhan Yu. 2020. Recall and learn: Fine-tuning deep pretrained language models with less forgetting. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7870–7881, Online. Association for Computational Linguistics.
  7. 7.Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018. What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2126–2136, Melbourne, Australia. Association for Computational Linguistics.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  9. 9.Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah A. Smith. 2020. Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping. CoRR, abs/2002.06305.
  10. 10.Steffen Eger, Andreas Rücklé, and Iryna Gurevych. 2019. Pitfalls in the evaluation of sentence embeddings. In Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019), pages 55–60, Florence, Italy. Association for Computational Linguistics.
  11. 11.Kawin Ethayarajh. 2019. How contextual are contextualized word representations? comparing the geometry of BERT, ELMo, and GPT-2 embeddings. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 55–65, Hong Kong, China. Association for Computational Linguistics.
  12. 12.Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2020. Investigating learning dynamics of BERT fine-tuning. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, pages 87–92, Suzhou, China. Association for Computational Linguistics.
  13. 13.Tianxing He, Jun Liu, Kyunghyun Cho, Myle Ott, Bing Liu, James Glass, and Fuchun Peng. 2021. Analyzing the forgetting problem in pretrain-finetuning of open-domain dialogue response models. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1121–1133, Online. Association for Computational Linguistics.
  14. 14.John Hewitt, Kawin Ethayarajh, Percy Liang, and Christopher D. Manning. 2021. Conditional probing: measuring usable information beyond a baseline. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 1626–1639. Association for Computational Linguistics.
  15. 15.John Hewitt and Percy Liang. 2019. Designing and interpreting probes with control tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2733–2743, Hong Kong, China. Association for Computational Linguistics.
  16. 16.Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019. What does BERT learn about the structure of language? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3651–3657, Florence, Italy. Association for Computational Linguistics.
  17. 17.Mael Jullien, Marco Valentino, and André Freitas. 2022. Do transformers encode a foundational ontology? probing abstract classes in natural language. CoRR, abs/2201.10262.
  18. 18.Nora Kassner and Hinrich Schütze. 2020. Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7811–7818, Online. Association for Computational Linguistics.
  19. 19.Najoung Kim, Roma Patel, Adam Poliak, Patrick Xia, Alex Wang, Tom McCoy, Ian Tenney, Alexis Ross, Tal Linzen, Benjamin Van Durme, Samuel R. Bowman, and Ellie Pavlick. 2019. Probing what different NLP tasks teach machines about function word comprehension. In Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (*SEM 2019), pages 235–249, Minneapolis, Minnesota. Association for Computational Linguistics.
  20. 20.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  21. 21.Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019. Revealing the dark secrets of BERT. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4365–4374, Hong Kong, China. Association for Computational Linguistics.
  22. 22.Katarzyna Krasnowska-Kieras and Alina Wróblewska. 2019. Empirical linguistic study of sentence embeddings. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5729–5739, Florence, Italy. Association for Computational Linguistics.
  23. 23.Artur Kulmizev, Vinit Ravishankar, Mostafa Abdou, and Joakim Nivre. 2020. Do neural language models show preferences for syntactic formalisms? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4077–4091, Online. Association for Computational Linguistics.
  24. 24.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. ALBERT: A lite BERT for self-supervised learning of language representations. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  25. 25.Bai Li, Zining Zhu, Guillaume Thomas, Yang Xu, and Frank Rudzicz. 2021. How is BERT surprised? layerwise detection of linguistic anomalies. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4215–4228, Online. Association for Computational Linguistics.
  26. 26.Xin Li and Dan Roth. 2002. Learning question classifiers. In COLING 2002: The 19th International Conference on Computational Linguistics.
  27. 27.Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019. Linguistic knowledge and transferability of contextual representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1073–1094, Minneapolis, Minnesota. Association for Computational Linguistics.
  28. 28.Amil Merchant, Elahe Rahimtoroghi, Ellie Pavlick, and Ian Tenney. 2020. What happens to BERT embeddings during fine-tuning? In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 33–44, Online. Association for Computational Linguistics.
  29. 29.David Mimno and Laure Thompson. 2017. The strange geometry of skip-gram with negative sampling. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2873–2878, Copenhagen, Denmark. Association for Computational Linguistics.
  30. 30.Marius Mosbach, Maksym Andriushchenko, and Dietrich Klakow. 2020a. On the stability of fine-tuning BERT: misconceptions, explanations, and strong baselines. CoRR, abs/2006.04884.
  31. 31.Marius Mosbach, Anna Khokhlova, Michael A. Hedderich, and Dietrich Klakow. 2020b. On the interplay between fine-tuning and sentence-level probing for linguistic knowledge in pre-trained transformers. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 68–82, Online. Association for Computational Linguistics.
  32. 32.Joakim Nivre, Marie-Catherine De Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajic, Christopher D Manning, Ryan McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, et al. 2016. Universal dependencies v1: A multilingual treebank collection. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 1659–1666.
  33. 33.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. Pytorch: An imperative style, high-performance deep learning library. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32, pages 8024–8035. Curran Associates, Inc.
  34. 34.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830.
  35. 35.Christian S Perone, Roberto Silveira, and Thomas S Paula. 2018. Evaluation of sentence embeddings in downstream and linguistic probing tasks. arXiv preprint arXiv:1806.06259.
  36. 36.Matthew E. Peters, Sebastian Ruder, and Noah A. Smith. 2019. To tune or not to tune? adapting pretrained representations to diverse tasks. In Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019), pages 7–14, Florence, Italy. Association for Computational Linguistics.
  37. 37.Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Xiaoyi Zhang, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, and Samuel R. Bowman. 2020. Intermediate-task transfer learning with pretrained language models: When and why does it work? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5231–5247, Online. Association for Computational Linguistics.
  38. 38.Abhilasha Ravichander, Yonatan Belinkov, and Eduard Hovy. 2021. Probing the probing paradigm: Does probing accuracy entail task relevance? In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 3363–3377, Online. Association for Computational Linguistics.
  39. 39.Nathan Schneider, Jena D. Hwang, Vivek Srikumar, Jakob Prange, Austin Blodgett, Sarah R. Moeller, Aviram Stern, Adi Bitan, and Omri Abend. 2018. Comprehensive supersense disambiguation of English prepositions and possessives. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 185–196, Melbourne, Australia. Association for Computational Linguistics.
  40. 40.Alon Talmor, Yanai Elazar, Yoav Goldberg, and Jonathan Berant. 2020. olmpics - on what language model pre-training captures. Trans. Assoc. Comput. Linguistics, 8:743–758.
  41. 41.Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R. Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das, and Ellie Pavlick. 2019. What do you learn from context? probing for sentence structure in contextualized word representations. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  42. 42.Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Well-read students learn better: On the importance of pre-training compact models. arXiv preprint arXiv:1908.08962.
  43. 43.Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh, and Matt Gardner. 2019. Do NLP models know numbers? probing numeracy in embeddings. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5307–5315, Hong Kong, China. Association for Computational Linguistics.
  44. 44.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355, Brussels, Belgium. Association for Computational Linguistics.
  45. 45.William F Whitney, Min Jae Song, David Brandfonbrener, Jaan Altosaar, and Kyunghyun Cho. 2021. Evaluating representations by the complexity of learning low-loss predictors. In Neural Compression: From Information Theory to Applications–Workshop@ ICLR 2021.
  46. 46.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  47. 47.Zhiyong Wu, Yun Chen, Ben Kao, and Qun Liu. 2020. Perturbed masking: Parameter-free probing for analyzing and interpreting BERT. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4166–4176, Online. Association for Computational Linguistics.
  48. 48.Yadollah Yaghoobzadeh, Katharina Kann, T. J. Hazen, Eneko Agirre, and Hinrich Schütze. 2019. Probing for semantic classes: Diagnosing the meaning content of word embeddings. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5740–5753, Florence, Italy. Association for Computational Linguistics.
  49. 49.Tianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q Weinberger, and Yoav Artzi. 2020. Revisiting few-sample bert fine-tuning. arXiv preprint arXiv:2006.05987.
  50. 50.Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021. Factual probing is [MASK]: Learning vs. learning to recall. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5017–5033, Online. Association for Computational Linguistics.
  51. 51.Yichu Zhou and Vivek Srikumar. 2021. DirectProbe: Studying representations without classifiers. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5070–5083, Online. Association for Computational Linguistics.

Citation

MLA
Zhou, Y., and V. Srikumar. “A Closer Look at How Fine-tuning Changes BERT”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 1046–61, https://doi.org/10.18653/v1/2022.acl-long.75.
APA
Zhou, Y., & Srikumar, V. (2022). A Closer Look at How Fine-tuning Changes BERT. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1046–1061. https://doi.org/10.18653/v1/2022.acl-long.75
Chicago
Zhou, Y., and V. Srikumar. 2022. “A Closer Look at How Fine-tuning Changes BERT”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1046–61. https://doi.org/10.18653/v1/2022.acl-long.75.
Harvard
Zhou, Y. and Srikumar, V. (2022) “A Closer Look at How Fine-tuning Changes BERT”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1046–1061. Available at: https://doi.org/10.18653/v1/2022.acl-long.75.
Vancouver
1. Zhou Y, Srikumar V (2022) A Closer Look at How Fine-tuning Changes BERT. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 1046–1061

BibTeX

@inproceedings{zhou-srikumar-2022-closer,
    title = "A Closer Look at How Fine-tuning Changes {BERT}",
    author = "Zhou, Yichu  and
      Srikumar, Vivek",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.75/",
    doi = "10.18653/v1/2022.acl-long.75",
    pages = "1046--1061"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/