Time Waits for No One! Analysis and Challenges of Temporal Misalignment

Kelvin LuuDaniel KhashabiSuchin GururanganKarishma MandyamNoah A. Smith

article2022NAACL124 citations

Establishes a multi-domain benchmark and regression-based metric to quantify how temporal misalignment degrades NLP model performance over time, proving that continued language model pretraining cannot substitute for finetuning on temporally aligned labeled data.

Listen

Modern natural language processing systems are commonly deployed in dynamic environments where language use and topics evolve over time. However, when these models are trained on data from one period and evaluated or used on data from another, temporal misalignment occurs. This disconnect causes systems to degrade silently after deployment, presenting significant risks to system reliability and operational performance. Understanding the speed, severity, and mechanics of this deterioration is critical for organizations that rely on automated text analysis.

The article establishes a systematic evaluation to quantify the impact of temporal misalignment across diverse applications. It assesses how performance decays over time and demonstrates the extent to which standard mitigation strategies—specifically adapting general language models on unannotated current text versus collecting fresh task-specific labeled data—can restore model accuracy.

To measure these effects, the researchers created an evaluation suite covering eight distinct tasks across four domains: social media (Twitter), scientific articles, newsroom journalism, and online food reviews (Yelp). The datasets span time ranges from five to over thirty years. By controlling training set sizes and varying the time gap between training and testing data, the study conducted over 500 experiments using the GPT-2 architecture. It introduced a standardized Temporal Degradation metric, which calculates the average annual rate of performance loss for each application.

The analysis reveals that temporal misalignment consistently degrades system performance, often at severe rates. In rapidly evolving domains like political social media analysis, model accuracy drops by as much as 40 F1 points across a five-year gap, exhibiting an average degradation rate of 7.72 points per year. In contrast, stable domains such as restaurant review sentiment analysis show minimal degradation, losing only around 0.26 points per year. Notably, performance deterioration occurs in both directions: newer models perform poorly on historical data just as older models fail on modern text. Crucially, continuing language model pretraining on unannotated contemporary text yields negligible performance gains. Instead, task-specific finetuning on temporally aligned labeled data drives almost all downstream recovery.

These findings indicate that organizations cannot rely on lightweight, unsupervised language model updates to keep production systems accurate over time. A failure to update labeled task data introduces severe operational risks, as performance can drop faster across a few years of misalignment than the margin of difference between competing state-of-the-art architectures. Standard static evaluation benchmarks may also provide an overly optimistic assessment of real-world model capabilities if training and testing sets share the same time period.

Organizations managing text-based models should implement continuous monitoring to detect performance drift and establish recurring pipelines to collect refreshed, labeled domain data. Where budgets are constrained, decision-makers must prioritize annotating new task-specific data over running unsupervised domain adaptation on raw text. Further technical research is needed to develop cost-effective continual learning algorithms and automated drift detection tools that minimize the manual burden of ongoing annotation.

While these conclusions are supported by a rigorous experimental design across hundreds of trials, readers should note that the study fixed dataset sizes to isolate temporal effects, meaning it did not measure performance when historical and modern data are accumulated together. Additionally, sudden real-world shocks, such as elections or pandemics, may accelerate degradation beyond the linear trends observed across the benchmark datasets.

arXiv: 2111.07408
Cover for Time Waits for No One! Analysis and Challenges of Temporal Misalignment

Abstract

When an NLP model is trained on text data from one time period and tested or deployed on data from another, the resulting temporal misalignment can degrade end-task performance. In this work, we establish a suite of eight diverse tasks across different domains (social media, science papers, news, and reviews) and periods of time (spanning five years or more) to quantify the effects of temporal misalignment. Our study is focused on the ubiquitous setting where a pretrained model is optionally adapted through continued domain-specific pretraining, followed by task-specific finetuning. We establish a suite of tasks across multiple domains to study temporal misalignment in modern NLP systems. We find stronger effects of temporal misalignment on task performance than have been previously reported. We also find that, while temporal adaptation through continued pretraining can help, these gains are small compared to task-specific finetuning on data from the target time period. Our findings motivate continued research to improve temporal robustness of NLP models.

Table of Contents

  • 1 Introduction
  • 2 Methodology Overview
  • 2.1 Learning Pipeline
  • 2.2 Evaluation Methodology
  • 2.3 Quantifying Temporal Degradation
  • 2.4 Domains, Tasks, and Datasets
  • 3 Empirical Results and Analysis
  • 3.1 Temporal Misalignment in Tasks
  • 3.2 Temporal Misalignment in LMs
  • 4 Limitations and Future Work
  • 5 Conclusion
  • Acknowledgments
  • References
  • Supplementary Material
  • A A Metric for Temporal Degradation
  • B Details of Model Development
  • C Data Collection
  • D Extended Results

Knowls

  1. Knowl 1 — Temporal Degradation Metric

    equation

    The Temporal Degradation (TD) score measures the expected average rate of model performance deterioration (such as in macro F1F_1, Rouge-L, or perplexity) per time unit of mismatch between training data timestamp t′t' and evaluation data timestamp tt.

    Let St′→tS_{t' \to t} denote the evaluation score of a model trained on data from timestamp t′t' and evaluated on data from timestamp tt. The modified performance difference D(t′→t)D(t' \to t) is defined as:

    D(t′→t)=−(St′→t−St→t)×sign(t′−t)D(t' \to t) = -(S_{t' \to t} - S_{t \to t}) \times \text{sign}(t' - t)

    The modification by sign(t′−t)\text{sign}(t' - t) ensures that as performance degrades relative to the temporally aligned baseline St→tS_{t \to t}, the value of D(t′→t)D(t' \to t) increases monotonically with training time distance, regardless of whether t′t' is in the past or the future relative to tt.

    For a fixed evaluation timestamp tt across a set of training timestamps TT, the degradation rate TD(t)TD(t) is computed via least-squares linear regression as the slope:

    TD(t)=∑t′∈T(D(t′→t)−Dˉ)(t′−tˉ′)∑t′∈T(t′−tˉ′)2TD(t) = \frac{\sum_{t' \in T} (D(t' \to t) - \bar{D})(t' - \bar{t}')}{\sum_{t' \in T} (t' - \bar{t}')^2}

    where tˉ′=1∣T∣∑t′∈Tt′\bar{t}' = \frac{1}{|T|} \sum_{t' \in T} t' and Dˉ=1∣T∣∑t′∈TD(t′→t)\bar{D} = \frac{1}{|T|} \sum_{t' \in T} D(t' \to t). The overall TD score for a task across all evaluation timestamps T′T' is the arithmetic mean:

    TD=1∣T′∣∑t∈T′TD(t)TD = \frac{1}{|T'|} \sum_{t \in T'} TD(t)

  2. Knowl 2 — Inadequacy of Temporal Pretraining Compared to Temporally Aligned Finetuning

    empirical result

    Continued domain-adaptive pretraining (DAPT) of a language model using unlabeled text from a target time period does not mitigate performance drops caused by temporally misaligned task supervision. In contrast, finetuning on labeled data temporally aligned with the evaluation period produces major performance gains.

    For example, in political affiliation classification on Twitter (POLIAFF):

    • A model finetuned on 2015 labeled data achieves an F1F_1 score of 91.4%91.4\% on 2015 test data, but drops to 48.4%48.4\% on 2020 test data.
    • Adapting the underlying language model via DAPT on unlabeled 2020 text before finetuning on 2015 labeled data only shifts the 2020 test F1F_1 from 48.4%48.4\% to 50.8%50.8\%.
    • Conversely, finetuning the default unadapted language model directly on 2020 labeled data yields an F1F_1 of 78.0%78.0\% on 2020 test data.

    Across multiple domains and tasks, varying the pretraining era under a fixed finetuning condition results in largely uniform task scores, demonstrating that unsupervised temporal domain adaptation cannot serve as a substitute for temporally aligned labeled training data.

  3. Knowl 3 — Downstream Task Degradation Rates and Bidirectionality Under Temporal Misalignment

    data/table

    Temporal misalignment between training and evaluation datasets leads to substantial performance degradation across diverse NLP tasks, with degradation occurring bidirectionally (models trained on past data degrade on future data, and models trained on future data degrade on past data).

    Domain Task Metric TD Score Correlation (rr)
    Twitter POLIAFF Macro F1F_1 7.72 0.98
    Twitter TWIERC Macro F1F_1 0.96 0.74
    Science SCIERC Macro F1F_1 0.67 0.93
    Science AIC Macro F1F_1 1.79 0.93
    News PUBCLS Macro F1F_1 5.46 0.85
    News NEWSUM Rouge-L 1.38 0.91
    News MFC Macro F1F_1 0.98 0.86
    Reviews YELPCLS Macro F1F_1 0.26 0.30

    The table reports TD scores (rate of performance degradation per unit time divergence) and Pearson correlation coefficients (rr) between training-evaluation distance and performance drop. In all tasks except Yelp review rating classification, degradation exhibits statistically significant linear correlation with temporal distance (r>0.5r > 0.5, p<0.05p < 0.05). Rapidly shifting tasks like political affiliation classification (POLIAFF) and publisher classification (PUBCLS) show degradation rates exceeding 5 to 7 F1F_1 points per time step.

  4. Knowl 4 — Suite of Benchmark Datasets for Evaluating NLP Temporal Misalignment

    experimental setup

    The temporal misalignment benchmark suite spans eight tasks across four distinct text domains, with temporal spans ranging from 5 to 36 years. For each task, training and evaluation sets are split into discrete, roughly equal-sized chronological partitions to isolate temporal drift from sample size effects:

    1. Twitter - Political Affiliation Classification (POLIAFF): 120k tweets (2015–2019) from U.S. politicians classified as Democrat vs. Republican.
    2. Twitter - Named Entity Type Classification (TWIERC): 8k tweets (2014–2019) classifying named entity mentions into Person, Organization, or Location.
    3. Science - Mention Type Classification (SCIERC): 8k scientific abstract mentions (1980–2016) mapped to Task, Method, Metric, Material, Other-Scientific-Term, or Generic.
    4. Science - AI Venue Classification (AIC): 16k papers (2009–2020) classified by publication venue (AAAI vs. ICML).
    5. News - Media Frames Classification (MFC): 20k news articles (2009–2016) classified across 15 media framing dimensions.
    6. News - Publisher Classification (PUBCLS): 67k articles (2009–2016) uniformly downsampled across three prolific publishers (Fox News, New York Times, Washington Post).
    7. News - Summarization (NEWSUM): 330k news article-summary pairs (2009–2016) evaluated using Rouge-L.
    8. Food Reviews - Review Rating Classification (YELPCLS): 126k restaurant reviews (2013–2019) predicting integer star ratings (1 to 5).
  5. Knowl 5 — Language Model Perplexity Degradation Across Domains

    empirical result

    When GPT-2 is adapted via continued domain-adaptive pretraining (DAPT) on time-stamped text collections within a domain, language model perplexity deteriorates as the evaluation time period diverges from the pretraining time period.

    The rate of language model temporal degradation varies substantially across domains:

    • Twitter: Perplexity TD score of 1.491.49, exhibiting the most rapid degradation (e.g., a model adapted on 2015 tweets has perplexity 24.824.8 on 2015 text but 30.930.9 on 2020 text).
    • News (Newsroom): Perplexity TD score of 0.490.49 (e.g., a model adapted on 2009–2010 text has perplexity 22.322.3 on 2009–2010 text and 20.320.3 on 2015–2016 text, whereas adapting on 2015–2016 text lowers 2015–2016 perplexity to 17.317.3).
    • Scientific Articles (Semantic Scholar): Perplexity TD score of 0.310.31.
    • Food Reviews (Yelp): Perplexity TD score of 0.190.19, displaying minimal temporal drift over multi-year spans.
  6. Knowl 6 — Marginal Label Distribution Drift Over Time

    empirical result

    Temporal misalignment induces significant shifts in the marginal distribution of target labels P(Y)P(Y) over time, independent of model architecture. Quantifying label drift via Kullback-Leibler (KL) divergence between each test time period and the initial period reveals substantial drift in specific tasks:

    • In political affiliation classification on Twitter (POLIAFF), Republican tweets outnumbered Democratic tweets by more than 2:1 in 2015, whereas Democratic tweets outnumbered Republican tweets by 2020, resulting in a label distribution KL divergence of 0.300.30 relative to 2015.
    • Media frames classification (MFC) reaches a label distribution KL divergence of 0.210.21 between 2009–2010 and 2015–2016.
    • In contrast, tasks such as named entity typing (SCIERC) and Yelp rating classification (YELPCLS) exhibit stable marginal label distributions over time, with maximum KL divergences of 0.010.01 and 0.110.11, respectively.
  7. Knowl 7 — Correlation Between Vocabulary Overlap and Task Degradation

    empirical result

    Lexical shift, measured as the percentage overlap between the top 10,000 most frequent unigrams of two time periods within a domain, decreases as temporal distance increases. Furthermore, lexical overlap strongly correlates with downstream task performance.

    The Pearson correlation coefficient (rr) between unigram vocabulary overlap and downstream task performance across time periods is:

    • POLIAFF (Twitter): r=0.84r = 0.84
    • MFC (News): r=0.80r = 0.80
    • AIC (Science): r=0.79r = 0.79
    • SCIERC (Science): r=0.72r = 0.72
    • NEWSUM (News): r=0.72r = 0.72
    • PUBCLS (News): r=0.65r = 0.65
    • TWIERC (Twitter): r=0.51r = 0.51
    • YELPCLS (Food Reviews): r=0.14r = 0.14
  8. Knowl 8 — Training Configuration for Temporal Adaptation and Finetuning

    experimental setup

    All experiments utilize GPT-2 as the base model using HuggingFace implementations on Quadro RTX 800 GPUs. Experimental parameters are standardized across settings:

    • Temporal Domain Adaptation (DAPT): Pretraining continues for 10,000 steps per temporal partition with a batch size of 32, block size of 1024 tokens, maximum learning rate of 5×10−55 \times 10^{-5}, and Adam optimizer parameters β1=0.9\beta_1 = 0.9, β2=0.999\beta_2 = 0.999, ϵ=10−8\epsilon = 10^{-8}.
    • Classification Finetuning: Models are finetuned for 50 epochs with a batch size of 32, maximum learning rate of 2×10−52 \times 10^{-5}, Adam optimizer (β1=0.9,β2=0.999,ϵ=10−8\beta_1 = 0.9, \beta_2 = 0.999, \epsilon = 10^{-8}), and sequence length 512.
    • Summarization Finetuning: NEWSUM models are finetuned for 10 epochs with a batch size of 8, learning rate of 2×10−52 \times 10^{-5}, sequence length 512, and generated using top-pp sampling (p=0.05p = 0.05), top-k=20k = 20, and temperature 1.0.

    All reported task performance figures are averaged across five independent runs using distinct random seeds.

  9. Knowl 9 — Limitations of Uniform Slice Temporal Misalignment Evaluation

    limitation

    The experimental evaluation isolates temporal misalignment by evaluating models trained exclusively on discrete, isolated time slices of fixed size, which presents several practical limitations:

    1. In practical deployments, practitioners often accumulate historical training data cumulatively across multiple successive time periods rather than discarding older data.
    2. In natural data streams, the volume and velocity of available data fluctuate over time rather than remaining constant across partitions.
    3. Language change and world events occur non-uniformly (e.g., sudden disruptions such as elections or pandemics cause abrupt semantic and topical shifts).
    4. Temporal misalignment can also impact annotation validity: annotating historical texts using contemporary perspectives risks introducing anachronistic label noise.

Coverage note — None was omitted. All primary empirical findings, formal metrics, benchmark task specifications, experimental hyperparameters, and stated limitations are covered.

References

  1. 1.Eduardo G Altmann, Janet B Pierrehumbert, and Adilson E Motter. 2009. Beyond word frequency: Bursts, lulls, and scaling in the temporal distributions of words. PLOS one, 4(11):e7678.
  2. 2.Waleed Ammar, Dirk Groeneveld, Chandra Bhagavatula, Iz Beltagy, Miles Crawford, Doug Downey, Jason Dunkelberger, Ahmed Elgohary, Sergey Feldman, Vu Ha, et al. 2018. Construction of the literature graph in semantic scholar. In NAACL.
  3. 3.Fan Bai, Alan Ritter, and Wei Xu. 2021. Pre-train or annotate? domain adaptation with a constrained budget. In EMNLP, pages 5002–5015.
  4. 4.Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. Scibert: A pretrained language model for scientific text. In EMNLP.
  5. 5.Andrew Eliot Borthwick. 1999. A maximum entropy approach to named entity recognition. New York University.
  6. 6.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165.
  7. 7.Ivy Cao, Zizhou Liu, Giannis Karamanolakis, Daniel Hsu, and Luis Gravano. 2021. Quantifying the effects of COVID-19 on restaurant reviews. In Proceedings of the International Workshop on Natural Language Processing for Social Media, Online. Association for Computational Linguistics.
  8. 8.Dallas Card, Amber E. Boydstun, Justin H. Gross, Philip Resnik, and Noah A. Smith. 2015. The media frames corpus: Annotations of frames across issues. In ACL.
  9. 9.Alexandra Chronopoulou, Matthew E. Peters, and Jesse Dodge. 2021. Efficient hierarchical domain adaptation for pretrained language models. ArXiv, abs/2112.08786.
  10. 10.Hal Daume III. 2007. Frustratingly easy domain adaptation. In ACL, pages 256–263.
  11. 11.Kushal Dave, Steve Lawrence, and David M. Pennock. 2003. Mining the peanut gallery: opinion extraction and semantic classification of product reviews. In WWW.
  12. 12.Cyprien de Masson d'Autume, Sebastian Ruder, Lingpeng Kong, and Dani Yogatama. 2019. Episodic memory in lifelong language learning. In nips.
  13. 13.Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W. Cohen. 2021. Time-aware language models as temporal knowledge bases. CoRR, abs/2106.15110.
  14. 14.Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, and Matt Gardner. 2021. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In EMNLP.
  15. 15.Jacob Eisenstein, Brendan O’Connor, Noah A Smith, and Eric P Xing. 2014. Diffusion of lexical change in social media. PloS one, 9(11).
  16. 16.Robert M. Entman. 1983. Framing: Toward clarification of a fractured paradigm. Journal of Communications.
  17. 17.Hila Gonen, Ganesh Jawahar, Djame Seddah, and Yoav Goldberg. 2020. Simple, interpretable and stable method for detecting words with usage change across corpora. In ACL.
  18. 18.Max Grusky, Mor Naaman, and Yoav Artzi. 2018. Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies. In NAACL.
  19. 19.Suchin Gururangan, Mike Lewis, Ari Holtzman, Noah A. Smith, and Luke Zettlemoyer. 2021. Demix layers: Disentangling domains for modular language modeling. CoRR, abs/2108.05036.
  20. 20.Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In ACL.
  21. 21.William L. Hamilton, Jure Leskovec, and Dan Jurafsky. 2016. Diachronic word embeddings reveal statistical laws of semantic change. In ACL.
  22. 22.Xiaochuang Han and Jacob Eisenstein. 2019. Unsupervised domain adaptation of contextualized embeddings for sequence labeling. In EMNLP.
  23. 23.Yufan Huang, Yanzhe Zhang, Jiaao Chen, Xuezhi Wang, and Diyi Yang. 2021. Continual learning for text classification with information disentanglement based regularization. In ACL.
  24. 24.Mohit Iyyer, Peter Enns, Jordan Boyd-Graber, and Philip Resnik. 2014. Political ideology detection using recursive neural networks. In ACL.
  25. 25.Kokil Jaidka, Niyati Chhaya, and Lyle Ungar. 2018. Diachronic degradation of language models: Insights from social media. In ACL.
  26. 26.Jing Jiang and ChengXiang Zhai. 2007. Instance weighting for domain adaptation in nlp. In ACL.
  27. 27.Xisen Jin, Dejiao Zhang, Henghui Zhu, Wei Xiao, Shang-Wen Li, Xiaokai Wei, Andrew Arnold, and Xiang Ren. 2021. Lifelong pretraining: Continually adapting language models to emerging corpora. arXiv preprint arXiv:2110.08534.
  28. 28.William Labov. 2011. Principles of linguistic change, volume 3: Cognitive and cultural factors, volume 36. John Wiley & Sons.
  29. 29.Angeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d’Autume, Sebastian Ruder, Dani Yogatama, et al. 2021. Pitfalls of static language modelling. arXiv preprint arXiv:2102.01951.
  30. 30.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Proc. of Text Summarization Branches Out.
  31. 31.Wei-Hao Lin, Eric P. Xing, and Alexander Hauptmann. 2008. A joint topic and perspective model for ideological discourse. In ECML/PKDD.
  32. 32.Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel S Weld. 2020. S2orc: The semantic scholar open research corpus. In ACL.
  33. 33.Jie Lu, Anjin Liu, Fan Dong, Feng Gu, Joao Gama, and Guangquan Zhang. 2019. Learning under concept drift: A review. IEEE Transactions on Knowledge and Data Engineering.
  34. 34.Yi Luan, Luheng He, Mari Ostendorf, and Hannaneh Hajishirzi. 2018. Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction. In EMNLP.
  35. 35.Afonso Mendes, Shashi Narayan, Sebastiao Miranda, Zita Marinho, Andre F. T. Martins, and Shay B. Cohen. 2019. Jointly extracting and compressing documents with summary state representations. In NAACL.
  36. 36.Bo Pang, Lillian Lee, and Shivakumar Vaithyanathan. 2002. Thumbs up? sentiment classification using machine learning techniques. In EMNLP.
  37. 37.Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In NAACL.
  38. 38.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 21:1–67.
  39. 39.Shruti Rijhwani and Daniel Preotiuc-Pietro. 2020. Temporally-informed analysis of named entity recognition. In ACL.
  40. 40.Paul Rottger and Janet Pierrehumbert. 2021. Temporal adaptation of bert and performance on downstream document classification: Insights from social media. In Findings of EMNLP, pages 2400–2412.
  41. 41.Maja R. Rudolph and David M. Blei. 2018. Dynamic embeddings for language evolution. WWW.
  42. 42.Tian Shi, Yaser Keneshloo, Naren Ramakrishnan, and Chandan K. Reddy. 2021. Neural abstractive text summarization with sequence-to-sequence models. ACM Transactions on Data Science.
  43. 43.Tian Shi, Ping Wang, and Chandan K. Reddy. 2019. LeafNATS: An open-source toolkit and live demo system for neural abstractive text summarization. In NAACL.
  44. 44.Hidetoshi Shimodaira. 2000. Improving predictive inference under covariate shift by weighting the loglikelihood function. Journal of Statistical Planning and Inference.
  45. 45.Ian Stewart and Jacob Eisenstein. 2018. Making “fetch” happen: The influence of social and linguistic context on nonstandard word growth and decline. In EMNLP.
  46. 46.Shivashankar Subramanian, Daniel King, Doug Downey, and Sergey Feldman. 2021. S2and: A benchmark and evaluation system for author name disambiguation. In ACM/IEEE Joint Conference on Digital Libraries (JCDL), pages 170–179. IEEE.
  47. 47.Nadine Tamburrini, Marco Cinnirella, Vincent AA Jansen, and John Bryden. 2015. Twitter users change word usage according to conversationpartner social identity. Social Networks, 40:84–89.
  48. 48.David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or fiction: Verifying scientific claims. In EMNLP.
  49. 49.Yunli Wang and Cyril Goutte. 2017. Detecting changes in twitter streams using temporal clusters of hashtags. In Proceedings of the Events and Stories in the News Workshop.
  50. 50.Geoffrey I. Webb, Loong Kuan Lee, Bart Goethals, and Francois Petitjean. 2018. Analyzing concept drift and shift from sample data. Data Mining and Knowledge Discovery.
  51. 51.Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, and Dan Klein. 2021. Detoxifying language models risks marginalizing minority voices. In NAACL.
  52. 52.Yi Yang and Jacob Eisenstein. 2016. Part-of-speech tagging for historical english. In NAACL.
  53. 53.Justine Zhang, Jonathan Chang, Cristian Danescu-Niculescu-Mizil, Lucas Dixon, Yiqing Hua, Dario Taraborelli, and Nithum Thain. 2018. Conversations gone awry: Detecting early signs of conversational failure. In ACL.
  54. 54.Kun Zhang, Bernhard Scholkopf, Krikamol Muandet, and Zhikun Wang. 2013. Domain adaptation under target and conditional shift. In ICML.
  55. 55.Michael J.Q. Zhang and Eunsol Choi. 2021. SituatedQA: Incorporating extra-linguistic contexts into QA. EMNLP.
  56. 56.Yi Zhang, Zachary Ives, and Dan Roth. 2020. “who said it, and why?” provenance for natural language claims. In ACL.
  57. 57.Xuhui Zhou, Maarten Sap, Swabha Swayamdipta, Yejin Choi, and Noah A. Smith. 2021. Challenges in automated debiasing for toxic language detection. In EACL.

Citation

MLA
Luu, K., et al. “Time Waits for No One! Analysis and Challenges of Temporal Misalignment”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 5944–58, https://doi.org/10.18653/v1/2022.naacl-main.435.
APA
Luu, K., Khashabi, D., Gururangan, S., Mandyam, K., & Smith, N. A. (2022). Time Waits for No One! Analysis and Challenges of Temporal Misalignment. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 5944–5958. https://doi.org/10.18653/v1/2022.naacl-main.435
Chicago
Luu, K., D. Khashabi, S. Gururangan, K. Mandyam, and N. A. Smith. 2022. “Time Waits for No One! Analysis and Challenges of Temporal Misalignment”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 5944–58. https://doi.org/10.18653/v1/2022.naacl-main.435.
Harvard
Luu, K. et al. (2022) “Time Waits for No One! Analysis and Challenges of Temporal Misalignment”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 5944–5958. Available at: https://doi.org/10.18653/v1/2022.naacl-main.435.
Vancouver
1. Luu K, Khashabi D, Gururangan S, Mandyam K, Smith NA (2022) Time Waits for No One! Analysis and Challenges of Temporal Misalignment. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 5944–5958

BibTeX

@inproceedings{luu-etal-2022-time,
    title = "Time Waits for No One! Analysis and Challenges of Temporal Misalignment",
    author = "Luu, Kelvin  and
      Khashabi, Daniel  and
      Gururangan, Suchin  and
      Mandyam, Karishma  and
      Smith, Noah A.",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.435/",
    doi = "10.18653/v1/2022.naacl-main.435",
    pages = "5944--5958"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/