Word Order Does Matter and Shuffled Language Models Know It

Mostafa AbdouVinit RavishankarArtur KulmizevAnders Søgaard

article2022ACL58 citationsBest Paper Award

Reveals how language models trained on scrambled text still recover natural word order through subword segmentation artifacts and statistical dependencies between sentence length and token frequencies, explaining why position embeddings remain vital even under shuffled pre-training.

Listen

Recent artificial intelligence research suggested that modern deep-learning language models do not strictly require natural word order during pre-training to perform well on standard language benchmarks. This raised widespread questions about how neural networks process human text and whether sequence structure is truly essential for automated comprehension. The article investigates why models pre-trained on shuffled text appear to perform so well, demonstrating that these models actually retain hidden word order information and that natural word order is indispensable for sophisticated language understanding.

To evaluate this question, the researchers conducted diagnostic probing experiments, attention mechanism analyses, and downstream evaluations across full-scale and controlled transformer models. They trained linear diagnostic models to predict word order and word positions using representations extracted from models pre-trained on natural text, shuffled text, and text lacking position markers. They also evaluated model performance across the SuperGLUE benchmark suite and adversarially filtered reasoning datasets, which remove spurious surface cues and require genuine syntactic and contextual understanding.

The investigation revealed several critical findings. First, models pre-trained on scrambled text retain significant knowledge of natural word order: simple linear classifiers predicted relative word order with over 71% accuracy, while models lacking position representations performed at chance levels near 50%. Second, this retention occurs because text scrambling was often performed before subword segmentation, leaving local word pieces intact, and because statistical dependencies between sentence length and vocabulary frequency leak structural information. Third, while shuffled models perform competitively on basic benchmarks, their performance drops significantly on complex, context-sensitive tasks, exhibiting an 8.1 point average decrease on challenging tasks and drops exceeding 12 points on reading comprehension and pronoun resolution. Finally, positional representations cannot be effectively introduced during downstream fine-tuning if omitted during initial pre-training.

These results indicate that sequence structure remains vital for high-level language comprehension and reasoning. Prior studies that minimized the importance of word order relied on evaluation benchmarks that allow models to succeed through shallow lexical heuristics rather than true comprehension. For organizations deploying natural language processing systems in high-stakes environments, relying on models that disregard sequence structure creates substantial operational risks in reasoning, coreference resolution, and semantic interpretation.

Organizations should maintain strict sequential integrity and appropriate subword pipelines during model pre-training. Furthermore, evaluation teams must incorporate adversarially filtered test suites that evaluate structural and syntactic reasoning rather than relying solely on standard aggregate benchmarks. While the findings provide high confidence regarding transformer behavior in English, future work should validate whether these dynamics generalize across generative architectures and morphologically diverse languages where word order plays different grammatical roles.

arXiv: 2203.10995
Cover for Word Order Does Matter and Shuffled Language Models Know It

Abstract

Recent studies have shown that language models pretrained and/or fine-tuned on randomly permuted sentences exhibit competitive performance on GLUE, putting into question the importance of word order information. Somewhat counter-intuitively, some of these studies also report that position embeddings appear to be crucial for models’ good performance with shuffled text. We probe these language models for word order information and investigate what position embeddings learned from shuffled text encode, showing that these models retain information pertaining to the original, naturalistic word order. We show this is in part due to a subtlety in how shuffling is implemented in previous work – before rather than after subword segmentation. Surprisingly, we find even Language models trained on text shuffled after subword segmentation retain some semblance of information about word order because of the statistical dependencies between sentence length and unigram probabilities. Finally, we show that beyond GLUE, a variety of language understanding tasks do require word order information, often to an extent that cannot be learned through fine-tuning.

Table of Contents

  • 1 Introduction
  • 2 Models
  • 3 Probing for word order
  • 4 Hidden word-order signals
  • 5 Attention analysis
  • 6 Evaluation beyond GLUE
  • 7 Other Findings
  • 8 On Word Order
  • 9 Conclusion
  • Acknowledgements
  • References
  • A Subword vs. word scrambling
  • B On biased sampling
  • C Full UD results

Knowls

  1. Knowl 1 — Probing Shuffled Pretrained Language Models for Naturalistic Word Order

    empirical result

    Probing frozen token representations from RoBERTa language models trained on natural text (ORIG), unigram-scrambled text (SHUF.N1), and natural text without position embeddings (NOPOS) reveals that models trained on shuffled text retain substantial linear word-order information.

    Evaluation is performed on Universal Dependencies English-GUM sentences (extlength≤30 ext{length} \le 30 tokens) using frozen representations extracted from the final Transformer encoder layer (mean-pooled over subwords) without model fine-tuning:

    1. Pairwise Precedence Classification: A logistic regression classifier takes concatenated representations m(x)⊕m(y)m(x) \oplus m(y) for word pair (x,y)(x, y) and predicts whether xx precedes yy in the original naturalistic sentence.
    2. Word Position Regression: A ridge regression model predicts the continuous absolute position p(x)p(x) of a word from its representation m(x)m(x), evaluated via 6-fold cross-validation with zero vocabulary overlap between train and held-out folds.
    Model Pairwise Accuracy (%) Regression
    2k train 5k train 10k train R2R^2
    ORIG 81.50 81.74 80.40 0.68
    SHUF.N1 65.96 64.98 71.82 0.60
    NOPOS 50.41 53.35 50.22 0.03

    SHUF.N1 achieves an R2R^2 of 0.60 (close to ORIG's 0.68) and pairwise classification accuracy up to 71.82%, whereas NOPOS achieves near-random performance (R2=0.03R^2 = 0.03, accuracy ≈50.2%–53.4%\approx 50.2\%\text{--}53.4\%). This demonstrates that position embeddings enable linear probes to recover word order from shuffled representations, and that pre-training on shuffled text does not eliminate word-order encoding.

  2. Knowl 2 — SuperGLUE and WinoGrande Evaluation of Shuffled Language Models

    data/table

    Evaluating full-scale RoBERTa language models pre-trained on various levels of scrambled text across SuperGLUE and WinoGrande shows that word-order information is essential for context-sensitive and adversarially filtered tasks, contrasting earlier findings on GLUE.

    The evaluated models are ORIG (unperturbed pre-training), SHUF.N1 through SHUF.N4 (nn-gram scrambling of size 1 to 4), NOPOS (pre-trained without position embeddings), and SHUF.CORPUS (pre-trained on sentences drawn solely from corpus-level unigram frequencies):

    Model BoolQ CB COPA MultiRC ReCoRD WiC WSC WinoGrande
    Acc F1 / Acc Acc F1a / EM F1 / Acc Acc Acc Acc
    ORIG 77.6 88.2 / 87.4 61.6 67.8 / 21.9 73.5 / 72.8 67.4 73.5 62.9
    SHUF.N1 72.4 79.7 / 82.5 59.7 66.2 / 15.0 61.1 / 60.4 63.0 62.9 55.7
    SHUF.N2 73.1 86.6 / 85.5 60.3 64.8 / 16.1 63.1 / 62.4 63.0 65.3 57.6
    SHUF.N4 73.5 87.9 / 87.1 60.8 66.2 / 18.2 64.6 / 63.9 62.4 65.3 59.53
    NOPOS 66.0 63.5 / 75.0 55.6 52.8 / 3.8 23.8 / 23.5 55.4 63.09 52.73
    SHUF.CORPUS 66.7 65.6 / 73.8 56.1 52.6 / 6.4 31.0 / 30.3 57.3 65.14 51.68

    For tasks solvable by surface heuristics (COPA, MultiRC), the performance difference between ORIG and SHUF.N1 is minor (mean gap of 1.75 points). However, for tasks requiring fine-grained syntactic roles and context sensitivity (BoolQ, CB, WiC, WSC), and especially those with adversarial filtering to eliminate statistical shortcuts (ReCoRD, WinoGrande), SHUF.N1 drops by an average of 8.1 points relative to ORIG, and NOPOS drops by 19.78 points, proving that downstream fine-tuning cannot recover word-order dependencies when surface heuristics are removed.

  3. Knowl 3 — Impact of Shuffling Granularity Relative to Subword Tokenization

    empirical result

    The order of operations between token scrambling and Byte-Pair Encoding (BPE) subword segmentation dictates how much local word-order information survives in the training data:

    1. Word-Level Shuffling (Shuffling Before BPE): Scrambling words before BPE tokenization keeps multi-subword tokens composing an individual word contiguous and correctly ordered. This exposes the model to authentic subword nn-grams during masked language modeling, allowing position embeddings to learn local offset attention and pairwise correlations (manifesting as translation-invariant bands in position embedding similarity matrices).
    2. Subword-Level Shuffling (Shuffling After BPE): Scrambling tokens after BPE tokenization permutes individual subwords independently, breaking intra-word subword sequences.

    Across 50,000 sentences, word-level shuffling produces a substantially larger cumulative subword bigram overlap with the unperturbed corpus compared to subword-level shuffling. In addition, short sentences exhibit a high probability of accidental bigram preservation under word-level shuffling.

  4. Knowl 4 — Sentence Length and Vocabulary Distribution as a Positional Signal

    empirical result

    Even when text is scrambled at the subword level (eliminating all natural nn-gram sequences), position embeddings can still acquire non-random structure due to the statistical dependence between vocabulary distributions and sentence lengths.

    In natural text, unigram probabilities vary with sentence length because distinct genres, registers, and formulaic expressions have characteristic sequence length distributions. When language models are trained on synthetic corpora generated from unigram distributions:

    1. Uniform Corpus Unigram Sampling: Tokens are sampled randomly from a single global unigram distribution matching sentence lengths of the BookCorpus. The resulting position embedding similarity matrix shows zero cross-position correlation.
    2. Length-Conditioned Disjoint Vocabulary Sampling: The vocabulary is split into subsets associated with short (<l< l) versus long (≥l\ge l) sentences. The resulting learned position embeddings develop pronounced, non-random cross-position Pearson correlation structures and sentence-boundary awareness.

    This confirms that the correlation between token distributions and sentence boundaries acts as a structural cue for position embeddings in the absence of explicit word order.

  5. Knowl 5 — Dependency Length Probing and Local Adjacency Bias in Scrambled Models

    empirical result

    Probing representations for dependency syntax across varying dependency lengths ∣i−j∣|i - j| reveals that shuffled language models possess a strong local adjacency bias.

    A bilinear probe trained on frozen representations from RoBERTa variants (ORIG, SHUF.N1, SHUF.N2, SHUF.N4, NOPOS, SHUF.CORPUS) using the English Web Treebank (UD-EWT) yields the following Unlabeled Attachment Score (UAS) differences relative to ORIG (Delta\\Delta UAS):

    1. For NOPOS, probing degradation increases roughly linearly with dependency length ∣i−j∣|i - j|, reflecting the general difficulty and lower frequency of long-range arcs.
    2. For SHUF.N1, the degradation Delta\\Delta UAS relative to ORIG is minimal for immediate adjacent dependencies (∣i−j∣=1|i - j| = 1), but widens substantially for dependencies of distance ∣i−j∣≥2|i - j| \ge 2.

    This confirms that models pre-trained on unigram-scrambled data retain local word-order information sufficient for resolving adjacent syntactic dependencies, but fail on non-local, long-range dependencies.

  6. Knowl 6 — Self-Attention Offset Distributions in Scrambled Language Models

    empirical result

    Analyzing the attention offset distribution in Transformer layers reveals that models pre-trained on scrambled text develop localized attention heads, unlike models trained without position embeddings.

    For each attention head and layer l∈{1,2,7,8,11,12}l \in \{1, 2, 7, 8, 11, 12\}, the offset δ=∣i−j∣\delta = |i - j| between a token at position ii and the token jj receiving maximum attention weight is measured across 100 BookCorpus sentences:

    1. SHUF.N1: Shows a sharp peak in attention frequency at offset δ=1\delta = 1 (adjacent tokens) starting at layer 0, mimicking a local convolutional window.
    2. NOPOS: Displays an essentially flat, uniform distribution of maximum attention offsets across all distances, indicating no preferential local routing.
    3. Shuffling Granularity: Models trained on data shuffled after subword segmentation exhibit flatter early-layer offset distributions than models shuffled before segmentation.

    This shows that position embeddings actively mediate sequence-adjacent attention routing even when trained exclusively on permuted inputs.

  7. Knowl 7 — Limitations of Positional Embedding Modifications at Fine-Tuning Time

    empirical result

    Altering or adding position embeddings after pre-training during the fine-tuning phase cannot restore or maintain model capabilities on GLUE benchmark tasks (MNLI, QNLI, RTE, SST-2, CoLA):

    1. Adding Learned Position Embeddings to NOPOS: Randomly initializing learnable position embeddings and adding them to a pre-trained NOPOS model prior to fine-tuning yields no performance improvement over vanilla NOPOS, except for a minor increase on MNLI. Injecting Gaussian noise without position embeddings produces the exact same score increase on MNLI, confirming that the improvement is a generic regularization effect rather than learned position encoding.
    2. Adding Sinusoidal Embeddings to NOPOS: Adding fixed sinusoidal embeddings to NOPOS prior to fine-tuning also produces no performance gain, indicating that Transformer encoder weights must be pre-trained jointly with positional representations to utilize them.
    3. Replacing Learned Embeddings in ORIG: Replacing learned position embeddings in an ORIG model with fixed sinusoidal embeddings prior to fine-tuning substantially hurts GLUE performance, indicating that Transformer encoder weights are co-adapted to the specific geometry of their pre-trained position embeddings and cannot adapt with fine-tuning data alone.
  8. Knowl 8 — Biased Unigram Sampling by Sentence Length

    algorithm

    A synthetic corpus generation algorithm designed to isolate the effect of sentence length and vocabulary correlation on position embedding learning in the absence of sequential word-order information:

    Input: Vocabulary VV of size ∣V∣=5000|V| = 5000, sentence collection S={s1,…,sM}S = \{s_1, \dots, s_M\}, split fidelity parameter p=0.80p = 0.80
    Output: Synthetic sentence dataset S′S'
    Partition VV into two halves V1,V2V_1, V_2 of size 2500 each, such that sum of unigram frequencies in V1V_1 equals that in V2V_2
    Determine median length threshold ll such that total token count in sentences with length <l< l equals token count in sentences with length ≥l\ge l
    Initialize S′←[]S' \leftarrow []
    for each sentence s∈Ss \in S do
        k←length(s)k \leftarrow \text{length}(s)
        s′←[]s' \leftarrow []
        for t=1t = 1 to kk do
            Sample r∼Uniform(0,1)r \sim \text{Uniform}(0, 1)
            if k<lk < l then
                if r<pr < p then
                    Sample token w∼Unigram(V1)w \sim \text{Unigram}(V_1)
                else
                    Sample token w∼Unigram(V2)w \sim \text{Unigram}(V_2)
            else
                if r<pr < p then
                    Sample token w∼Unigram(V2)w \sim \text{Unigram}(V_2)
                else
                    Sample token w∼Unigram(V1)w \sim \text{Unigram}(V_1)
            Append ww to s′s'
        Append s′s' to S′S'
    return S′S'

Coverage note — No substantial contributed material was omitted. All probing results, architectural analyses of subword versus word shuffling, sentence length correlations, SuperGLUE/WinoGrande downstream evaluations, dependency arc length probing, and post-pretraining embedding manipulation experiments are covered.

References

  1. 1.Matteo Alleman, Jonathan Mamou, Miguel A Del Rio, Hanlin Tang, Yoon Kim, and SueYeon Chung. 2021. Syntactic perturbations reveal representational correlates of hierarchical phrase structure in pretrained language models. arXiv preprint arXiv:2104.07578.
  2. 2.Jörg Bahlmann, Antoni Rodriguez-Fornells, Michael Rotte, and Thomas F Münte. 2007. An fmri study of canonical and noncanonical word order in german. Human brain mapping, 28(10):940–949.
  3. 3.Thomas G Bever. 1970. The cognitive basis for linguistic structures. Cognition and the development of language.
  4. 4.Chris M. Bishop. 1995. Training with noise is equivalent to tikhonov regularization. Neural Computation, 7(1):108–116.
  5. 5.Alexander Camuto, Matthew Willetts, Umut ¸Sim¸sekli, Stephen Roberts, and Chris Holmes. 2021. Explicit Regularisation in Gaussian Noise Injections. arXiv:2007.07368 [cs, stat].
  6. 6.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044.
  7. 7.Louis Clouatre, Prasanna Parthasarathi, Amal Zouaq, and Sarath Chandar. 2021. Demystifying neural language models’ insensitivity to word-order. arXiv preprint arXiv:2107.13955.
  8. 8.Bernard Comrie. 1989. Language universals and linguistic typology: Syntax and morphology. University of Chicago press.
  9. 9.Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. 2020. On the Relationship between Self-Attention and Convolutional Layers. arXiv:1911.03584 [cs, stat].
  10. 10.Joseph H Danks and Sam Glucksberg. 1971. Psychological scaling of adjective orders. Journal of Memory and Language, 10(1):63.
  11. 11.Marie-Catherine De Marneffe, Mandy Simons, and Judith Tonhauser. 2019. The commitmentbank: Investigating projection in naturally occurring discourse. In proceedings of Sinn und Bedeutung, volume 23, pages 107–124.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  13. 13.Nai Ding, Lucia Melloni, Hang Zhang, Xing Tian, and David Poeppel. 2016. Cortical tracking of hierarchical linguistic structures in connected speech. Nature neuroscience, 19(1):158–164.
  14. 14.Timothy Dozat, Peng Qi, and Christopher D. Manning. 2017. Stanford’s graph-based neural dependency parser at the CoNLL 2017 shared task. In Proceedings of the CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies, pages 20–30, Vancouver, Canada. Association for Computational Linguistics.
  15. 15.Evelina Fedorenko, Terri L Scott, Peter Brunner, William G Coon, Brianna Pritchett, Gerwin Schalk, and Nancy Kanwisher. 2016. Neural correlate of the construction of sentence meaning. Proceedings of the National Academy of Sciences, 113(41):E6256–E6262.
  16. 16.Fernanda Ferreira, Karl GD Bailey, and Vittoria Ferraro. 2002. Good-enough representations in language comprehension. Current directions in psychological science, 11(1):11–15.
  17. 17.Angela D Friederici, Axel Mecklinger, Kevin M Spencer, Karsten Steinhauer, and Emanuel Donchin. 2001. Syntactic parsing preferences and their online revisions: A spatio-temporal analysis of event-related brain potentials. Cognitive Brain Research, 11(2):305–323.
  18. 18.Angela D Friederici, Martin Meyer, and D Yves Von Cramon. 2000. Auditory language comprehension: an event-related fmri study on the processing of syntactic and lexical information. Brain and language, 74(2):289–300.
  19. 19.Edward Gibson, Leon Bergen, and Steven T Piantadosi. 2013. Rational integration of noisy evidence and prior semantic expectations in sentence interpretation. Proceedings of the National Academy of Sciences, 110(20):8051–8056.
  20. 20.Ashim Gupta, Giorgi Kvernadze, and Vivek Srikumar. 2021. Bert & family eat word salad: Experiments with text understanding. arXiv preprint arXiv:2101.03453.
  21. 21.Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018. Annotation artifacts in natural language inference data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 107–112, New Orleans, Louisiana. Association for Computational Linguistics.
  22. 22.John Hewitt and Percy Liang. 2019. Designing and interpreting probes with control tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2733–2743, Hong Kong, China. Association for Computational Linguistics.
  23. 23.Huiyuan Jin and Haitao Liu. 2017. How will text size influence the length of its linguistic constituents? Poznan Studies in Contemporary Linguistics, 53.
  24. 24.Marcel A Just and Patricia A Carpenter. 1980. A theory of reading: From eye fixations to comprehension. Psychological review, 87(4):329.
  25. 25.Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018. Looking beyond the surface:a challenge set for reading comprehension over multiple sentences. In NAACL.
  26. 26.Artur Kulmizev and Joakim Nivre. 2021. Schrödinger’s Tree – On Syntax and Neural Language Models. arXiv preprint arXiv:2110.08887.
  27. 27.Yulia Lerner, Christopher J Honey, Lauren J Silbert, and Uri Hasson. 2011. Topographic mapping of a hierarchy of temporal receptive windows using a narrated story. Journal of Neuroscience, 31(8):2906–2915.
  28. 28.Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012. The winograd schema challenge. In Thirteenth International Conference on the Principles of Knowledge Representation and Reasoning.
  29. 29.Yongjie Lin, Yi Chern Tan, and Robert Frank. 2019. Open sesame: Getting inside bert’s linguistic knowledge. arXiv preprint arXiv:1906.01698.
  30. 30.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692 [cs].
  31. 31.Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3428–3448, Florence, Italy. Association for Computational Linguistics.
  32. 32.Francis Mollica, Matthew Siegelman, Evgeniia Diachek, Steven T Piantadosi, Zachary Mineroff, Richard Futrell, Hope Kean, Peng Qian, and Evelina Fedorenko. 2020. Composition is the core driver of the language-selective network. Neurobiology of Language, 1(1):104–134.
  33. 33.Joe O’Connor and Jacob Andreas. 2021. What context features can transformer language models use? arXiv preprint arXiv:2106.08367.
  34. 34.Christophe Pallier, Anne-Dominique Devauchelle, and Stanislas Dehaene. 2011. Cortical representation of the constituent structure of sentences. Proceedings of the National Academy of Sciences, 108(6):2522–2527.
  35. 35.Isabel Papadimitriou, Ethan A Chi, Richard Futrell, and Kyle Mahowald. 2021. Deep subjecthood: Higher-order grammatical features in multilingual bert. arXiv preprint arXiv:2101.11043.
  36. 36.Ankur P Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. 2016. A decomposable attention model for natural language inference. arXiv preprint arXiv:1606.01933.
  37. 37.Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2227–2237, New Orleans, Louisiana. Association for Computational Linguistics.
  38. 38.Thang M Pham, Trung Bui, Long Mai, and Anh Nguyen. 2020. Out of order: How important is the sequential order of words in a sentence in natural language understanding tasks? arXiv preprint arXiv:2012.15180.
  39. 39.Mohammad Taher Pilehvar and Jose Camacho-Collados. 2018. Wic: the word-in-context dataset for evaluating context-sensitive meaning representations. arXiv preprint arXiv:1808.09121.
  40. 40.Tiago Pimentel, Naomi Saphra, Adina Williams, and Ryan Cotterell. 2020. Pareto probing: Trading off accuracy for complexity. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3138–3153, Online. Association for Computational Linguistics.
  41. 41.Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018. Hypothesis only baselines in natural language inference. In Proceedings of the Seventh Joint Conference on Lexical and Computational Semantics, pages 180–191, New Orleans, Louisiana. Association for Computational Linguistics.
  42. 42.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  43. 43.Vinit Ravishankar, Artur Kulmizev, Mostafa Abdou, Anders Søgaard, and Joakim Nivre. 2021. Attention can reflect syntactic structure (if you let it). In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 3031–3045, Online. Association for Computational Linguistics.
  44. 44.Vinit Ravishankar and Anders Søgaard. 2021. The impact of positional encodings on multilingual compression. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 763–777, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  45. 45.Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. 2011. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In 2011 AAAI Spring Symposium Series.
  46. 46.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020. Winogrande: An adversarial winograd schema challenge at scale. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8732–8740.
  47. 47.Bengt Sigurd, Mats Eeg-Olofsson, and Joost Van Weijer. 2004. Word length, sentence length and frequency – zipf revisited. Studia Linguistica, 58(1):37–52.
  48. 48.Natalia Silveira, Timothy Dozat, Marie-Catherine De Marneffe, Samuel R Bowman, Miriam Connor, John Bauer, and Christopher D Manning. 2014. A gold standard dependency corpus for english. In LREC, pages 2897–2904. Citeseer.
  49. 49.Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela. 2021. Masked language modeling and the distributional hypothesis: Order word matters pre-training for little. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2888–2913, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  50. 50.Koustuv Sinha, Prasanna Parthasarathi, Joelle Pineau, and Adina Williams. 2020. Unnatural language inference. arXiv preprint arXiv:2101.00010.
  51. 51.Matthew J Traxler. 2014. Trends in syntactic parsing: Anticipation, bayesian estimation, and good-enough parsing. Trends in cognitive sciences, 18(11):605–611.
  52. 52.Paul Trichelair, Ali Emami, Jackie Chi Kit Cheung, Adam Trischler, Kaheer Suleman, and Fernando Diaz. 2018. On the evaluation of common-sense reasoning in natural language understanding. arXiv preprint arXiv:1811.01778.
  53. 53.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008.
  54. 54.Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5797–5808, Florence, Italy. Association for Computational Linguistics.
  55. 55.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019. SuperGLUE: A stickier benchmark for general-purpose language understanding systems. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  56. 56.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355, Brussels, Belgium. Association for Computational Linguistics.
  57. 57.Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang, Hao Yang, Qun Liu, and Jakob Grue Simonsen. 2021. ON POSITION EMBEDDINGS IN BERT. page 21.
  58. 58.Yu-An Wang and Yun-Nung Chen. 2020. What do position embeddings learn? an empirical study of pre-trained language model positional encoding. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6840–6849, Online. Association for Computational Linguistics.
  59. 59.Yilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart, and Stefano Ermon. 2020. A theory of usable information under computational constraints. arXiv preprint arXiv:2002.10689.
  60. 60.Amir Zeldes. 2017. The gum corpus: Creating multilayer resources in the classroom. Language Resources and Evaluation, 51(3):581–612.
  61. 61.Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme. 2018. Record: Bridging the gap between human and machine commonsense reading comprehension. arXiv preprint arXiv:1810.12885.
  62. 62.Yuan Zhang, Jason Baldridge, and Luheng He. 2019. PAWS: Paraphrase adversaries from word scrambling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1298–1308, Minneapolis, Minnesota. Association for Computational Linguistics.
  63. 63.Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In The IEEE International Conference on Computer Vision (ICCV).

Citation

MLA
Abdou, M., et al. “Word Order Does Matter and Shuffled Language Models Know It”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 6907–19, https://doi.org/10.18653/v1/2022.acl-long.476.
APA
Abdou, M., Ravishankar, V., Kulmizev, A., & Søgaard, A. (2022). Word Order Does Matter and Shuffled Language Models Know It. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6907–6919. https://doi.org/10.18653/v1/2022.acl-long.476
Chicago
Abdou, M., V. Ravishankar, A. Kulmizev, and A. Søgaard. 2022. “Word Order Does Matter and Shuffled Language Models Know It”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6907–19. https://doi.org/10.18653/v1/2022.acl-long.476.
Harvard
Abdou, M. et al. (2022) “Word Order Does Matter and Shuffled Language Models Know It”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6907–6919. Available at: https://doi.org/10.18653/v1/2022.acl-long.476.
Vancouver
1. Abdou M, Ravishankar V, Kulmizev A, Søgaard A (2022) Word Order Does Matter and Shuffled Language Models Know It. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6907–6919

BibTeX

@inproceedings{abdou-etal-2022-word,
    title = "Word Order Does Matter and Shuffled Language Models Know It",
    author = "Abdou, Mostafa  and
      Ravishankar, Vinit  and
      Kulmizev, Artur  and
      S{\o}gaard, Anders",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.476/",
    doi = "10.18653/v1/2022.acl-long.476",
    pages = "6907--6919"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/