Word Order Does Matter and Shuffled Language Models Know It
Mostafa AbdouVinit RavishankarArtur KulmizevAnders Søgaard
Reveals how language models trained on scrambled text still recover natural word order through subword segmentation artifacts and statistical dependencies between sentence length and token frequencies, explaining why position embeddings remain vital even under shuffled pre-training.
Recent artificial intelligence research suggested that modern deep-learning language models do not strictly require natural word order during pre-training to perform well on standard language benchmarks. This raised widespread questions about how neural networks process human text and whether sequence structure is truly essential for automated comprehension. The article investigates why models pre-trained on shuffled text appear to perform so well, demonstrating that these models actually retain hidden word order information and that natural word order is indispensable for sophisticated language understanding.
To evaluate this question, the researchers conducted diagnostic probing experiments, attention mechanism analyses, and downstream evaluations across full-scale and controlled transformer models. They trained linear diagnostic models to predict word order and word positions using representations extracted from models pre-trained on natural text, shuffled text, and text lacking position markers. They also evaluated model performance across the SuperGLUE benchmark suite and adversarially filtered reasoning datasets, which remove spurious surface cues and require genuine syntactic and contextual understanding.
The investigation revealed several critical findings. First, models pre-trained on scrambled text retain significant knowledge of natural word order: simple linear classifiers predicted relative word order with over 71% accuracy, while models lacking position representations performed at chance levels near 50%. Second, this retention occurs because text scrambling was often performed before subword segmentation, leaving local word pieces intact, and because statistical dependencies between sentence length and vocabulary frequency leak structural information. Third, while shuffled models perform competitively on basic benchmarks, their performance drops significantly on complex, context-sensitive tasks, exhibiting an 8.1 point average decrease on challenging tasks and drops exceeding 12 points on reading comprehension and pronoun resolution. Finally, positional representations cannot be effectively introduced during downstream fine-tuning if omitted during initial pre-training.
These results indicate that sequence structure remains vital for high-level language comprehension and reasoning. Prior studies that minimized the importance of word order relied on evaluation benchmarks that allow models to succeed through shallow lexical heuristics rather than true comprehension. For organizations deploying natural language processing systems in high-stakes environments, relying on models that disregard sequence structure creates substantial operational risks in reasoning, coreference resolution, and semantic interpretation.
Organizations should maintain strict sequential integrity and appropriate subword pipelines during model pre-training. Furthermore, evaluation teams must incorporate adversarially filtered test suites that evaluate structural and syntactic reasoning rather than relying solely on standard aggregate benchmarks. While the findings provide high confidence regarding transformer behavior in English, future work should validate whether these dynamics generalize across generative architectures and morphologically diverse languages where word order plays different grammatical roles.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). Its HANS benchmark shows how standard NLI scores can hide reliance on superficial cues, motivating the source’s use of adversarial tests for genuine language understanding.
- Paper: Self-Attention with Relative Position Representations, Peter Shaw et al. (2018). This study establishes how Transformers encode relative word positions in attention, providing useful grounding for the source’s analysis of positional information.
- Paper: BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Jacob Devlin et al. (2019). BERT’s Transformer pre-training and GLUE evaluations provide essential context for understanding the source’s experiments on pre-trained models and downstream benchmarks.
- Paper: What Does BERT Learn about the Structure of Language?, Ganesh Jawahar et al. (2019). Its probing analysis of linguistic structure in BERT prepares readers for the source’s tests of what word-order information remains encoded in model representations.
- Paper: Premise Order Matters in Reasoning with Large Language Models, Xinyun Chen et al. (2024). Building on the importance of sequence structure for understanding, this study tests how rearranging logically equivalent premises disrupts large language models’ reasoning.
- Paper: When transformers learn "impossible" languages, what do they learn?, Ram Janarthan et al. (2026). This later study extends the question of how altered word order affects language models by testing grammatical sensitivity and generation across systematically modified languages.
