Supportive pretraining data refers to a subset of training examples within a machine learning model pretraining corpus that directly facilitates and enhances specific emergent capabilities, such as in-context learning and downstream task adaptation. Rather than being selected purely based on domain relevance or topic similarity to target tasks, supportive pretraining data is typically identified by analyzing its functional contribution to model behavior through gradient-based attribution or data influence methods. These high-utility examples often possess distinct compositional traits, such as an enriched concentration of rare or long-tail tokens and challenging contextual structures that require the model to resolve long-range dependencies during unsupervised pretraining.