Unified speech-text pre-training is a machine learning framework in which speech audio and written text are jointly trained within a shared model architecture to align their representations across modalities. By processing acoustic signals and text corpora simultaneously through combinations of self-supervised and supervised objectives, this approach bridges the structural gap between continuous audio features and discrete linguistic tokens. Integrating rich linguistic knowledge from large text corpora directly into speech representations allows models to build a unified semantic space, substantially improving generalization and performance on downstream tasks such as automatic speech recognition, speech translation, and cross-modal language understanding.