Multilingual pre-trained language models (multilingual PLMs) are neural network models trained on extensive text datasets spanning multiple languages to learn shared linguistic and semantic representations across diverse language families. Built predominantly on Transformer architectures, these models utilize self-supervised learning objectives across multilingual corpora, allowing them to capture universal syntactic and semantic patterns without relying exclusively on parallel translation data. This unified cross-lingual representation facilitates cross-lingual transfer learning, enabling the models to be fine-tuned on tasks such as text classification, named entity recognition, question answering, and translation in high-resource languages and subsequently applied to low-resource languages with minimal or zero task-specific training data.