A student transformer is a transformer-based neural network model that is trained through knowledge distillation by learning from one or more teacher models. In this framework, the student transformer optimizes its parameters not only on ground-truth training data but also by mimicking the output probabilities, intermediate feature representations, or architectural inductive biases transferred from the teacher models. This training strategy enables the student transformer to inherit useful task-specific knowledge, improve data efficiency, and achieve strong predictive performance while often maintaining a smaller parameter size and lower computational footprint for practical deployment.