Multi-text training loss is a contrastive loss function used in vision-language pre-training that simultaneously aligns an individual visual input with multiple corresponding text descriptions. Rather than pairing an image with only a single caption or randomly selecting a single text augmentation per training iteration, this formulation treats multiple diverse textual descriptions, such as original captions and rewritten variants, as positive matches for the same image. By optimizing a multi-positive contrastive objective across the batch, multi-text training loss leverages richer linguistic diversity in vocabulary and syntactic phrasing, thereby mitigating text overfitting and improving cross-modal representation alignment and zero-shot transfer performance.