MS-COCO captions refers to a large-scale collection of human-written natural language descriptions associated with images in the Microsoft Common Objects in Context dataset. Typically providing five distinct descriptive sentences for each image, the dataset captures complex everyday scenes containing multiple objects, background contexts, and interactions. It is widely used in artificial intelligence, computer vision, and multimodal natural language processing as a standard benchmark for training and evaluating models across tasks such as automatic image caption generation, vision-language alignment, cross-modal retrieval, and text-to-image synthesis.