Text-based person search is a computer vision and cross-modal retrieval task that involves identifying and locating images of a specific individual within a visual database using a natural language textual description as the search query. Often referred to as text-to-image person re-identification, this process enables automated systems to find target pedestrians based on descriptions of physical attributes, such as clothing style, color, carried accessories, and bodily appearance, without requiring an existing reference photo. The primary objective is to align high-level language semantics with visual feature representations, bridging the cross-modal domain gap to accurately match textual descriptions with corresponding individuals across diverse surveillance viewpoints, lighting conditions, poses, and backgrounds.