Spatial prompt learning is a machine learning adaptation technique that incorporates spatial location cues, such as coordinates, bounding boxes, or regional indicators, as prompts to guide pretrained foundation models. Instead of updating the entire set of model parameters, this method introduces learnable or contextual spatial tokens that direct the attention of vision-language and visual models toward specific regions of interest and the relational interactions between entities. By grounding high-level semantic representations in precise positional information, spatial prompt learning enhances the ability of models to perform fine-grained visual reasoning, isolate targeted objects, and interpret complex interactions across open-ended visual environments.