Cross-scale attention is an attention mechanism in deep learning architectures, particularly vision transformers, that enables direct interaction and information exchange between feature representations extracted at different spatial scales or levels of granularity. While conventional attention mechanisms typically process inputs at a uniform resolution, cross-scale attention allows a model to simultaneously relate fine-grained local details with broad, coarse-grained context. By bridging multiple receptive fields across the network, this approach enhances the model ability to recognize patterns and objects across varying sizes while maintaining both precise spatial boundaries and high-level semantic information.