CrossFormer is a vision transformer architecture designed to capture and integrate visual features across multiple spatial scales for computer vision applications. Unlike conventional vision transformers that rely on uniform-scale token embeddings or discard fine-grained details to reduce computational costs, CrossFormer uses a cross-scale embedding layer that blends image patches of different sizes into unified representations. It processes these multi-scale features using a long-short distance attention mechanism, which divides self-attention into local interactions for fine details and long-range interactions for broader context. Coupled with a dynamic position bias that allows the model to flexibly adapt to variable input image dimensions, the architecture serves as a versatile backbone for tasks such as image classification, object detection, and semantic segmentation.