A multiscale feature pyramid is a hierarchical representation in computer vision neural networks that structures visual information across multiple spatial resolutions and channel dimensions. In this architecture, early network layers operate at high spatial resolutions with lower channel capacity to capture fine-grained, low-level details such as edges and textures, while deeper layers progressively reduce spatial resolution and expand channel capacity to represent complex, high-level semantic context. By organizing representations into a tiered multi-resolution structure, multiscale feature pyramids enable visual models to efficiently detect, segment, and recognize visual patterns or objects across varying sizes and scales in images and video.