An outer transformer is a neural network component within hierarchical or nested vision transformer architectures that models global relationships among standard-sized image patches. In architectures where visual data is divided into patch-level tokens and further subdivided into finer sub-patch tokens, the outer transformer operates at the coarser patch level to capture long-range dependencies across the entire visual sequence. It receives patch representations that incorporate fine-grained local features extracted by a corresponding inner transformer, applying self-attention across the patch sequence to synthesize global spatial context and generate robust high-level representations.