An inner transformer is a neural network component within nested vision transformer architectures designed to process fine-grained, local information within individual image patches. While standard visual transformers divide an image into larger patches and model relationships across the entire image, the inner transformer operates at a sub-patch level inside each primary patch, treating smaller subdivisions as local visual tokens. It computes self-attention among these sub-tokens to extract detailed intra-patch textures and structures, which are then linearly projected and merged into the higher-level representations handled by an outer transformer. This hierarchical structure enables the model to capture fine-grained local visual details alongside global context without incurring prohibitive computational costs.