Transformer in Transformer is a visual neural network architecture that captures both fine-grained local details and global structural relationships by utilizing nested transformer modules. Standard vision transformers partition an input image into a sequence of coarse patches and apply self-attention exclusively across those patches, which can miss localized sub-patch features. To address this, Transformer in Transformer further divides each coarse patch into smaller sub-patches, employing an inner transformer block to compute self-attention among the sub-patches within a single patch while an outer transformer block models the interactions across all coarse patches. By combining these nested inner and outer representations, the architecture effectively learns multi-scale visual features for tasks such as image recognition without incurring substantial computational overhead.