Attention decomposition is a technique in neural network architecture where a dense, multidimensional self-attention computation is factorized into lower-dimensional sub-operations to improve computational efficiency. In transformer and attention-based models, computing global interactions across an entire multidimensional input, such as a two-dimensional image, typically incurs quadratic computational and memory costs relative to the sequence length. Attention decomposition resolves this scalability bottleneck by separating the attention mechanism along distinct coordinate axes, such as performing sequential or factorized passes over horizontal and vertical dimensions. This structured reduction allows models to capture long-range contextual relationships and integrate explicit spatial priors with linear complexity, avoiding the prohibitive overhead of full pairwise token interaction.