A convolution-free architecture is a neural network design that processes spatial or multidimensional data, such as images or multimodal inputs, without using standard convolutional layers. Unlike traditional convolutional neural networks that rely on localized sliding filters to extract features, convolution-free models depend on alternative processing mechanisms, most notably self-attention layers in transformer models or fully connected multilayer perceptrons. By dispensing with convolutional operations and their inherent localized inductive biases, these architectures are capable of capturing global context and modeling long-range dependencies across entire input representations and between diverse data modalities directly from the initial stages of feature extraction.