Multi-head cross attention is a neural network mechanism in transformer architectures that allows a model to selectively focus on and integrate relevant information from one data sequence or modality into another. Unlike self-attention, in which queries, keys, and values are derived from the same input, cross attention generates its queries from a primary input stream while taking its keys and values from a separate context or external source. By projecting these inputs into multiple parallel attention heads, the mechanism enables the network to simultaneously learn diverse contextual relationships and correspondences across distinct representation subspaces before concatenating and projecting the results into a unified output.