Guided deformable attention is a neural network attention mechanism in computer vision that dynamically samples and aggregates features from flexible, data-dependent spatial locations under the guidance of contextual reference representations. Rather than computing pairwise interactions across every location in a fixed grid, it predicts coordinate offsets to focus attention on a sparse set of highly relevant candidate points. In this guided approach, auxiliary or previously inferred features direct the prediction of sampling offsets, facilitating precise alignment across misaligned regions, such as consecutive video clips or frames. By concentrating attention weights exclusively on these guided sampling positions, the mechanism achieves effective feature fusion and motion compensation while maintaining computational efficiency and manageable memory usage.