Image-to-video transfer learning is a machine learning paradigm in computer vision where representations learned by a model pre-trained on static images are adapted to analyze and process video data. Because training large video models from scratch is computationally expensive and constrained by the relative scarcity of annotated video datasets, this cross-modality approach leverages the rich spatial feature extractors developed from large-scale image collections. The framework typically incorporates temporal modeling mechanisms to capture motion patterns and dependencies across consecutive frames, enabling effective performance on downstream video understanding tasks such as action recognition and video retrieval.