A spatio-temporal adapter is a lightweight neural network module designed to adapt pre-trained static image models for video understanding tasks through parameter-efficient fine-tuning. By integrating spatial and temporal reasoning mechanisms into a compact architecture, the adapter can be inserted into the layers of a frozen visual backbone to capture dynamic motion and cross-frame relationships without modifying the core pre-trained weights. This modular approach allows models lacking native temporal awareness to process sequential video data while updating only a small fraction of parameters, significantly reducing computational and storage costs compared to full model retraining.