SVFormer is a deep learning framework that applies vision transformer architecture to semi-supervised video action recognition, enabling computer vision models to identify actions using a small amount of labeled video data alongside extensive unlabeled videos. To generate reliable supervisory signals from unlabeled footage, the framework employs an exponential moving average teacher-student pseudo-labeling mechanism. It also integrates video-specific data augmentation techniques, such as Tube TokenMix, which blends video clips using temporally aligned token masks, and temporal warping, which alters frame pacing to account for complex motion speeds and temporal variations across video sequences.