Low-latency online video recognition is the computer vision process of identifying actions, behaviors, or visual events within continuous, live video streams with minimal processing delay. Unlike offline video recognition methods that process entire pre-recorded clips using access to future frames, online recognition operates under causal constraints, analyzing incoming frames sequentially as they are captured. Achieving low latency requires computationally efficient temporal modeling and inference architectures that can process frames in real time, enabling immediate decision-making in interactive and time-critical applications such as autonomous navigation, robotics, smart surveillance, and real-time human-computer interaction.