Long-term video recognition is a computer vision task focused on identifying, classifying, and understanding actions, events, or complex activities across extended temporal durations rather than within isolated, short video clips. While conventional video recognition models typically analyze brief segments spanning only a few seconds due to memory and computational bottlenecks, long-term recognition requires systems to capture and reason over long-range temporal context spanning minutes, hours, or full-length videos. Successfully performing this task involves tracking sequential dependencies, contextual changes, and temporally distant interactions over time, typically through techniques such as persistent memory caching, hierarchical modeling, and efficient temporal aggregation architectures.