Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action Localization
Huan RenWenfei YangTianzhu ZhangYongdong Zhang
Proposes a proposal-based multiple instance learning framework for weakly-supervised temporal action localization that aligns training and testing objectives by directly classifying candidate action proposals rather than individual video segments.
Identifying and localizing specific human actions within long, unedited videos is essential for high-impact applications such as automated surveillance, content summarization, and search. While traditional artificial intelligence models require precise, time-consuming annotations for every action instance, weakly-supervised methods lower deployment costs by training models using only overall video-level category labels. However, existing standard approaches score short video snippets individually during training but evaluate entire action proposals during testing. This mismatch leads to poor localization because isolated snippets often lack sufficient context to distinguish complex activities.
The article develops and evaluates a Proposal-based Multiple Instance Learning framework that eliminates this discrepancy by directly classifying complete candidate action proposals during both training and evaluation.
To evaluate this framework, the authors conducted extensive experiments using standard benchmark video datasets, including THUMOS14 (over 400 untrimmed videos) and ActivityNet versions 1.2 and 1.3 (spanning up to nearly 20,000 videos across up to 200 categories). The system first generates initial candidate action and background proposals, extracts context-rich features by contrasting each proposal with its immediate surrounding temporal regions, evaluates proposal completeness using automatically generated pseudo-labels, and enforces ranking consistency across visual appearance and motion streams.
The key findings demonstrate that this direct proposal-based approach significantly outperforms prior methods. On the THUMOS14 benchmark, the framework established a new state-of-the-art weakly-supervised detection accuracy, achieving an average mean Average Precision of 46.5%, which increased to 47.0% when combined with initial segment predictions (outperforming the previous best benchmark by 1.9 percentage points). On ActivityNet 1.2 and 1.3, the system achieved leading average accuracies of 26.5% and 25.5%, respectively. Ablation analyses showed that incorporating background proposals during training improved detection accuracy by 5.3 percentage points, while outer-inner contrastive feature extraction boosted accuracy by 5.5 percentage points compared to unextended boundaries.
These results indicate that video understanding systems can achieve high temporal precision without expensive frame-by-frame labeling, substantially lowering data annotation costs and engineering timelines. Furthermore, the findings demonstrate that localization bottlenecks in weakly-supervised video models stem primarily from how proposals are scored rather than how initial boundaries are generated.
Organizations implementing automated video analytics should consider adopting direct proposal-scoring frameworks and incorporating surrounding temporal context to refine action detection. For next steps, teams should pilot this architecture on domain-specific video streams to evaluate its performance under real-world noise. Confidence in these findings is high across standard public benchmarks, though performance boundaries remain constrained by the quality of initial proposal generation and potential visual ambiguities across diverse real-world operating environments.
- Paper: Temporal Segment Networks: Towards Good Practices for Deep Action Recognition, Limin Wang et al. (2016). It introduces the segment-based two-stream feature modeling paradigm that serves as the standard foundational input representation for temporal action localization.
- Paper: Bridging the Gap between Classification and Localization for Weakly Supervised Object Localization, Eunji Kim et al. (2022). It examines the core structural discrepancy between classification objectives and precise boundary localization under weak supervision.
- Paper: Two-Stream Convolutional Networks for Action Recognition in Videos, Karen Simonyan et al. (2014). It establishes the complementary spatial RGB and temporal optical flow architecture leveraged for multi-modal action detection.
- Paper: TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection, Hao Sun et al. (2024). It extends temporal boundary localization concepts by coupling proposal retrieval and highlight detection through reciprocal transformer interactions.
- Paper: TRACE: Temporal Grounding Video LLM via Causal Event Modeling, Yongxin Guo 0001 et al. (2025). It builds on temporal grounding mechanisms by formalizing video action boundaries and saliency into structured multimodal LLM event modeling.
- Paper: TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning, Xiangyu Zeng 0004 et al. (2025). It generalizes temporal action localization principles to long-video understanding via instruction-tuned temporal grounding.
