Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action Localization

Huan RenWenfei YangTianzhu ZhangYongdong Zhang

article2023CVPR50 citations

Proposes a proposal-based multiple instance learning framework for weakly-supervised temporal action localization that aligns training and testing objectives by directly classifying candidate action proposals rather than individual video segments.

Listen

Identifying and localizing specific human actions within long, unedited videos is essential for high-impact applications such as automated surveillance, content summarization, and search. While traditional artificial intelligence models require precise, time-consuming annotations for every action instance, weakly-supervised methods lower deployment costs by training models using only overall video-level category labels. However, existing standard approaches score short video snippets individually during training but evaluate entire action proposals during testing. This mismatch leads to poor localization because isolated snippets often lack sufficient context to distinguish complex activities.

The article develops and evaluates a Proposal-based Multiple Instance Learning framework that eliminates this discrepancy by directly classifying complete candidate action proposals during both training and evaluation.

To evaluate this framework, the authors conducted extensive experiments using standard benchmark video datasets, including THUMOS14 (over 400 untrimmed videos) and ActivityNet versions 1.2 and 1.3 (spanning up to nearly 20,000 videos across up to 200 categories). The system first generates initial candidate action and background proposals, extracts context-rich features by contrasting each proposal with its immediate surrounding temporal regions, evaluates proposal completeness using automatically generated pseudo-labels, and enforces ranking consistency across visual appearance and motion streams.

The key findings demonstrate that this direct proposal-based approach significantly outperforms prior methods. On the THUMOS14 benchmark, the framework established a new state-of-the-art weakly-supervised detection accuracy, achieving an average mean Average Precision of 46.5%, which increased to 47.0% when combined with initial segment predictions (outperforming the previous best benchmark by 1.9 percentage points). On ActivityNet 1.2 and 1.3, the system achieved leading average accuracies of 26.5% and 25.5%, respectively. Ablation analyses showed that incorporating background proposals during training improved detection accuracy by 5.3 percentage points, while outer-inner contrastive feature extraction boosted accuracy by 5.5 percentage points compared to unextended boundaries.

These results indicate that video understanding systems can achieve high temporal precision without expensive frame-by-frame labeling, substantially lowering data annotation costs and engineering timelines. Furthermore, the findings demonstrate that localization bottlenecks in weakly-supervised video models stem primarily from how proposals are scored rather than how initial boundaries are generated.

Organizations implementing automated video analytics should consider adopting direct proposal-scoring frameworks and incorporating surrounding temporal context to refine action detection. For next steps, teams should pilot this architecture on domain-specific video streams to evaluate its performance under real-world noise. Confidence in these findings is high across standard public benchmarks, though performance boundaries remain constrained by the quality of initial proposal generation and potential visual ambiguities across diverse real-world operating environments.

arXiv: 2305.17861
Cover for Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action Localization

Abstract

Weakly-supervised temporal action localization aims to localize and recognize actions in untrimmed videos with only video-level category labels during training. Without instance-level annotations, most existing methods follow the Segment-based Multiple Instance Learning (S-MIL) framework, where the predictions of segments are supervised by the labels of videos. However, the objective for acquiring segment-level scores during training is not consistent with the target for acquiring proposal-level scores during testing, leading to suboptimal results. To deal with this problem, we propose a novel Proposal-based Multiple Instance Learning (P-MIL) framework that directly classifies the candidate proposals in both the training and testing stages, which includes three key designs: 1) a surrounding contrastive feature extraction module to suppress the discriminative short proposals by considering the surrounding contrastive information, 2) a proposal completeness evaluation module to inhibit the low-quality proposals with the guidance of the completeness pseudo labels, and 3) an instance-level rank consistency loss to achieve robust detection by leveraging the complementarity of RGB and FLOW modalities. Extensive experimental results on two challenging benchmarks including THUMOS14 and ActivityNet demonstrate the superior performance of our method. Our code is available at github.com/RenHuan1999/CVPR2023-P-MIL.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Our Method
  • 3.1. Candidate Proposal Generation
  • 3.2. Proposal Feature Extraction and Classification
  • 3.3. Proposal Refinement
  • 3.4. Network Training and Inference
  • 3.5. Discussions
  • 4. Experiment
  • 4.1. Datasets and Evaluation Metrics
  • 4.2. Implementation Details
  • 4.3. Comparison with State-of-the-art Methods
  • 4.4. Ablation Studies
  • 5. Conclusion
  • 6. Acknowledgement
  • References

Knowls

  1. Knowl 1 — Proposal-based Multiple Instance Learning (P-MIL) Framework for WTAL

    model/method

    Weakly-supervised temporal action localization (WTAL) models based on Segment-based Multiple Instance Learning (S-MIL) suffer from an objective inconsistency: the network classifies individual segments during training but must score multi-segment candidate proposals during inference. Proposal-based Multiple Instance Learning (P-MIL) eliminates this discrepancy by operating directly on candidate proposals in both training and testing via a two-stage pipeline:

    1. Candidate Proposal Generation: A standard two-branch S-MIL model is trained with video-level classification and attention sparsity losses on snippet features to output a temporal attention sequence and class activation sequences (CAS). Action proposals PactP_{act} and background proposals PbkgP_{bkg} are extracted by thresholding the attention sequence.
    2. Proposal-level Learning and Classification: Candidate proposals are represented as contrastive feature vectors using boundary expansion. A proposal classification head predicts foreground attention weights and category probabilities, which are aggregated via top-kk pooling to predict video-level labels under standard video-level supervision.
    3. Proposal Refinement: Low-quality or over-complete proposals are suppressed using pseudo-label-guided completeness evaluation, and candidate proposal relative rankings are regularized via cross-modal rank consistency between RGB and optical flow streams.
  2. Knowl 2 — Surrounding Contrastive Feature Extraction (SCFE)

    model/method

    Because video-level classification losses encourage classifiers to focus on the most discriminative short segments rather than full action durations, the Surrounding Contrastive Feature Extraction (SCFE) module explicitly encodes surrounding temporal context into proposal representations.

    Given snippet-level video features XS∈RT×DX_S \in \mathbb{R}^{T \times D} (where TT is the number of snippets and DD is feature dimensionality) and a candidate proposal Pi=(si,ei)P_i = (s_i, e_i) with temporal duration li=ei−sil_i = e_i - s_i:

    1. The proposal boundaries are expanded outward by a fraction α\alpha of its duration on both sides, partitioning the neighborhood into three contiguous temporal intervals: left context [si−αli,si][s_i - \alpha l_i, s_i], inner action [si,ei][s_i, e_i], and right context [ei,ei+αli][e_i, e_i + \alpha l_i] (with α=0.25\alpha = 0.25).
    2. 1D RoIAlign followed by max-pooling is applied over XSX_S for each region (using output RoI bin sizes of 2, 8, and 2 for left, inner, and right regions, respectively), producing three DD-dimensional feature vectors: Xil∈RDX_i^l \in \mathbb{R}^D, Xin∈RDX_i^n \in \mathbb{R}^D, and Xir∈RDX_i^r \in \mathbb{R}^D.
    3. The final proposal feature Xi∈RDX_i \in \mathbb{R}^D is obtained by subtracting the outer context features from the inner feature and projecting the concatenated representation through a fully connected layer (FCFC):

    Xi=FC(Cat(Xin−Xil, Xin, Xin−Xir))X_i = FC\left(\text{Cat}(X_i^n - X_i^l, \, X_i^n, \, X_i^n - X_i^r)\right)

    where Cat(⋅)\text{Cat}(\cdot) denotes concatenation along the channel dimension.

  3. Knowl 3 — Proposal Completeness Evaluation (PCE) Module

    model/method

    The Proposal Completeness Evaluation (PCE) module suppresses over-complete proposals containing irrelevant background by supervising a completeness prediction head with pseudo ground-truth Intersection over Union (IoU) labels.

    1. Pseudo Instance Mining: Given candidate proposals P={Pi}i=1MP = \{P_i\}_{i=1}^M and their predicted proposal foreground attention weights A∈RM×1A \in \mathbb{R}^{M \times 1}, a high-confidence proposal subset Q={Pi∣A(i)≥γ⋅max⁡(A)}Q = \{P_i \mid A(i) \ge \gamma \cdot \max(A)\} is formed with threshold parameter γ=0.8\gamma = 0.8. A non-maximum suppression (NMS) routine iteratively selects the proposal in QQ with the highest attention weight as a pseudo instance, removes all proposals overlapping with it from QQ, and repeats until QQ is empty, yielding a set of pseudo instances G={(sj,ej)}j=1NG = \{(s_j, e_j)\}_{j=1}^N.
    2. Completeness Pseudo-Label Assignment: For each candidate proposal Pi∈PP_i \in P, its completeness pseudo label q(i)∈[0,1]q(i) \in [0, 1] is defined as its maximum temporal IoU with any pseudo instance in GG:

    q(i)=max⁡j∈{1,…,N}IoU(Pi,Gj)q(i) = \max_{j \in \{1, \dots, N\}} \text{IoU}(P_i, G_j)

    1. Completeness Loss: A completeness branch (two fully connected layers with sigmoid activation) predicts completeness scores q^∈RM×1\hat{q} \in \mathbb{R}^{M \times 1} in parallel with the classification head, trained via Mean Squared Error (MSE):

    Lcomp=1M∑i=1M(q(i)−q^(i))2L_{comp} = \frac{1}{M} \sum_{i=1}^M \left(q(i) - \hat{q}(i)\right)^2

  4. Knowl 4 — Instance-level Rank Consistency (IRC) Loss

    equation

    The Instance-level Rank Consistency (IRC) loss enforces consistency between the relative score distributions predicted by RGB and optical flow streams across clusters of overlapping proposals, improving candidate ranking robustness for test-time Non-Maximum Suppression (NMS).

    1. Candidate proposals with attention weight above the video mean, A(i)≥mean(A)A(i) \ge \text{mean}(A), form a filtered set RR. For each anchor proposal r∈Rr \in R, all candidate proposals in PP overlapping with rr form a cluster Ωr\Omega_r of size Nr=∣Ωr∣N_r = |\Omega_r|.
    2. For ground-truth video action class cc, the unnormalized classification scores corresponding to the proposals in Ωr\Omega_r are retrieved from the base classification scores SbaseS_{base} of the RGB and FLOW branches, denoted pr,cRGB∈RNrp_{r,c}^{RGB} \in \mathbb{R}^{N_r} and pr,cFLOW∈RNrp_{r,c}^{FLOW} \in \mathbb{R}^{N_r}.
    3. The intra-cluster score distributions are normalized via softmax for modality ∗∈{RGB,FLOW}* \in \{RGB, FLOW\}:

    Dr,c∗=softmax(pr,c∗)D_{r,c}^* = \text{softmax}(p_{r,c}^*)

    1. The IRC loss is the bidirectional Kullback-Leibler (KL) divergence averaged over all anchor clusters r∈Rr \in R:

    LIRC=1∣R∣∑r∈R(KL(Dr,cFLOW∥Dr,cRGB)+KL(Dr,cRGB∥Dr,cFLOW))L_{IRC} = \frac{1}{|R|} \sum_{r \in R} \left( \text{KL}(D_{r,c}^{FLOW} \parallel D_{r,c}^{RGB}) + \text{KL}(D_{r,c}^{RGB} \parallel D_{r,c}^{FLOW}) \right)

    where for probability vectors Dr,ct,Dr,cs∈RNrD_{r,c}^t, D_{r,c}^s \in \mathbb{R}^{N_r}, the discrete KL divergence is:

    KL(Dr,ct∥Dr,cs)=−∑i=1NrDr,ct(i)log⁡Dr,cs(i)Dr,ct(i)\text{KL}(D_{r,c}^t \parallel D_{r,c}^s) = -\sum_{i=1}^{N_r} D_{r,c}^t(i) \log \frac{D_{r,c}^s(i)}{D_{r,c}^t(i)}

  5. Knowl 5 — Candidate Proposal Generation via Multi-Threshold Attention

    model/method

    In the initial stage of P-MIL, an S-MIL model is trained on non-overlapping 16-frame segment features XS∈RT×DX_S \in \mathbb{R}^{T \times D}. A category-agnostic attention branch outputs an attention sequence A∈RT×1A \in \mathbb{R}^{T \times 1}, and a classification branch predicts base snippet-level Class Activation Sequences (CAS) Sbase∈RT×(C+1)S_{base} \in \mathbb{R}^{T \times (C+1)}, where index C+1C+1 represents background. Multiplying SbaseS_{base} elementwise with AA along the temporal dimension yields background-suppressed CAS Ssupp=Sbase⊙AS_{supp} = S_{base} \odot A.

    Top-kk pooling followed by softmax over SbaseS_{base} and SsuppS_{supp} produces video-level predictions y^base,y^supp∈RC+1\hat{y}_{base}, \hat{y}_{supp} \in \mathbb{R}^{C+1}, optimized by:

    LtotalSMIL=−∑c=1C+1(ybase(c)log⁡y^base(c)+ysupp(c)log⁡y^supp(c))+λnorm1T∑t=1T∣A(t)∣L_{total}^{SMIL} = -\sum_{c=1}^{C+1} \left( y_{base}(c) \log \hat{y}_{base}(c) + y_{supp}(c) \log \hat{y}_{supp}(c) \right) + \lambda_{norm} \frac{1}{T} \sum_{t=1}^T |A(t)|

    where ybase=[y,1]y_{base} = [y, 1] and ysupp=[y,0]y_{supp} = [y, 0] for video label y∈{0,1}Cy \in \{0, 1\}^C.

    From the learned attention sequence AA, candidate temporal intervals are extracted by applying multiple thresholds:

    • Action proposals Pact={(si,ei)}i=1M1P_{act} = \{(s_i, e_i)\}_{i=1}^{M_1} are continuous temporal intervals where A(t)≥θactA(t) \ge \theta_{act}, sampled across θact∈{0.1,0.2,…,0.9}\theta_{act} \in \{0.1, 0.2, \dots, 0.9\}.
    • Background proposals Pbkg={(si,ei)}i=1M2P_{bkg} = \{(s_i, e_i)\}_{i=1}^{M_2} are continuous intervals where A(t)<θbkgA(t) < \theta_{bkg}, sampled across θbkg∈{0.3,0.5,0.7}\theta_{bkg} \in \{0.3, 0.5, 0.7\}.

    The training proposal set is P=Pact∪PbkgP = P_{act} \cup P_{bkg} (M=M1+M2M = M_1 + M_2 proposals), enabling explicit foreground-background contrast during proposal classification. Inference exclusively processes PactP_{act}.

  6. Knowl 6 — P-MIL Network Training and Inference Pipeline

    algorithm

    The unified training and inference procedure for Proposal-based Multiple Instance Learning (P-MIL) is structured as follows:

    Input: Video segment features XS∈RT×DX_S \in \mathbb{R}^{T \times D} for RGB and FLOW modalities, ground-truth video label y∈{0,1}Cy \in \{0, 1\}^C, candidate proposals P=Pact∪PbkgP = P_{act} \cup P_{bkg} of total count MM.
    Output: Predicted action instances {(ci,si,ei,s(i))}\{(c_i, s_i, e_i, s(i))\}.
    Training Stage:
      Extract proposal features XP∈RM×DX_P \in \mathbb{R}^{M \times D} via SCFE using RoIAlign and outer-inner contrast.
      Predict proposal attention weights A∈RM×1A \in \mathbb{R}^{M \times 1} and base class scores Sbase∈RM×(C+1)S_{base} \in \mathbb{R}^{M \times (C+1)}.
      Compute background-suppressed proposal scores Ssupp=Sbase⊙AS_{supp} = S_{base} \odot A.
      Aggregate SbaseS_{base} and SsuppS_{supp} via top-k temporal pooling to obtain video-level scores y^base,y^supp∈RC+1\hat{y}_{base}, \hat{y}_{supp} \in \mathbb{R}^{C+1}.
      Compute video classification loss Lcls=−∑c=1C+1[ybase(c)log⁡y^base(c)+ysupp(c)log⁡y^supp(c)]L_{cls} = -\sum_{c=1}^{C+1} [y_{base}(c) \log \hat{y}_{base}(c) + y_{supp}(c) \log \hat{y}_{supp}(c)].
      Generate completeness pseudo labels q∈RMq \in \mathbb{R}^M via PCE and predict completeness scores q^∈RM\hat{q} \in \mathbb{R}^M.
      Compute proposal completeness loss Lcomp=1M∑i=1M(q(i)−q^(i))2L_{comp} = \frac{1}{M} \sum_{i=1}^M (q(i) - \hat{q}(i))^2.
      Compute Instance-level Rank Consistency loss LIRCL_{IRC} across overlapping clusters between RGB and FLOW branches.
      Update network parameters using total loss Ltotal=Lcls+λcompLcomp+λIRCLIRCL_{total} = L_{cls} + \lambda_{comp} L_{comp} + \lambda_{IRC} L_{IRC} (with λcomp=20\lambda_{comp}=20, λIRC=2\lambda_{IRC}=2).
    Inference Stage:
      Filter active video classes: retain category c∈{1,…,C}c \in \{1, \dots, C\} if y^supp(c)≥θcls\hat{y}_{supp}(c) \ge \theta_{cls} (with θcls=0.2\theta_{cls} = 0.2).
      For each retained category cc and each action proposal Pi=(si,ei)∈PactP_i = (s_i, e_i) \in P_{act}:
        Compute detection confidence score s(i)=Ssupp(i,c)⋅q^(i)s(i) = S_{supp}(i, c) \cdot \hat{q}(i).
      Apply class-wise Soft-NMS on scored proposals to remove overlapping duplicate predictions.
  7. Knowl 7 — Temporal Action Localization Performance on THUMOS14

    data/table

    Localization performance of P-MIL compared with weakly-supervised and fully-supervised methods on the THUMOS14 test set. All methods use two-stream I3D features pretrained on Kinetics-400. Performance is evaluated using mean Average Precision (mAP, %) at temporal IoU thresholds 0.1 to 0.7, as well as average mAP across ranges 0.1:0.5, 0.3:0.7, and 0.1:0.7.

    Could not parse LaTeX table

    P-MIL achieves 39.8%39.8\% [email protected] and 46.5%46.5\% average [email protected]:0.7, outperforming existing WTAL baselines. Fusing the detection predictions of the S-MIL model and the P-MIL model (Ours*) further increases average mAP to 47.0%47.0\%.

  8. Knowl 8 — Temporal Action Localization Performance on ActivityNet 1.2 and 1.3

    data/table

    Localization results of P-MIL on the ActivityNet1.2 (100 classes) and ActivityNet1.3 (200 classes) validation sets using I3D features. Evaluation metric is mAP (%) at IoU thresholds 0.5, 0.75, 0.95, and average mAP across thresholds 0.5:0.05:0.95 (AVG).

    Could not parse LaTeX table

    On ActivityNet 1.2, P-MIL achieves 25.5%25.5\% average mAP alone and 26.5%26.5\% average mAP when fused with S-MIL. On ActivityNet 1.3, P-MIL achieves 23.9%23.9\% average mAP alone and 25.5%25.5\% average mAP when fused with S-MIL.

  9. Knowl 9 — Ablation Analysis of P-MIL Architecture and Proposal Operations

    empirical result

    Ablation experiments conducted on the THUMOS14 benchmark evaluate the contribution of individual design choices across mAP@IoU (0.1, 0.3, 0.5, 0.7) and average mAP (AVG, 0.1:0.7):

    1. Candidate Proposal Generation: Training the proposal classifier only on foreground action proposals (PactP_{act}) results in 41.2%41.2\% average mAP. Augmenting training with background proposals (Pact∪PbkgP_{act} \cup P_{bkg}) increases average mAP to 46.5%46.5\% (+5.3%+5.3\%), verifying the necessity of background negative samples during proposal classification.
    2. Direct vs. Indirect Proposal Scoring: On the identical set of candidate proposals:
      • S-MIL indirect segment CAS aggregation yields 43.6%43.6\% average mAP.
      • P-MIL direct proposal classification yields 46.5%46.5\% average mAP (+2.9%+2.9\% gain).
      • Late fusion of S-MIL and P-MIL scores reaches 47.0%47.0\% average mAP.
      • An upper bound scoring candidate proposals by their ground-truth IoU achieves 67.4%67.4\% average mAP, showing that proposal localization quality is high and scoring accuracy is the primary performance bottleneck.
    3. Feature Extraction Variants:
      • Proposal representation without temporal boundary extension: 41.0%41.0\% average mAP.
      • Simple concatenation of left context, inner, and right context features: 41.9%41.9\% average mAP (+0.9%+0.9\%).
      • Outer-inner contrast representation Cat(Xin−Xil,Xin,Xin−Xir)\text{Cat}(X_i^n - X_i^l, X_i^n, X_i^n - X_i^r): 46.5%46.5\% average mAP (+5.5%+5.5\% over non-extended baseline).
    4. Proposal Refinement Components:
      • Baseline P-MIL without refinement: 45.2%45.2\% average mAP.
      • With Proposal Completeness Evaluation (PCE): 45.9%45.9\% average mAP (+0.7%+0.7\%).
      • With Instance-level Rank Consistency (IRC): 46.0%46.0\% average mAP (+0.8%+0.8\%).
      • With both PCE and IRC: 46.5%46.5\% average mAP (+1.3%+1.3\% total improvement).

Coverage note — None was omitted; all key contributions, mathematical formulations, algorithmic steps, benchmark results, and ablation studies from the paper are represented.

References

  1. 1.Hakan Bilen and Andrea Vedaldi. Weakly supervised deep detection networks. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 2846–2854, 2016.
  2. 2.Navaneeth Bodla, Bharat Singh, Rama Chellappa, and Larry S Davis. Soft-nms–improving object detection with one line of code. In Proceedings of the IEEE International Conference on Computer Vision, pages 5561–5569, 2017.
  3. 3.Shyamal Buch, Victor Escorcia, Bernard Ghanem, Li Fei-Fei, and Juan Carlos Niebles. End-to-end, single-stream temporal action detection in untrimmed videos. In Proceedings of the British Machine Vision Conference, volume 2, page 7, 2017.
  4. 4.Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 961–970, 2015.
  5. 5.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision, pages 213–229. Springer, 2020.
  6. 6.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6299–6308, 2017.
  7. 7.Yu-Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold, David A Ross, Jia Deng, and Rahul Sukthankar. Rethinking the faster r-cnn architecture for temporal action localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1130–1139, 2018.
  8. 8.Feng Cheng and Gedas Bertasius. Tallformer: Temporal action localization with long-memory transformer. In Proceedings of the European Conference on Computer Vision, pages 503–521, 2022.
  9. 9.Adrien Gaidon, Zaid Harchaoui, and Cordelia Schmid. Temporal localization of actions with actoms. IEEE transactions on Pattern Analysis and Machine Intelligence, 35(11):2782–2795, 2013.
  10. 10.Junyu Gao, Mengyuan Chen, and Changsheng Xu. Fine-grained temporal contrastive learning for weakly-supervised temporal action localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19999–20009, 2022.
  11. 11.Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Region-based convolutional networks for accurate object detection and segmentation. IEEE transactions on Pattern Analysis and Machine Intelligence, 38(1):142–158, 2015.
  12. 12.Guoqiang Gong, Xinghan Wang, Yadong Mu, and Qi Tian. Learning temporal co-attention models for unsupervised video action localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9819–9828, 2020.
  13. 13.Bo He, Xitong Yang, Le Kang, Zhiyu Cheng, Xin Zhou, and Abhinav Shrivastava. Asm-loc: Action-aware segment modeling for weakly-supervised temporal action localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13925–13935, 2022.
  14. 14.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Gir-´ shick. Mask r-cnn. In Proceedings of the IEEE International Conference on Computer Vision, pages 2961–2969, 2017.
  15. 15.Fa-Ting Hong, Jia-Chang Feng, Dan Xu, Ying Shan, and Wei-Shi Zheng. Cross-modal consensus network for weakly supervised temporal action localization. In Proceedings of the 29th ACM International Conference on Multimedia, pages 1591–1599, 2021.
  16. 16.Linjiang Huang, Liang Wang, and Hongsheng Li. Foreground-action consistency network for weakly supervised temporal action localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8002–8011, 2021.
  17. 17.Linjiang Huang, Liang Wang, and Hongsheng Li. Weakly supervised temporal action localization via representative snippet knowledge propagation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3272–3281, 2022.
  18. 18.Haroon Idrees, Amir R Zamir, Yu-Gang Jiang, Alex Gorban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah. The thumos challenge on action recognition for videos “in the wild”. Computer Vision and Image Understanding, 155:1–23, 2017.
  19. 19.Ashraful Islam, Chengjiang Long, and Richard Radke. A hybrid attention mechanism for weakly-supervised temporal action localization. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 1637–1645, 2021.
  20. 20.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. Internation Conference on Learning Representations, 2014.
  21. 21.Pilhyeon Lee, Youngjung Uh, and Hyeran Byun. Background suppression network for weakly-supervised temporal action localization. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 11320–11327, 2020.
  22. 22.Yong Jae Lee, Joydeep Ghosh, and Kristen Grauman. Discovering important people and objects for egocentric video summarization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1346–1353, 2012.
  23. 23.Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg. Tvqa: Localized, compositional video question answering. Proceedings of the Conference on Empirical Methods in Natural Language Processing, 2018.
  24. 24.Jingjing Li, Tianyu Yang, Wei Ji, Jue Wang, and Li Cheng. Exploring denoised cross-video contrast for weakly-supervised temporal action localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19914–19924, 2022.
  25. 25.Tianwei Lin, Xiao Liu, Xin Li, Errui Ding, and Shilei Wen. Bmn: Boundary-matching network for temporal action proposal generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3889–3898, 2019.
  26. 26.Daochang Liu, Tingting Jiang, and Yizhou Wang. Completeness modeling and context separation for weakly supervised temporal action localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1298–1307, 2019.
  27. 27.Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg. Ssd: Single shot multibox detector. In European Conference on Computer Vision, pages 21–37, 2016.
  28. 28.Yuan Liu, Jingyuan Chen, Zhenfang Chen, Bing Deng, Jianqiang Huang, and Hanwang Zhang. The blessings of unlabeled background in untrimmed videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6176–6185, 2021.
  29. 29.Ziyi Liu, Le Wang, Qilin Zhang, Zhanning Gao, Zhenxing Niu, Nanning Zheng, and Gang Hua. Weakly supervised temporal action localization through contrast based evaluation networks. In Proceedings of the IEEE International Conference on Computer Vision, pages 3899–3908, 2019.
  30. 30.Fuchen Long, Ting Yao, Zhaofan Qiu, Xinmei Tian, Jiebo Luo, and Tao Mei. Gaussian temporal awareness networks for action localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 344–353, 2019.
  31. 31.Wang Luo, Tianzhu Zhang, Wenfei Yang, Jingen Liu, Tao Mei, Feng Wu, and Yongdong Zhang. Action unit memory network for weakly supervised temporal action localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9969–9979, 2021.
  32. 32.Zhekun Luo, Devin Guillory, Baifeng Shi, Wei Ke, Fang Wan, Trevor Darrell, and Huijuan Xu. Weakly-supervised action localization with expectation-maximization multiinstance learning. In European Conference on Computer Vision, pages 729–745, 2020.
  33. 33.Oded Maron and Tomas Lozano-P´erez. A framework for multiple-instance learning. Advances in Neural Information Processing Systems, 10, 1997.
  34. 34.Niluthpol Chowdhury Mithun, Sujoy Paul, and Amit K Roy-Chowdhury. Weakly supervised video moment retrieval from text queries. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11592–11601, 2019.
  35. 35.Sanath Narayan, Hisham Cholakkal, Munawar Hayat, Fahad Shahbaz Khan, Ming-Hsuan Yang, and Ling Shao. D2-net: Weakly-supervised action localization via discriminative embeddings and denoised activations. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13608–13617, 2021.
  36. 36.Phuc Nguyen, Ting Liu, Gautam Prasad, and Bohyung Han. Weakly supervised action localization by sparse temporal pooling network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6752–6761, 2018.
  37. 37.Phuc Xuan Nguyen, Deva Ramanan, and Charless C Fowlkes. Weakly-supervised action localization with background modeling. In Proceedings of the IEEE International Conference on Computer Vision, pages 5502–5511, 2019.
  38. 38.Sujoy Paul, Sourya Roy, and Amit K Roy-Chowdhury. W-talc: Weakly-supervised temporal activity localization and classification. In Proceedings of the European Conference on Computer Vision, pages 563–579, 2018.
  39. 39.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in Neural Information Processing Systems, 28, 2015.
  40. 40.Baifeng Shi, Qi Dai, Yadong Mu, and Jingdong Wang. Weakly-supervised action localization by generative attention modeling. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1009–1019, 2020.
  41. 41.Zheng Shou, Hang Gao, Lei Zhang, Kazuyuki Miyazawa, and Shih-Fu Chang. Autoloc: Weakly-supervised temporal action localization in untrimmed videos. In Proceedings of the European Conference on Computer Vision, pages 154–171, 2018.
  42. 42.Zheng Shou, Dongang Wang, and Shih-Fu Chang. Temporal action localization in untrimmed videos via multi-stage cnns. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1049–1058, 2016.
  43. 43.Peng Tang, Xinggang Wang, Xiang Bai, and Wenyu Liu. Multiple instance detection network with online instance classifier refinement. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 2843–2851, 2017.
  44. 44.Sarvesh Vishwakarma and Anupam Agrawal. A survey on activity recognition and behavior understanding in video surveillance. The Visual Computer, 29(10):983–1009, 2013.
  45. 45.Limin Wang, Yuanjun Xiong, Dahua Lin, and Luc Van Gool. Untrimmednets for weakly supervised action recognition and detection. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 4325–4334, 2017.
  46. 46.Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. Temporal segment networks: Towards good practices for deep action recognition. In European Conference on Computer Vision, pages 20–36, 2016.
  47. 47.Yuetian Weng, Zizheng Pan, Mingfei Han, Xiaojun Chang, and Bohan Zhuang. An efficient spatio-temporal pyramid transformer for action detection. In European Conference on Computer Vision, pages 358–375. Springer, 2022.
  48. 48.Kun Xia, Le Wang, Sanping Zhou, Nanning Zheng, and Wei Tang. Learning to refactor action and co-occurrence features for temporal action localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13884–13893, 2022.
  49. 49.Bo Xiong, Yannis Kalantidis, Deepti Ghadiyaram, and Kristen Grauman. Less is more: Learning highlight detection from video duration. In Proceedings of the IEEE/CVF conference on Computer Vision and Pattern Recognition, pages 1258–1267, 2019.
  50. 50.Huijuan Xu, Abir Das, and Kate Saenko. R-c3d: Region convolutional 3d network for temporal activity detection. In Proceedings of the IEEE International Conference on Computer Vision, pages 5783–5792, 2017.
  51. 51.Mengmeng Xu, Chen Zhao, David S Rojas, Ali Thabet, and Bernard Ghanem. G-tad: Sub-graph localization for temporal action detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10156–10165, 2020.
  52. 52.Ke Yang, Dongsheng Li, and Yong Dou. Towards precise end-to-end weakly supervised object detection network. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8372–8381, 2019.
  53. 53.Le Yang, Houwen Peng, Dingwen Zhang, Jianlong Fu, and Junwei Han. Revisiting anchor mechanisms for temporal action localization. IEEE Transactions on Image Processing, 29:8535–8548, 2020.
  54. 54.Wenfei Yang, Tianzhu Zhang, Zhendong Mao, Yongdong Zhang, Qi Tian, and Feng Wu. Multi-scale structure-aware network for weakly supervised temporal action detection. IEEE Transactions on Image Processing, 30:5848–5861, 2021.
  55. 55.Wenfei Yang, Tianzhu Zhang, Xiaoyuan Yu, Tian Qi, Yongdong Zhang, and Feng Wu. Uncertainty guided collaborative training for weakly supervised temporal action detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 53–63, 2021.
  56. 56.Wenfei Yang, Tianzhu Zhang, Yongdong Zhang, and Feng Wu. Local correspondence network for weakly supervised temporal sentence grounding. IEEE Transactions on Image Processing, 30:3252–3262, 2021.
  57. 57.Wenfei Yang, Tianzhu Zhang, Yongdong Zhang, and Feng Wu. Uncertainty guided collaborative training for weakly supervised and unsupervised temporal action localization. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.
  58. 58.Christopher Zach, Thomas Pock, and Horst Bischof. A duality based approach for realtime tv-l 1 optical flow. In Joint Pattern Recognition Symposium, pages 214–223, 2007.
  59. 59.Runhao Zeng, Wenbing Huang, Mingkui Tan, Yu Rong, Peilin Zhao, Junzhou Huang, and Chuang Gan. Graph convolutional networks for temporal action localization. In Proceedings of the IEEE International Conference on Computer Vision, pages 7094–7103, 2019.
  60. 60.Yuanhao Zhai, Le Wang, Wei Tang, Qilin Zhang, Junsong Yuan, and Gang Hua. Two-stream consensus networks for weakly-supervised temporal action localization. In Proceedings of the European Conference on Computer Vision, August 2020.
  61. 61.Can Zhang, Meng Cao, Dongming Yang, Jie Chen, and Yuexian Zou. Cola: Weakly-supervised temporal action localization with snippet contrastive learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recogintion, 2021.
  62. 62.Chenlin Zhang, Jianxin Wu, and Yin Li. Actionformer: Localizing moments of actions with transformers. In European Conference on Computer Vision, 2022.
  63. 63.Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Xiaoou Tang, and Dahua Lin. Temporal action detection with structured segment networks. In Proceedings of the IEEE International Conference on Computer Vision, pages 2914–2923, 2017.
  64. 64.Jia-Xing Zhong, Nannan Li, Weijie Kong, Tao Zhang, Thomas H Li, and Ge Li. Step-by-step erasion, one-by-one collection: A weakly supervised temporal action detector. In Proceedings of the ACM Multimedia Conference on Multimedia Conference, pages 35–44, 2018.
  65. 65.Zixin Zhu, Wei Tang, Le Wang, Nanning Zheng, and Gang Hua. Enriching local and global contexts for temporal action localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13516–13525, 2021.

Citation

MLA
Ren, H., et al. “Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action Localization”. arXiv, 2023, http://arxiv.org/abs/2305.17861v1.
APA
Ren, H., Yang, W., Zhang, T., & Zhang, Y. (2023). Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action Localization. arXiv. http://arxiv.org/abs/2305.17861v1
Chicago
Ren, H., W. Yang, T. Zhang, and Y. Zhang. 2023. “Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action Localization”. arXiv. http://arxiv.org/abs/2305.17861v1.
Harvard
Ren, H. et al. (2023) “Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action Localization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.17861v1.
Vancouver
1. Ren H, Yang W, Zhang T, Zhang Y (2023) Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action Localization. arXiv

BibTeX

@article{ren2023proposal,
  title = {Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action Localization},
  author = {Ren, Huan and Yang, Wenfei and Zhang, Tianzhu and Zhang, Yongdong},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.17861v1},
  eprint = {2305.17861}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE