Efficient Movie Scene Detection using State-Space Transformers

Md Mohaiminul IslamMahmudul HasanKishan Shamsundar AthreyTony BraskichGedas Bertasius

article2023CVPR82 citations

Proposes TranS4mer, a hybrid architecture combining structured state-space sequence modeling with self-attention to detect movie scene boundaries across long video sequences while cutting GPU memory usage by three times compared to standard transformers.

Listen

Accurately identifying movie scenes is critical for understanding overarching video storylines and powers commercial applications such as automated video search, preview creation, and non-disruptive ad placement. Traditional video models struggle with this task because they are limited to short-range temporal windows. While standard attention-based Transformer models can model context, their computation and memory costs scale quadratically, making them prohibitively slow and resource-intensive when applied across long video sequences.

The article demonstrates an end-to-end architecture called TranS4mer that resolves this bottleneck by efficiently capturing both short- and long-range dependencies for movie scene boundary detection. TranS4mer introduces a dual-action building block that pairs standard self-attention with structured state-space sequence modeling. To maintain efficiency, the model restricts self-attention exclusively to frames within individual uninterrupted camera shots, which models local elements while keeping compute costs low. It then uses linear-complexity gated state-space layers across shots to capture long-range narrative context over entire movie segments.

The approach was evaluated on three established scene detection benchmarks—MovieNet (1,100 full-length films), BBC Planet Earth, and OVSD—as well as multiple long-range video classification datasets. TranS4mer consistently achieved state-of-the-art results, outperforming the prior leading method by 3.38% average precision on MovieNet, 4.66% on BBC, and 7.36% on OVSD. Crucially, TranS4mer proved to be 2.68 times faster and required 3.38 times less memory than previous state-of-the-art convolutional networks, while running 2 times faster with 3 times less memory than standard Transformer baselines. It also generalized effectively, securing top performance across long-form movie clip classification and procedural activity recognition.

These findings prove that combining localized attention with global state-space modeling removes the computational barriers of long-video processing without sacrificing accuracy. For industry platforms handling massive video catalogs, TranS4mer provides a pathway to deploy automated, high-precision scene detection at substantially reduced infrastructure and compute costs. Future efforts supported by the article involve expanding the architecture to multi-modal tasks, including natural language video grounding, automated movie summarization, and trailer generation. While the findings are strong, practical production deployments will need to account for domain variations beyond structured film datasets.

Cover for Efficient Movie Scene Detection using State-Space Transformers

Abstract

The ability to distinguish between different movie scenes is critical for understanding the storyline of a movie. However, accurately detecting movie scenes is often challenging as it requires the ability to reason over very long movie segments. This contrasts with most existing video recognition models, which are typically designed for short-range video analysis. This work proposes a State-Space Transformer model that can efficiently capture dependencies in long movie videos for accurate movie scene detection. Our model, called TranS4mer, is built using a novel S4A building block, combining the strengths of structured state-space sequence (S4) and self-attention (A) layers. Given a sequence of frames divided into movie shots (uninterrupted periods where the camera position does not change), the S4A block first applies self-attention to capture short-range intra-shot dependencies. Afterward, the state-space operation in the S4A block aggregates long-range inter-shot cues. The final TranS4mer model, which can be trained end-to-end, is obtained by stacking the S4A blocks one after the other multiple times. Our proposed TranS4mer outperforms all prior methods in three movie scene detection datasets, including MovieNet, BBC, and OVSD, while being 2× faster and requiring 3× less GPU memory than standard Transformer models. We will release our code and models.

Table of Contents

  • 1. Introduction
  • 2. Related works
  • 3. Background
  • 4. Technical Approach
  • 4.1. The TranS4mer Model
  • 4.2. Training and Loss Functions
  • 4.3. Implementation Details
  • 5. Experimental Setup
  • 5.1. Datasets
  • 5.2. Evaluation Metrics
  • 5.3. Baselines
  • 6. Results and Analysis
  • 6.1. Main Results on the MovieNet Dataset
  • 6.2. Scene Detection on Other Datasets
  • 6.3. Long-range Video Classification
  • 6.4. Ablation Studies
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — State-Space Transformer (TranS4mer) Architecture for Video Scene Detection

    model/method

    The State-Space Transformer (TranS4mer) is an end-to-end architecture designed for movie scene boundary detection. It processes long video sequences by decoupling temporal modeling into short-range intra-shot representations and long-range inter-shot context.

    Given an input video segment Vi∈RN×K×3×H×W\mathbf{V}_i \in \mathbb{R}^{N \times K \times 3 \times H \times W} consisting of N=2m+1N = 2m + 1 consecutive shots with KK sampled RGB frames of spatial dimensions H×WH \times W per shot, TranS4mer operates as follows:

    1. Patch Tokenization: Each frame is partitioned into P=HWp2P = \frac{HW}{p^2} non-overlapping spatial patches of size p×pp \times p. Patches are mapped through a trainable linear projection layer to a hidden dimension DD, and a learnable positional embedding is added, producing an input tensor Vi∈RN×K×P×D\mathbf{V}_i \in \mathbb{R}^{N \times K \times P \times D}.

    2. Backbone Feature Processing: The token sequence is processed through a stack of BB State-Space Self-Attention (S4A) blocks. Each S4A block consists of an intra-shot Multi-Head Attention (MHA) module applied independently across tokens within each individual shot, followed by an inter-shot Gated State-Space (GS4) module applied along the unrolled sequence of all shots.

    3. Output Heads: The sequence of class (extCLS ext{CLS}) tokens corresponding to all NN shots is extracted from the final S4A block and routed to two linear output heads: a boundary prediction head (which computes the probability that the central shot sis_i is a scene boundary) and a contrastive head (which maps shot representations into a metric space for shot-scene similarity learning).

  2. Knowl 2 — State-Space Self-Attention (S4A) Block

    model/method

    The State-Space Self-Attention (S4A) block is the core building block of the TranS4mer architecture. It combines Multi-Head Self-Attention (MHA) for local intra-shot modeling and Gated State-Space (GS4) operations for global inter-shot sequence modeling.

    Let Vi∈RN×K×P×D\mathbf{V}_i \in \mathbb{R}^{N \times K \times P \times D} denote the input token tensor, representing NN shots, KK frames per shot, PP patches per frame, and feature dimension DD. The block consists of two sequential modules:

    1. Intra-Shot Module: Processes the tokens of each shot independently. For a shot S=[x1,…,xL]S = [\mathbf{x}_1, \dots, \mathbf{x}_L] with length L=K×PL = K \times P and token vectors xj∈RD\mathbf{x}_j \in \mathbb{R}^D, the intra-shot updates with Layer Normalization (LN) and Multi-Layer Perceptron (MLP) are computed as: x′=MHA(LN(xin))+xin\mathbf{x}' = \text{MHA}(\text{LN}(\mathbf{x}_{in})) + \mathbf{x}_{in} xout=MLP(LN(x′))+x′\mathbf{x}_{out} = \text{MLP}(\text{LN}(\mathbf{x}')) + \mathbf{x}' This intra-shot attention scales with computational cost O(NL2)=O(NK2P2)\mathcal{O}(N L^2) = \mathcal{O}(N K^2 P^2), avoiding the quadratic O(N2L2)\mathcal{O}(N^2 L^2) cost of global self-attention across all shots.

    2. Inter-Shot Module: Takes the output tensor Vi′∈RN×K×P×D\mathbf{V}'_i \in \mathbb{R}^{N \times K \times P \times D} from the intra-shot module and flattens it into an extended sequence of length L′=N×K×PL' = N \times K \times P, denoted [z1,…,zL′][\mathbf{z}_1, \dots, \mathbf{z}_{L'}]. It applies a Gated S4 (GS4) layer followed by an MLP with residual connections: z′=GS4(LN(zin))+zin\mathbf{z}' = \text{GS4}(\text{LN}(\mathbf{z}_{in})) + \mathbf{z}_{in} zout=MLP(LN(z′))+z′\mathbf{z}_{out} = \text{MLP}(\text{LN}(\mathbf{z}')) + \mathbf{z}' Because the GS4 layer operates with linear computational and memory complexity with respect to L′L', the inter-shot module models long temporal dependencies across distant shots efficiently.

  3. Knowl 3 — Gated State-Space Sequence (GS4) Layer Formulation

    equation

    In the inter-shot module of the TranS4mer model, the Gated State-Space (GS4) layer combines continuous-time structured state-space modeling with a gating mechanism. Given an input token sequence x∈RL′×D\mathbf{x} \in \mathbb{R}^{L' \times D}:

    u=ϕ(Wux),v=ϕ(Wvx)\mathbf{u} = \phi(\mathbf{W}_u \mathbf{x}), \qquad \mathbf{v} = \phi(\mathbf{W}_v \mathbf{x}) h=S4(u),u^=Whx\mathbf{h} = \text{S4}(\mathbf{u}), \qquad \hat{\mathbf{u}} = \mathbf{W}_h \mathbf{x} y=Wy(u^⊙v)\mathbf{y} = \mathbf{W}_y (\hat{\mathbf{u}} \odot \mathbf{v})

    where:

    • Wu,Wv,Wh,Wy∈RD×D\mathbf{W}_u, \mathbf{W}_v, \mathbf{W}_h, \mathbf{W}_y \in \mathbb{R}^{D \times D} are learnable projection weight matrices.
    • ϕ(⋅)\phi(\cdot) represents the Gaussian Error Linear Unit (GELU) activation function.
    • S4(⋅)\text{S4}(\cdot) denotes a Structured State-Space sequence operator whose continuous-time state transition matrix A\mathbf{A} is structured (diagonal and low-rank) to enable convolution kernel computation in closed form.
    • ⊙\odot represents element-wise multiplication (Hadamard product).
    • y∈RL′×D\mathbf{y} \in \mathbb{R}^{L' \times D} is the contextualized output sequence.
  4. Knowl 4 — Problem Formulation of Movie Scene Boundary Detection

    definition

    Movie scene boundary detection is defined as a sequence classification task over a hierarchical video representation:

    1. A movie video V\mathcal{V} is partitioned into TT non-overlapping, contiguous video shots V={s1,s2,…,sT}\mathcal{V} = \{s_1, s_2, \dots, s_T\}, where each shot sis_i contains frames captured continuously by a single camera.
    2. Each shot sis_i is represented by KK uniformly sampled RGB frames: si={f1,…,fK}s_i = \{\mathbf{f}_1, \dots, \mathbf{f}_K\}, where each frame fj∈R3×H×W\mathbf{f}_j \in \mathbb{R}^{3 \times H \times W} has height HH and width WW.
    3. To classify a target shot sis_i, a context window of N=2m+1N = 2m + 1 adjacent shots centered at sis_i is extracted: Vi={si−m,…,si,…,si+m}\mathbf{V}_i = \{s_{i-m}, \dots, s_i, \dots, s_{i+m}\}.
    4. The objective is to predict a binary label yi∈{0,1}y_i \in \{0, 1\} for shot sis_i, where yi=1y_i = 1 indicates that the shot marks a movie scene boundary (a transition between semantically distinct story units) and yi=0y_i = 0 indicates it does not.
  5. Knowl 5 — Self-Supervised Pretraining with Shot-Scene Contrastive and Pseudo-Boundary Objectives

    model/method

    TranS4mer is pretrained in a self-supervised manner using pseudo-scene boundaries generated by the Dynamic Time Warping (DTW) algorithm on unlabeled movie videos.

    Given a context window Vi={si−m,…,si,…,si+m}\mathbf{V}_i = \{s_{i-m}, \dots, s_i, \dots, s_{i+m}\} and a DTW-identified pseudo-boundary shot si∗s_{i^*}, the sequence is split into two pseudo-scenes: VL={si−m,…,si∗}\mathcal{V}_L = \{s_{i-m}, \dots, s_{i^*}\} and VR={si∗+1,…,si+m}\mathcal{V}_R = \{s_{i^*+1}, \dots, s_{i+m}\}. Let r\mathbf{r} denote a shot representation obtained by applying a linear projection to the shot's CLS token, and let rˉL\bar{\mathbf{r}}_L and rˉR\bar{\mathbf{r}}_R denote the scene representations formed by mean-pooling the shot CLS tokens within VL\mathcal{V}_L and VR\mathcal{V}_R, respectively. The model is optimized using the sum of two objectives:

    1. Shot-Scene Contrastive Loss (LC\mathcal{L}_C): LC=lC(ri−m,rˉL)+lC(ri+m,rˉR)\mathcal{L}_C = l_C(\mathbf{r}_{i-m}, \bar{\mathbf{r}}_L) + l_C(\mathbf{r}_{i+m}, \bar{\mathbf{r}}_R) where for any shot embedding r\mathbf{r} and scene embedding rˉ\bar{\mathbf{r}}: lC(r,rˉ)=−log⁡S(r,rˉ)S(r,rˉ)+∑rnS(rn,rˉ)+∑rˉnS(r,rˉn)l_C(\mathbf{r}, \bar{\mathbf{r}}) = -\log \frac{\mathcal{S}(\mathbf{r}, \bar{\mathbf{r}})}{\mathcal{S}(\mathbf{r}, \bar{\mathbf{r}}) + \sum_{\mathbf{r}_n} \mathcal{S}(\mathbf{r}_n, \bar{\mathbf{r}}) + \sum_{\bar{\mathbf{r}}_n} \mathcal{S}(\mathbf{r}, \bar{\mathbf{r}}_n)} with cosine similarity S(x,y)=exp⁡(x⊤y∥x∥∥y∥)\mathcal{S}(\mathbf{x}, \mathbf{y}) = \exp\left(\frac{\mathbf{x}^\top \mathbf{y}}{\|\mathbf{x}\| \|\mathbf{y}\|}\right), and rn,rˉn\mathbf{r}_n, \bar{\mathbf{r}}_n denote negative shot and scene representations from the mini-batch.

    2. Pseudo-Boundary Loss (LB\mathcal{L}_B): LB=−log⁡(ρb(ri∗))−log⁡(1−ρb(rbˉ))\mathcal{L}_B = -\log \left( \rho_b(\mathbf{r}_{i^*}) \right) - \log \left( 1 - \rho_b(\mathbf{r}_{\bar{b}}) \right) where ρb(⋅)\rho_b(\cdot) is a linear classifier outputting the boundary probability, ri∗\mathbf{r}_{i^*} is the pseudo-boundary shot embedding, and rbˉ\mathbf{r}_{\bar{b}} is the embedding of a randomly selected non-boundary shot.

  6. Knowl 6 — Movie Scene Boundary Detection Performance on MovieNet

    data/table

    The performance of TranS4mer on the MovieNet test set was evaluated under both unsupervised and supervised settings. Metrics include Average Precision (AP), mean Intersection over Union (mIoU), Area Under the Receiver Operating Characteristic Curve (AUC-ROC), F1-Score, GPU memory consumption (in GB), and training throughput (in samples per second).

    Method AP (%) mIoU (%) AUC-ROC (%) F1 (%) Memory (GB) Samples/s
    Unsupervised
    GraphCut 14.10 29.70 - - - -
    SCSA 14.70 30.50 - - - -
    DP 15.50 32.00 - - - -
    Story Graph 25.10 35.70 - - - -
    Grouping 33.60 37.20 - - - -
    BaSSL 31.55 39.36 71.67 32.55 34.28 0.96
    Transformer 32.16 37.24 70.23 31.24 30.28 1.27
    TimeSformer 32.47 37.76 70.95 31.45 28.12 1.47
    Vanilla S4 33.34 38.12 71.86 32.21 15.62 1.83
    TranS4mer 34.45 39.60 73.25 33.41 10.13 2.57
    Supervised
    Siamese 35.80 39.60 - - - -
    MS-LSTM 46.50 46.20 - - - -
    LGSS 47.10 48.80 - - - -
    ViS4mer 55.13 48.27 88.74 46.15 - -
    ShotCoL 53.40 - - - 34.28 0.96
    BaSSL 57.40 50.69 90.54 47.02 34.28 0.96
    Transformer 58.81 51.21 90.84 47.88 30.28 1.27
    TimeSformer 59.62 50.75 90.66 48.02 28.12 1.47
    Vanilla S4 59.71 51.32 90.96 47.85 15.62 1.83
    TranS4mer 60.78 51.91 91.89 48.36 10.13 2.57

    In the supervised setting, TranS4mer achieves 60.78% AP, outperforming the prior state-of-the-art method (BaSSL) by +3.38% AP while consuming 3.38×3.38\times less GPU memory (10.13 GB vs. 34.28 GB) and training 2.68×2.68\times faster (2.57 vs. 0.96 samples/s). In the unsupervised setting, TranS4mer achieves the highest score across all metrics (34.45% AP).

  7. Knowl 7 — Movie Scene Detection Performance on BBC and OVSD Benchmarks

    data/table

    TranS4mer was evaluated on two additional scene boundary detection benchmarks: the BBC dataset (11 episodes of Planet Earth, 670 scenes, 4.9K shots) and the Open Video Scene Dataset (OVSD, 21 short films, 300 scenes, 10K shots). Average Precision (AP) results are compared against prior methods:

    BBC Dataset OVSD Dataset
    Method AP (%) Method AP (%)
    BaSSL 39.98 BaSSL 28.68
    Transformer 41.86 Transformer 33.12
    TimeSformer 42.23 TimeSformer 33.87
    Vanilla S4 42.56 Vanilla S4 34.21
    TranS4mer 43.64 TranS4mer 36.04

    TranS4mer achieves state-of-the-art results on both datasets, outperforming BaSSL by +4.66% AP on BBC and +7.36% AP on OVSD, demonstrating strong cross-dataset generalization.

  8. Knowl 8 — Ablation on S4A Architectural Components and S4 Variants

    data/table

    Ablation experiments conducted on MovieNet validate the design choices of the S4A block in TranS4mer, including module composition, layer placement, layer sparsity, and state-space formulation:

    (a) S4A Modules (b) S4 Layer Depths (c) S4 Sparsity (d) S4 Variants
    Configuration AP (%) Layers AP (%) Frequency AP (%) Variant AP (%)
    Intra-Shot Only 57.80 Layers 1–6 58.31 Every 2nd 59.82 Vanilla S4 59.71
    Inter-Shot Only 55.59 Layers 7–12 58.82 Every 4th 58.01 Diagonal S4 (DS4) 60.13
    Both (S4A) 60.78 Layers 1–12 60.78 All Layers 60.78 Gated S4 (GS4) 60.78

    Key observations:

    • Removing the inter-shot module reduces AP by 2.98%, whereas removing the intra-shot module reduces AP by 5.19%, confirming their complementary roles.
    • Placing S4 layers across all 12 blocks performs better than placing them only in early (1--6) or late (7--12) layers, or subsampling every 2nd or 4th layer.
    • Gated S4 (GS4) outperforms standard S4 (59.71%) and Diagonal S4 (60.13%), achieving the best performance at 60.78% AP.
  9. Knowl 9 — Computational Efficiency and Contextual Extent Scaling

    empirical result

    Experiments analyzing the effect of contextual window length (number of input shots) on performance and resource requirements reveal the scaling advantages of TranS4mer over standard Transformer baselines:

    • Temporal Extent Impact: On MovieNet, increasing the context window from 9 shots to 25 shots improves boundary detection accuracy from 53.29% AP to 60.78% AP (+7.49% AP). Expanding the window further to 33 shots leads to performance saturation without additional gains.
    • Accuracy under Long Sequences: While a standard Transformer performs comparably to or better than TranS4mer on very short sequences (e.g., 9 shots), TranS4mer outperforms the Transformer by +2.07% AP when scaling to 33 shots.
    • Speed and Memory Efficiency: As sequence length increases, the training speed of the standard Transformer degrades substantially while GPU memory consumption grows quadratically. At 33 input shots, TranS4mer achieves a 2.7×2.7\times higher training throughput and requires 3.38×3.38\times less GPU memory compared to the standard Transformer baseline.
  10. Knowl 10 — Transfer Performance on Long-Range Video Classification Benchmarks

    data/table

    To assess generalizability beyond scene detection, TranS4mer was evaluated on three long-form video understanding benchmarks: Long-form Video Understanding (LVU, 30K videos across 7 tasks), Breakfast Actions (1,712 videos of 10 cooking activities), and COIN (11,827 videos of 180 procedural tasks). Video clips average 2--3 minutes in length.

    (a) Top-1 Accuracy (%) on LVU Benchmark Tasks
    Method Relation Speak Scene Director Genre Writer Year
    ObjTrans. 53.10 39.40 56.90 51.20 54.60 34.50 39.10
    ViS4mer 57.14 40.79 67.44 62.61 54.71 48.80 44.75
    TranS4mer 59.52 39.21 70.93 63.86 55.85 46.93 45.45
    (b) Breakfast Activity Dataset (c) COIN Procedural Dataset
    Model Pretrain Samples Top-1 Acc. (%) Model Pretrain Samples Top-1 Acc. (%)
    GHRM 306K 75.50 TSN 306K 73.40
    Dist.Sup. 136M 89.90 Dist.Sup. 136M 90.00
    ViS4mer 495K 88.17 ViS4mer 495K 88.41
    TranS4mer 495K 90.27 TranS4mer 495K 89.23

    TranS4mer obtains the best performance in 5 of 7 tasks on LVU, achieves state-of-the-art accuracy on Breakfast (90.27%), and ranks second on COIN (89.23%) while using orders of magnitude less pretraining data (495K samples vs. 136M samples used by Distant Supervision).

Coverage note — Standard background derivations of generic continuous-time State-Space Models (SSMs) and the external Dynamic Time Warping (DTW) pseudo-boundary generation algorithm were omitted as non-contributed background material.

References

  1. 1.Dzmitry Bahdanau, Kyunghyun Cho, et al. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv: 1409.0473, 2014. 1, 2
  2. 2.Lorenzo Baraldi, Costantino Grana, and Rita Cucchiara. Analysis and re-use of videos in educational digital libraries with automatic scene detection. In Italian Research Conference on Digital Libraries, pages 155–164. Springer, 2015. 2
  3. 3.Lorenzo Baraldi, Costantino Grana, and Rita Cucchiara. A deep siamese network for scene detection in broadcast videos. In Proceedings of the 23rd ACM international conference on Multimedia, pages 1199–1202, 2015. 2, 5, 6, 7
  4. 4.Lorenzo Baraldi, Costantino Grana, and Rita Cucchiara. Shot and scene detection via hierarchical clustering for re-using broadcast video. In International conference on computer analysis of images and patterns, pages 801–811. Springer, 2015. 2
  5. 5.BBC. Planet earth. https://www.bbc.co.uk/programmes/b006mywy. 5
  6. 6.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In Proceedings of the International Conference on Machine Learning (ICML), July 2021. 6, 7
  7. 7.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. 1, 2
  8. 8.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020. 2
  9. 9.Brandon Castellano. Pyscenedetect: Intelligent scene cut detection and video splitting tool, 2018. 3
  10. 10.Vasileios T Chasanis, Aristidis C Likas, and Nikolaos P Galatsanos. Scene detection in videos using shot clustering and sequence alignment. IEEE transactions on multimedia, 11(1):89–100, 2008. 2, 6
  11. 11.Shixing Chen, Xiaohan Nie, David Fan, Dongqing Zhang, Vimal Bhat, and Raffay Hamid. Shot contrastive self-supervised learning for scene boundary detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9796–9805, 2021. 1, 2, 3, 5, 6
  12. 12.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. Rethinking attention with performers. arXiv preprint arXiv:2009.14794, 2020. 2
  13. 13.Costas Cotsaces, Nikos Nikolaidis, and Ioannis Pitas. Video shot detection and condensed representation. a review. IEEE signal processing magazine, 23(2):28–37, 2006. 3
  14. 14.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019. 1, 2
  15. 15.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 1, 2
  16. 16.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 2, 4, 5, 6, 7, 8
  17. 17.Karan Goel, Albert Gu, Chris Donahue, and Christopher Re.´ It’s raw! audio generation with state-space models. International Conference on Machine Learning (ICML), 2022. 2
  18. 18.Albert Gu, Karan Goel, and Christopher Re. Efficiently ´ modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396, 2021. 2, 3, 6, 7
  19. 19.Albert Gu, Isys Johnson, Karan Goel, Khaled Kamal Saab, Tri Dao, Atri Rudra, and Christopher Re. Combining recur- ´ rent, convolutional, and continuous-time models with linear state space layers. In Thirty-Fifth Conference on Neural Information Processing Systems, 2021. 2
  20. 20.Ankit Gupta. Diagonal state spaces are as effective as structured state spaces. arXiv preprint arXiv:2203.14343, 2022. 2, 8
  21. 21.Bo Han and Weiguo Wu. Video scene segmentation using a novel boundary evaluation criterion and dynamic programming. In 2011 IEEE International conference on multimedia and expo, pages 1–6. IEEE, 2011. 6
  22. 22.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 6
  23. 23.Dan Hendrycks and Kevin Gimpel. Bridging nonlinearities and stochastic regularizers with gaussian error linear units. CoRR, abs/1606.08415, 3, 2016. 3
  24. 24.Weizhe Hua, Zihang Dai, Hanxiao Liu, and Quoc Le. Transformer quality in linear time. In International Conference on Machine Learning, pages 9099–9117. PMLR, 2022. 3
  25. 25.Qingqiu Huang, Yu Xiong, Anyi Rao, Jiaze Wang, and Dahua Lin. Movienet: A holistic dataset for movie understanding. In European Conference on Computer Vision, pages 709–727. Springer, 2020. 2, 5, 6
  26. 26.Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Franc¸ois Fleuret. Transformers are rnns: Fast autoregressive transformers with linear attention. In International Conference on Machine Learning, pages 5156–5165. PMLR, 2020. 2
  27. 27.Ephraim Katz. The film encyclopedia. Thomas Y. Crowell, 1979. 3
  28. 28.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 5
  29. 29.Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. Reformer: The efficient transformer. arXiv preprint arXiv:2001.04451, 2020. 2
  30. 30.Hilde Kuehne, Ali Arslan, and Thomas Serre. The language of actions: Recovering the syntax and semantics of goal-directed human activities. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 780–787, 2014. 2, 7
  31. 31.Xudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach, Shih-Fu Chang, and Lorenzo Torresani. Learning to recognize procedural activities with distant supervision. arXiv preprint arXiv:2201.10990, 2022. 7
  32. 32.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019. 1, 2
  33. 33.Harsh Mehta, Ankit Gupta, Ashok Cutkosky, and Behnam Neyshabur. Long range language modeling via gated state spaces. arXiv preprint arXiv:2206.13947, 2022. 2, 3, 8
  34. 34.Md Mohaiminul Islam and Gedas Bertasius. Long movie clip classification with state-space video models. European Conference on Computer Vision (ECCV), 2022. 2, 6, 7
  35. 35.Jonghwan Mun, Minchul Shin, Gunsoo Han, Sangho Lee, Seongsu Ha, Joonseok Lee, and Eun-Sol Kim. Boundary-aware self-supervised learning for video scene segmentation. arXiv preprint arXiv:2201.05277, 2022. 1, 2, 3, 5, 6, 7
  36. 36.Mandela Patrick, Dylan Campbell, Yuki Asano, Ishan Misra, Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and Joao F Henriques. Keeping your eye on the ball: Tra- ˜ jectory attention in video transformers. Advances in neural information processing systems, 34:12493–12506, 2021. 2
  37. 37.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019. 1, 2
  38. 38.Anyi Rao, Linning Xu, Yu Xiong, Guodong Xu, Qingqiu Huang, Bolei Zhou, and Dahua Lin. A local-to-global approach to multi-modal movie scene segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10146–10155, 2020. 2, 3, 6
  39. 39.Zeeshan Rasheed and Mubarak Shah. Detection and representation of scenes in videos. IEEE transactions on Multimedia, 7(6):1097–1105, 2005. 2, 6
  40. 40.Daniel Rotman, Dror Porat, and Gal Ashour. Robust and efficient video scene detection using optimal sequential grouping. In 2016 IEEE international symposium on multimedia (ISM), pages 275–280. IEEE, 2016. 2, 5, 6, 7
  41. 41.Yong Rui, Thomas S Huang, and Sharad Mehrotra. Constructing table-of-content for videos. Multimedia systems, 7(5):359–368, 1999. 3
  42. 42.Panagiotis Sidiropoulos, Vasileios Mezaris, Ioannis Kompatsiaris, Hugo Meinedo, Miguel Bugalho, and Isabel Trancoso. Temporal video segmentation to scenes using high-level audiovisual features. IEEE Transactions on Circuits and Systems for Video Technology, 21(8):1163–1177, 2011. 2, 3
  43. 43.Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. Coin: A large-scale dataset for comprehensive instructional video analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1207–1216, 2019. 2, 7
  44. 44.Yansong Tang, Jiwen Lu, and Jie Zhou. Comprehensive instructional video analysis: The coin dataset and performance evaluation. IEEE transactions on pattern analysis and machine intelligence, 43(9):3138–3153, 2020. 7
  45. 45.Makarand Tapaswi, Martin Bauml, and Rainer Stiefelhagen. Storygraphs: visualizing character interactions as a timeline. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 827–834, 2014. 6
  46. 46.Paul-Louis Thirard and Lorenzo Codelli. Robert sklar. film, an international history of the medium, 1993;; kristin thompson, david bordwell. film history, an introduction, 1994. 1895, revue d’histoire du cinema ´ , 17(1):170–170, 1994. 2, 3
  47. 47.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 1, 2
  48. 48.Chao-Yuan Wu and Philipp Krahenbuhl. Towards long-form video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1884–1894, 2021. 2, 7
  49. 49.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32, 2019. 1, 2
  50. 50.Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. Big bird: Transformers for longer sequences. Advances in Neural Information Processing Systems, 33:17283–17297, 2020. 2
  51. 51.Jiaming Zhou, Kun-Yu Lin, Haoxin Li, and Wei-Shi Zheng. Graph-based high-order relation modeling for long-term action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8984–8993, 2021. 7
  52. 52.Xiang Sean Zhou, Yong Rui, and Thomas S Huang. Constructing table-of-content for videos. In Exploration of Visual Data, pages 53–73. Springer, 2003. 2

Citation

MLA
Islam, M. M., et al. “Efficient Movie Scene Detection Using State-Space Transformers”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 18749–58, https://doi.org/10.1109/CVPR52729.2023.01798.
APA
Islam, M. M., Hasan, M., Athrey, K. S., Braskich, T., & Bertasius, G. (2023). Efficient Movie Scene Detection using State-Space Transformers. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18749–18758. https://doi.org/10.1109/CVPR52729.2023.01798
Chicago
Islam, M. M., M. Hasan, K. S. Athrey, T. Braskich, and G. Bertasius. 2023. “Efficient Movie Scene Detection Using State-Space Transformers”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18749–58. https://doi.org/10.1109/CVPR52729.2023.01798.
Harvard
Islam, M.M. et al. (2023) “Efficient Movie Scene Detection using State-Space Transformers”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 18749–18758. Available at: https://doi.org/10.1109/CVPR52729.2023.01798.
Vancouver
1. Islam MM, Hasan M, Athrey KS, Braskich T, Bertasius G (2023) Efficient Movie Scene Detection using State-Space Transformers. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 18749–18758

BibTeX

@inproceedings{Islam_2023, title={Efficient Movie Scene Detection using State-Space Transformers}, url={http://dx.doi.org/10.1109/CVPR52729.2023.01798}, DOI={10.1109/cvpr52729.2023.01798}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Islam, Md Mohaiminul and Hasan, Mahmudul and Athrey, Kishan Shamsundar and Braskich, Tony and Bertasius, Gedas}, year={2023}, month=June, pages={18749–18758} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/