Dense-Captioning Events in Videos

Ranjay KrishnaKenji HataFrederic RenLi Fei-FeiJuan Carlos Niebles

article2017ICCV1,640 citations

Introduces the task and benchmark of dense video captioning alongside a context-aware model that simultaneously localizes and describes multiple temporal events across untrimmed videos in a single pass.

Listen

Modern automated video analysis systems frequently struggle with capturing the rich, multi-layered actions found in everyday, unconstrained videos. While traditional computer vision models assign broad, discrete action categories or describe short clips with a single sentence, natural videos typically contain multiple distinct, overlapping, and temporally interdependent events spanning seconds to several minutes. Without the capability to pinpoint exactly when events happen and describe them in detail, video-driven applications face significant limitations in automated search, content cataloging, and intelligent surveillance.

To address this limitation, the article introduces the unified task of dense-captioning events in videos, aiming to simultaneously detect when specific events occur and describe each event in natural language. The primary objective is to demonstrate an integrated computational architecture that performs event proposal and natural-language captioning across both short and long video sequences in a single processing pass, while incorporating contextual dependencies between past, concurrent, and future events.

To evaluate this framework, the authors created and deployed the ActivityNet Captions benchmark, comprising 20,000 untrimmed open-domain videos totaling 849 hours and annotated with 100,000 temporally localized descriptions. The technical approach couples a multi-scale temporal proposal module—sampling video frame features across multiple time strides—with a language generation module. This language network uses an attention mechanism to pool contextual hidden representations from preceding and succeeding events, enabling it to describe individual occurrences while maintaining narrative consistency across the entire video timeline.

The findings confirm that integrating temporal context substantially improves caption quality and event localization. When generating captions on predicted events, the full context-aware model achieved a CIDEr score of 17.29, representing an approximate 40% relative improvement over the baseline model lacking context (12.34). Incorporating multi-stride sampling resolved gradient decay issues in long video analysis and consistently outperformed single-stride event detection, particularly as the number of proposed events scaled. Furthermore, the architecture improved related tasks: in video retrieval benchmarks, adding context increased top-50 retrieval recall from 0.32 to 0.65 and cut the median retrieval rank by more than half, moving from rank 78 to 34.

These results indicate that automated video understanding benefits considerably from moving beyond isolated frame classification toward unified, context-aware sequence modeling. For organizations managing massive digital media archives, video surveillance feeds, or streaming platforms, dense-captioning frameworks offer a clear pathway to high-precision indexing and queryable video databases without requiring expensive human re-annotation. In addition, an online variant of the model that conditions solely on past events showed strong performance, demonstrating the technical feasibility of deploying dense captioning directly within real-time streaming pipelines.

Organizations developing automated video processing pipelines should prioritize architectures that capture temporal context and support streaming event detection. Future engineering work should concentrate on refining event boundary precision, as highly overlapping or redundant event proposals currently risk generating repetitive text descriptions. Confidence in the core findings remains strong given the extensive 20,000-video empirical benchmark, though performance may vary when applying the system to specialized, fine-grained activity domains that diverge from open-domain internet video datasets.

Cover for Dense-Captioning Events in Videos

Abstract

Most natural videos contain numerous events. For example, in a video of a "man playing a piano", the video might also contain "another man dancing" or "a crowd clapping". We introduce the task of dense-captioning events, which involves both detecting and describing events in a video. We propose a new model that is able to identify all events in a single pass of the video while simultaneously describing the detected events with natural language. Our model introduces a variant of an existing proposal module that is designed to capture both short as well as long events that span minutes. To capture the dependencies between the events in a video, our model introduces a new captioning module that uses contextual information from past and future events to jointly describe all events. We also introduce ActivityNet Captions, a large-scale benchmark for dense-captioning events. ActivityNet Captions contains 20k videos amounting to 849 video hours with 100k total descriptions, each with it's unique start and end time. Finally, we report performances of our model for dense-captioning events, video retrieval and localization.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Dense-captioning events model
  • 3.1 Event proposal module
  • 3.2 Captioning module with context
  • 3.3 Implementation details.
  • 4 ActivityNet Captions dataset
  • 4.1 Dataset statistics
  • 4.2 Temporal agreement amongst annotators
  • 5 Experiments
  • 5.1 Dense-captioning events
  • 5.2 Event localization
  • 5.3 Video and paragraph retrieval
  • 6 Conclusion
  • 7 Supplementary material
  • 7.1 Comparison to other datasets.
  • 7.2 Detailed dataset statistics
  • 7.3 Dataset collection process
  • 7.4 Annotation details
  • References

Knowls

  1. Knowl 1 — Dense-Captioning Events Task Formulation

    definition

    The task of dense-captioning events in videos requires a model to simultaneously detect all events occurring in an untrimmed video and generate a natural language description for each detected event.

    Formally, given an input video sequence v={vt}t=0T−1v = \{v_t\}_{t=0}^{T-1} of TT ordered frames, the system outputs a set of localized sentence descriptions S={si}i=1MS = \{s_i\}_{i=1}^M. Each description is defined as a tuple:

    si=(tistart,tiend,{wi,j}j=1Li)s_i = \left(t_i^{\text{start}}, t_i^{\text{end}}, \{w_{i,j}\}_{j=1}^{L_i}\right)

    where tistart,tiend∈[0,T−1]t_i^{\text{start}}, t_i^{\text{end}} \in [0, T-1] denote the temporal start and end bounds of the event (tistart<tiendt_i^{\text{start}} < t_i^{\text{end}}), and {wi,j}j=1Li\{w_{i,j}\}_{j=1}^{L_i} is a variable-length sequence of words describing the event, with each word wi,jw_{i,j} drawn from a vocabulary VV. Unlike standard single-sentence video captioning, the events can span widely varying time durations, can occur sequentially or concurrently with temporal overlap, and require joint temporal boundary estimation and text generation.

  2. Knowl 2 — ActivityNet Captions Benchmark

    definition

    ActivityNet Captions is a large-scale open-domain benchmark for dense-captioning events, built on top of the 20k untrimmed videos from the ActivityNet dataset. It contains 849 total hours of video paired with 100k temporally localized sentence descriptions.

    Key statistics of the benchmark include:

    • Video Duration: Videos average 180 seconds in length, with the longest videos exceeding 10 minutes.
    • Sentence Distribution: Each video contains an average of 3.65±1.793.65 \pm 1.79 temporally localized sentences (normally distributed), with each additional minute of video adding approximately 1 sentence description.
    • Sentence Length: Sentences contain an average of 13.48±6.3313.48 \pm 6.33 words, and the aggregate paragraph for a video averages 40±2640 \pm 26 words.
    • Temporal Coverage & Overlap: Each sentence describes an average of 36 seconds (31% of the video duration). The entire annotated paragraph covers on average 94.6% of the video duration. Approximately 10% of temporal event intervals overlap with one another.
    • Linguistic Distribution: The dataset exhibits an action-centric distribution with a high proportion of verbs and coreference pronouns (e.g., he, she) relative to dense image captioning datasets.
    • Annotator Agreement: Independent paragraphs collected from distinct crowd workers exhibit an average temporal Intersection over Union (tIoU) of 70.2% with maximal overlapping sentence combinations.
  3. Knowl 3 — Multi-Scale Single-Pass Event Proposal Module

    model/method

    The event proposal module detects events of diverse temporal lengths in a single forward pass without running sliding windows over the video.

    1. Feature Extraction: Video frames are partitioned into non-overlapping clips of δ=16\delta = 16 frames. A 3D convolutional neural network (C3D) extracts a sequence of D=500D = 500 dimensional semantic features:

    ft=F(vt:vt+δ)∈R500f_t = F(v_t : v_{t+\delta}) \in \mathbb{R}^{500}

    1. Multi-Stride Temporal Recurrence: To capture both short and long events (up to multiple minutes), the sequence of features is sampled at multiple strides (strides of 1, 2, 4, and 8) in parallel and fed into a proposal Long Short-Term Memory (LSTM) network. The LSTM accumulates temporal evidence across time.

    2. Proposal Generation: At each recurrent time step, the model outputs KK candidate proposals with temporal boundary offsets and confidence scores:

    P={(tistart,tiend,scorei,hi)}P = \left\{\left(t_i^{\text{start}}, t_i^{\text{end}}, \text{score}_i, h_i\right)\right\}

    where hih_i is the hidden state vector of the proposal LSTM at the detection step, serving as the visual feature representation of proposed event ii. Overlapping proposals are preserved rather than suppressed via non-maximum suppression to allow the detection of simultaneous events.

  4. Knowl 4 — Temporal Context-Aware Attention Captioning Module

    model/method

    To capture dependencies across related events in a video, the captioning module attends over past and future event representations when generating the description for a target event.

    For a candidate event ii with proposal LSTM hidden representation hih_i and time boundaries [tistart,tiend][t_i^{\text{start}}, t_i^{\text{end}}], all other proposed events j≠ij \neq i are partitioned into two context sets based on temporal finish times: past events (tjend<tiendt_j^{\text{end}} < t_i^{\text{end}}) and future/concurrent events (tjend≥tiendt_j^{\text{end}} \ge t_i^{\text{end}}).

    An attention weight wjw_j measuring the relevance of event jj to event ii is calculated using a learned attention vector aia_i:

    ai=wahi+baa_i = w_a h_i + b_a

    wj=ai⊤hjw_j = a_i^\top h_j

    where waw_a and bab_a are learned parameters. The past context representation hipasth_i^{\text{past}} and future context representation hifutureh_i^{\text{future}} are computed via normalized weighted sums:

    hipast=1Zpast∑j≠i1[tjend<tiend]wjhj,Zpast=∑j≠i1[tjend<tiend]h_i^{\text{past}} = \frac{1}{Z^{\text{past}}} \sum_{j \neq i} \mathbf{1}[t_j^{\text{end}} < t_i^{\text{end}}] w_j h_j, \quad Z^{\text{past}} = \sum_{j \neq i} \mathbf{1}[t_j^{\text{end}} < t_i^{\text{end}}]

    hifuture=1Zfuture∑j≠i1[tjend≥tiend]wjhj,Zfuture=∑j≠i1[tjend≥tiend]h_i^{\text{future}} = \frac{1}{Z^{\text{future}}} \sum_{j \neq i} \mathbf{1}[t_j^{\text{end}} \ge t_i^{\text{end}}] w_j h_j, \quad Z^{\text{future}} = \sum_{j \neq i} \mathbf{1}[t_j^{\text{end}} \ge t_i^{\text{end}}]

    The concatenated vector [hipast;hi;hifuture][h_i^{\text{past}}; h_i; h_i^{\text{future}}] is passed as input to a 2-layer language LSTM (512 hidden units per layer) to generate the sequence of words for event ii. In an online streaming configuration, hifutureh_i^{\text{future}} is omitted.

  5. Knowl 5 — Multi-Task Loss and Alternating Optimization for Dense Captioning

    model/method

    The dense event captioning model is trained end-to-end using a joint multi-task loss combining proposal confidence estimation and token generation:

    L=λ1Lcap+λ2Lprop\mathcal{L} = \lambda_1 \mathcal{L}_{\text{cap}} + \lambda_2 \mathcal{L}_{\text{prop}}

    where λ1=1.0\lambda_1 = 1.0 and λ2=0.1\lambda_2 = 0.1.

    • Lprop\mathcal{L}_{\text{prop}} is a weighted binary cross-entropy loss evaluated over proposal confidence scores for varying proposal lengths.
    • Lcap\mathcal{L}_{\text{cap}} is the standard cross-entropy loss over ground-truth words for event proposals exhibiting high Intersection over Union (IoU) with ground-truth segments, normalized by batch size and sentence length.

    Optimization Protocol:

    • Training alternates every 500 iterations between optimizing the proposal module and the language module.
    • The captioning module is initially trained for 10 epochs with neighboring context masked out before enabling context feature inputs.
    • Optimization uses Stochastic Gradient Descent (SGD) with momentum 0.9, a batch size of 1, and initial learning rates of 1×10−21 \times 10^{-2} for the language model and 1×10−31 \times 10^{-3} for the proposal module.
    • Maximum sentence length is capped at 30 words; sentences are generated during inference using beam search with a beam size of 5.
  6. Knowl 6 — Evaluation Metric for Dense-Captioning Events

    experimental setup

    The joint localization and description ability of dense video captioning models is evaluated using average precision across multiple temporal Intersection over Union (tIoU) thresholds.

    For an input video, the model outputs temporal proposals ranked by confidence. The top 1,000 proposals are matched against ground-truth temporal intervals at tIoU thresholds of 0.3, 0.5, and 0.7. For each threshold, standard natural language generation precision metrics—Bleu@1 (B@1B@1), Bleu@2 (B@2B@2), Bleu@3 (B@3B@3), Bleu@4 (B@4B@4), METEOR (MM), and CIDEr (CC)—are computed over the matched predictions and averaged.

    To evaluate language generation independently of proposal quality, models are also evaluated in an oracle setting using ground-truth event segment boundaries.

  7. Knowl 7 — Dense-Captioning Events Performance on ActivityNet Captions

    data/table

    The table below compares the performance of baseline methods and variants of the context-aware model on ActivityNet Captions. Performance is evaluated under two settings: using ground-truth (GT) proposals to evaluate language generation in isolation, and using proposals predicted by the multi-scale proposal module (learned proposals) for end-to-end dense captioning.

    With GT proposals With learned proposals
    Model B@1 B@2 B@3 B@4 M C B@1 B@2 B@3 B@4 M C
    LSTM-YT 18.22 7.43 3.24 1.24 6.56 14.86 - - - - - -
    S2VT 20.35 8.99 4.60 2.62 7.85 20.97 - - - - - -
    H-RNN 19.46 8.78 4.34 2.53 8.02 20.18 - - - - - -
    no context (ours) 20.35 8.99 4.60 2.62 7.85 20.97 12.23 3.48 2.10 0.88 3.76 12.34
    online−-attn (ours) 21.92 9.88 5.21 3.06 8.50 22.19 15.20 5.43 2.52 1.34 4.18 14.20
    online (ours) 22.10 10.02 5.66 3.10 8.88 22.94 17.10 7.34 3.23 1.89 4.38 15.30
    full−-attn (ours) 26.34 13.12 6.78 3.87 9.36 24.24 15.43 5.63 2.74 1.72 4.42 15.29
    full (ours) 26.45 13.48 7.12 3.98 9.46 24.56 17.95 7.69 3.86 2.20 4.82 17.29
    • Baselines that global-mean-pool video features (LSTM-YT) perform poorly over long video sequences. Hierarchical RNNs (H-RNN) improve only slightly as they focus on object-level features.
    • Incorporating past context (online models) improves CIDEr over no-context from 20.97 to 22.94 (GT proposals) and from 12.34 to 15.30 (learned proposals).
    • Incorporating bidirectional context (full model) yields the best scores across all metrics (CIDEr 24.56 with GT proposals, 17.29 with learned proposals).
    • Learned attention weighting outperforms unweighted mean pooling (−attn-\text{attn} variants), particularly in the learned proposal regime where filtering irrelevant proposals is necessary.
  8. Knowl 8 — Video and Paragraph Retrieval Benchmark Performance

    data/table

    The contextual representations learned by the event captioning pipeline were evaluated on video retrieval (finding the video matching a given paragraph) and paragraph retrieval (finding the paragraph matching a given video) using a max-margin ranking loss.

    Video retrieval Paragraph retrieval
    Model R@1 R@5 R@50 Med. rank R@1 R@5 R@50 Med. rank
    LSTM-YT 0.00 0.04 0.24 102 0.00 0.07 0.38 98
    no context 0.05 0.14 0.32 78 0.07 0.18 0.45 56
    online (ours) 0.10 0.32 0.60 36 0.17 0.34 0.70 33
    full (ours) 0.14 0.32 0.65 34 0.18 0.36 0.74 32
    • Incorporating temporal event context substantially reduces the median retrieval rank from 78 (no context) to 34 (full model) for video retrieval, and from 56 to 32 for paragraph retrieval.
    • The online context model performs nearly on par with the full bidirectional model in retrieval tasks (median rank 36 vs 34 in video retrieval, 33 vs 32 in paragraph retrieval).
  9. Knowl 9 — Impact of Context on Sequential Event Captioning

    empirical result

    Evaluating caption generation by sentence order within a video reveals distinct effects of unidirectional versus bidirectional context:

    • First Event Captioning: In the no-context baseline, the 1st sentence achieves B@4=4.51B@4 = 4.51, M=9.34M = 9.34, C=31.56C = 31.56. In the online model (which has no prior context for the first event), performance is B@4=4.77B@4 = 4.77, M=8.10M = 8.10, C=30.92C = 30.92. In the full model (which incorporates future context), B@4B@4 increases to 5.525.52 and MM increases to 10.0310.03 (C=29.92C = 29.92).
    • Subsequent Event Captioning: Improvements from past context primarily manifest in subsequent sentences. For the 2nd sentence, CIDEr increases from 19.3719.37 (no context) to 20.1720.17 (full context) and B@4B@4 increases from 1.871.87 to 2.332.33. For the 3rd sentence, CIDEr increases from 19.3619.36 (no context) to 20.0120.01 (full context).
    • Overall, incorporating future event information allows past events to reference downstream actions (e.g., mentioning mixing bowls or ingredients before they are combined later in the video).
  10. Knowl 10 — Caption Duplication on Highly Overlapping Temporal Proposals

    limitation

    When the multi-scale proposal module predicts multiple candidate events that have a high degree of temporal overlap, the contextual captioning model often struggles to discriminate between the subtle visual differences of the overlapping segments. Consequently, the language model can generate identical or near-duplicate sentences for distinct proposed intervals. In addition, when applied to videos containing rare or uncharacteristic event progressions, temporal context from surrounding events can introduce noisy priors that mislead the captioning network.

Coverage note — None was omitted; all key contributions including task definition, dataset statistics, multi-scale proposal and context-attention architectures, loss formulations, training scheme, evaluation protocol, dense captioning and retrieval results, contextual ablation by sentence position, and model limitations are covered.

References

  1. 1.A. Alahi, K. Goel, V. Ramanathan, A. Robicquet, L. Fei-Fei, and S. Savarese. Social lstm: Human trajectory prediction in crowded spaces. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 961–971, 2016.
  2. 2.C. Baldassano, J. Chen, A. Zadbood, J. W. Pillow, U. Hasson, and K. A. Norman. Discovering event structure in continuous narrative perception and memory. bioRxiv, page 081018, 2016.
  3. 3.O. Boiman and M. Irani. Detecting irregularities in images and in video. International journal of computer vision, 74(1):17–31, 2007.
  4. 4.F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. C. Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 961–970, 2015.
  5. 5.F. Caba Heilbron, J. C. Niebles, and B. Ghanem. Fast temporal activity proposals for efficient detection of human actions in untrimmed videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1914–1923, 2016.
  6. 6.D. L. Chen and W. B. Dolan. Collecting highly parallel data for paraphrase evaluation. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics (ACL-2011), Portland, OR, June 2011.
  7. 7.P. Das, C. Xu, R. F. Doell, and J. J.. Corso. A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2013.
  8. 8.J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell. Long-term recurrent convolutional networks for visual recognition and description. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2625–2634, 2015.
  9. 9.O. Duchenne, I. Laptev, J. Sivic, F. Bach, and J. Ponce. Automatic annotation of human actions in video. In Computer Vision, 2009 IEEE 12th International Conference on, pages 1491–1498. IEEE, 2009.
  10. 10.V. Escorcia, F. C. Heilbron, J. C. Niebles, and B. Ghanem. Daps: Deep action proposals for action understanding. In European Conference on Computer Vision, pages 768–784. Springer, 2016.
  11. 11.A. Gaidon, Z. Harchaoui, and C. Schmid. Temporal localization of actions with actoms. IEEE transactions on pattern analysis and machine intelligence, 35(11):2782–2795, 2013.
  12. 12.A. Geiger, P. Lenz, C. Stiller, and R. Urtasun. Vision meets robotics: The kitti dataset. International Journal of Robotics Research (IJRR), 2013.
  13. 13.G. Gkioxari and J. Malik. Finding action tubes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 759–768, 2015.
  14. 14.D. B. Goldman, B. Curless, D. Salesin, and S. M. Seitz. Schematic storyboarding for video visualization and editing. In ACM Transactions on Graphics (TOG), volume 25, pages 862–871. ACM, 2006.
  15. 15.A. Gorban, H. Idrees, Y.-G. Jiang, A. Roshan Zamir, I. Laptev, M. Shah, and R. Sukthankar. THUMOS challenge: Action recognition with a large number of classes. http://www.thumos.info/, 2015.
  16. 16.M. Gygli, H. Grabner, and L. Van Gool. Video summarization by learning submodular mixtures of objectives. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3090–3098, 2015.
  17. 17.S. Ji, W. Xu, M. Yang, and K. Yu. 3d convolutional neural networks for human action recognition. IEEE transactions on pattern analysis and machine intelligence, 35(1):221–231, 2013.
  18. 18.J. Johnson, A. Karpathy, and L. Fei-Fei. Densecap: Fully convolutional localization networks for dense captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4565–4574, 2016.
  19. 19.S. Karaman, L. Seidenari, and A. Del Bimbo. Fast saliency based pooling of fisher encoded dense trajectories. In ECCV THUMOS Workshop, volume 1, 2014.
  20. 20.A. Karpathy and L. Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3128–3137, 2015.
  21. 21.A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei. Large-scale video classification with convolutional neural networks. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 1725–1732, 2014.
  22. 22.J. Krause, J. Johnson, R. Krishna, and L. Fei-Fei. A hierarchical approach for generating descriptive image paragraphs. In Computer Vision and Patterm Recognition (CVPR), 2017.
  23. 23.R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei. Visual genome: Connecting language and vision using crowdsourced dense image annotations. In International Journal on Computer Vision (IJCV), 2017.
  24. 24.R. A. Krishna, K. Hata, S. Chen, J. Kravitz, D. A. Shamma, L. Fei-Fei, and M. S. Bernstein. Embracing error to enable rapid crowdsourcing. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems, pages 3167–3179. ACM, 2016.
  25. 25.H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre. Hmdb: a large video database for human motion recognition. In Computer Vision (ICCV), 2011 IEEE International Conference on, pages 2556–2563. IEEE, 2011.
  26. 26.I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld. Learning realistic human actions from movies. In Computer Vision and Pattern Recognition, 2008. CVPR 2008. IEEE Conference on, pages 1–8. IEEE, 2008.
  27. 27.W. Liu, T. Mei, Y. Zhang, C. Che, and J. Luo. Multi-task deep visual-semantic embedding for video thumbnail selection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3707–3715, 2015.
  28. 28.M. Marszałek, I. Laptev, and C. Schmid. Actions in context. In IEEE Conference on Computer Vision & Pattern Recognition, 2009.
  29. 29.T. Mikolov, M. Karafiát, L. Burget, J. Cernocký, and S. Khudanpur. Recurrent neural network based language model. In Interspeech, volume 2, page 3, 2010.
  30. 30.B. Ni, V. R. Paramathayalan, and P. Moulin. Multiple granularity analysis for fine-grained action detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 756–763, 2014.
  31. 31.J. C. Niebles, C.-W. Chen, and L. Fei-Fei. Modeling temporal structure of decomposable motion segments for activity classification. In European conference on computer vision, pages 392–405. Springer, 2010.
  32. 32.D. Oneata, J. Verbeek, and C. Schmid. Efficient action localization with approximately normalized fisher vectors. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2545–2552, 2014.
  33. 33.M. Otani, Y. Nakashima, E. Rahtu, J. Heikkilä, and N. Yokoya. Learning joint representations of videos and sentences with web image search. In European Conference on Computer Vision, pages 651–667. Springer, 2016.
  34. 34.Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui. Jointly modeling embedding and translation to bridge video and language. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4594–4602, 2016.
  35. 35.H. Pirsiavash and D. Ramanan. Parsing videos of actions with segmental grammars. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 612–619, 2014.
  36. 36.M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal. Grounding action descriptions in videos. Transactions of the Association for Computational Linguistics (TACL), 1:25–36, 2013.
  37. 37.A. Rohrbach, M. Rohrbach, W. Qiu, A. Friedrich, M. Pinkal, and B. Schiele. Coherent multi-sentence video description with variable level of detail. In German Conference on Pattern Recognition, pages 184–195. Springer, 2014.
  38. 38.A. Rohrbach, M. Rohrbach, and B. Schiele. The long-short story of movie description. In German Conference on Pattern Recognition, pages 209–221. Springer, 2015.
  39. 39.A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele. A dataset for movie description. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
  40. 40.M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele. A database for fine grained activity detection of cooking activities. In Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on, pages 1194–1201. IEEE, 2012.
  41. 41.N. Salehi, L. C. Irani, M. S. Bernstein, A. Alkhatib, E. Ogbe, K. Milland, et al. We are dynamo: Overcoming stalling and friction in collective action for crowd workers. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems, pages 1621–1630. ACM, 2015.
  42. 42.C. Schuldt, I. Laptev, and B. Caputo. Recognizing human actions: A local svm approach. In Pattern Recognition, 2004. ICPR 2004. Proceedings of the 17th International Conference on, volume 3, pages 32–36. IEEE, 2004.
  43. 43.G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta. Hollywood in homes: Crowdsourcing data collection for activity understanding. In European Conference on Computer Vision, 2016.
  44. 44.Y. Song, J. Vallmitjana, A. Stent, and A. Jaimes. Tvsum: Summarizing web videos using titles. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5179–5187, 2015.
  45. 45.K. Soomro, A. R. Zamir, and M. Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012.
  46. 46.Y. Tian, R. Sukthankar, and M. Shah. Spatiotemporal deformable part models for action detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2642–2649, 2013.
  47. 47.A. Torabi, C. Pal, H. Larochelle, and A. Courville. Using descriptive video services to create a large data source for video annotation research. arXiv preprint arXiv:1503.01070, 2015.
  48. 48.A. Vahdat, B. Gao, M. Ranjbar, and G. Mori. A discriminative key pose sequence model for recognizing human interactions. In Computer Vision Workshops (ICCV Workshops), 2011 IEEE International Conference on, pages 1729–1736. IEEE, 2011.
  49. 49.S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko. Sequence to sequence-video to text. In Proceedings of the IEEE International Conference on Computer Vision, pages 4534–4542, 2015.
  50. 50.S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko. Translating videos to natural language using deep recurrent neural networks. arXiv preprint arXiv:1412.4729, 2014.
  51. 51.L. Wang, Y. Qiao, and X. Tang. Action recognition and detection by combining motion and appearance features. THUMOS14 Action Recognition Challenge, 1:2, 2014.
  52. 52.L. Wang, Y. Qiao, and X. Tang. Video action detection with relational dynamic-poselets. In European Conference on Computer Vision, pages 565–580. Springer, 2014.
  53. 53.W. Wolf. Key frame selection by motion analysis. In Acoustics, Speech, and Signal Processing, 1996. ICASSP-96. Conference Proceedings., 1996 IEEE International Conference on, volume 2, pages 1228–1231. IEEE, 1996.
  54. 54.H. Xu, S. Venugopalan, V. Ramanishka, M. Rohrbach, and K. Saenko. A multi-scale multiple instance video description network. arXiv preprint arXiv:1505.05914, 2015.
  55. 55.J. Xu, T. Mei, T. Yao, and Y. Rui. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5288–5296, 2016.
  56. 56.K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio. Show, attend and tell: Neural image caption generation with visual attention. In ICML, volume 14, pages 77–81, 2015.
  57. 57.R. Xu, C. Xiong, W. Chen, and J. J. Corso. Jointly modeling deep video and compositional text to bridge vision and language in a unified framework. In AAAI, volume 5, page 6, 2015.
  58. 58.J. Yamato, J. Ohya, and K. Ishii. Recognizing human action in time-sequential images using hidden markov model. In Computer Vision and Pattern Recognition, 1992. Proceedings CVPR’92., 1992 IEEE Computer Society Conference on, pages 379–385. IEEE, 1992.
  59. 59.H. Yang, B. Wang, S. Lin, D. Wipf, M. Guo, and B. Guo. Unsupervised extraction of video highlights via robust recurrent auto-encoders. In Proceedings of the IEEE International Conference on Computer Vision, pages 4633–4641, 2015.
  60. 60.L. Yang, K. Tang, J. Yang, and L.-J. Li. Dense captioning with joint inference and visual context. arXiv preprint arXiv:1611.06949, 2016.
  61. 61.L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville. Describing videos by exploiting temporal structure. In Proceedings of the IEEE international conference on computer vision, pages 4507–4515, 2015.
  62. 62.T. Yao, T. Mei, and Y. Rui. Highlight detection with pairwise deep ranking for first-person video summarization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 982–990, 2016.
  63. 63.S. Yeung, A. Fathi, and L. Fei-Fei. Videoset: Video summary evaluation through text. arXiv preprint arXiv:1406.5824, 2014.
  64. 64.H. Yu, J. Wang, Z. Huang, Y. Yang, and W. Xu. Video paragraph captioning using hierarchical recurrent neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4584–4593, 2016.
  65. 65.H. J. Zhang, J. Wu, D. Zhong, and S. W. Smoliar. An integrated system for content-based video retrieval and browsing. Pattern recognition, 30(4):643–658, 1997.

Citation

MLA
Krishna, R., et al. “Dense-Captioning Events in Videos”. arXiv, 2017, http://arxiv.org/abs/1705.00754v1.
APA
Krishna, R., Hata, K., Ren, F., Fei-Fei, L., & Niebles, J. C. (2017). Dense-Captioning Events in Videos. arXiv. http://arxiv.org/abs/1705.00754v1
Chicago
Krishna, R., K. Hata, F. Ren, L. Fei-Fei, and J. C. Niebles. 2017. “Dense-Captioning Events in Videos”. arXiv. http://arxiv.org/abs/1705.00754v1.
Harvard
Krishna, R. et al. (2017) “Dense-Captioning Events in Videos”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1705.00754v1.
Vancouver
1. Krishna R, Hata K, Ren F, Fei-Fei L, Niebles JC (2017) Dense-Captioning Events in Videos. arXiv

BibTeX

@article{krishna2017dense,
  title = {Dense-Captioning Events in Videos},
  author = {Krishna, Ranjay and Hata, Kenji and Ren, Frederic and Fei-Fei, Li and Niebles, Juan Carlos},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1705.00754v1},
  eprint = {1705.00754}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE