A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation

Federico PerazziJordi Pont-TusetBrian McWilliamsLuc Van GoolM. GrossAlexander Sorkine‐Hornung

article2016CVPR2,415 citations

Introduces the densely annotated DAVIS video-segmentation benchmark and complementary spatial, contour, and temporal metrics that expose the strengths and weaknesses of current methods.

Listen

Separating moving foreground objects from background video at pixel-level accuracy is essential for visual effects, video editing, tracking, and action recognition. However, progress in video object segmentation has lagged behind still-image segmentation, primarily due to the lack of high-resolution, standardized, and densely annotated benchmarks. Existing datasets are constrained by low spatial resolution, sparse frame annotations, or limited sequence diversity, leading to algorithm overfitting and saturated performance on outdated metrics.

To address this bottleneck, the article introduces DAVIS (Densely Annotated VIdeo Segmentation), a high-definition benchmark dataset and evaluation methodology. The article set out to rigorously benchmark and analyze the performance, strengths, and failure modes of state-of-the-art video object segmentation algorithms under realistic and diverse video conditions.

The benchmark comprises 50 Full HD 1080p video sequences totaling 3,455 frames, each manually annotated with pixel-accurate ground truth segmentation masks. The sequences cover four diverse content categories (humans, animals, vehicles, and objects) and systematically incorporate 15 challenging real-world attributes, including occlusions, motion blur, fast motion, and dynamic backgrounds. Using these data, the article evaluated 12 state-of-the-art unsupervised and semi-supervised algorithms alongside two preprocessing methods across three complementary metrics: region similarity (Jaccard index), contour accuracy (F-measure), and temporal stability.

The empirical findings demonstrate that current state-of-the-art segmentation methods leave substantial room for improvement. The top-performing methods, Non-Local Consensus (NLC) and Fully Connected Proposals (FCP), achieved average region similarity scores of only 0.641 and 0.631, respectively, falling short of the 0.70 threshold considered necessary for production-quality accuracy. Furthermore, sequential frame-propagation techniques suffered severe performance decay over time, whereas non-local, proposal-based methods maintained higher temporal stability and recall. Specific scene challenges caused sharp performance drops across most models; for instance, appearance changes degraded the performance of color-reliant models by nearly 50%, while dynamic backgrounds and fast motion severely disrupted optical-flow and trajectory-clustering methods.

These findings indicate that existing video segmentation algorithms remain too brittle and computationally demanding for fully automated commercial deployment in post-production and large-scale video processing pipelines. To make meaningful operational progress, development must move away from simple sequential frame propagation and localized motion clustering toward architectures that leverage robust object proposals, non-local spatio-temporal connectivity, and high-resolution contour refinement.

Organizations developing or deploying video segmentation tools should prioritize proposal-driven, non-locally connected models like FCP or NLC when handling unconstrained video, while utilizing boundary-focused refinement techniques to ensure contour precision. Future research must specifically target computational efficiency and memory footprints, as heavy preprocessing requirements for motion estimation and boundary extraction currently limit the real-time utility of these methods on high-resolution Full HD video.

  • Paper: The Pascal Visual Object Classes Challenge: A Retrospective, M. Everingham et al. (2014). Reading the PASCAL VOC retrospective first clarifies the benchmark evolution and annotation practices that directly enabled modern video segmentation datasets like DAVIS.
  • Paper: Object Tracking Benchmark, Yi Wu et al. (2015). Understanding this foundational object tracking benchmark provides essential context for the rigorous evaluation protocols and performance metrics later adopted in video segmentation.
Cover for A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation

Abstract

Over the years, datasets and benchmarks have proven their fundamental importance in computer vision research, enabling targeted progress and objective comparisons in many fields. At the same time, legacy datasets may impend the evolution of a field due to saturated algorithm performance and the lack of contemporary, high quality data. In this work we present a new benchmark dataset and evaluation methodology for the area of video object segmentation. The dataset, named DAVIS (Densely Annotated VIdeo Segmentation), consists of fifty high quality, Full HD video sequences, spanning multiple occurrences of common video object segmentation challenges such as occlusions, motion-blur and appearance changes. Each video is accompanied by densely annotated, pixel-accurate and per-frame ground truth segmentation. In addition, we provide a comprehensive analysis of several state-of-the-art segmentation approaches using three complementary metrics that measure the spatial extent of the segmentation, the accuracy of the silhouette contours and the temporal coherence. The results uncover strengths and weaknesses of current approaches, opening up promising directions for future works.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. Datasets
  • 2.2. Algorithms
  • 3. Dataset Description
  • 4. Experimental Validation
  • 4.1. Metrics Selection
  • 4.2. Metrics Validation
  • 5. Evaluated Algorithms
  • 6. Quantitative Evaluation
  • 6.1. Error Measure Statistics
  • 6.2. Attributes-based Evaluation
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — DAVIS Benchmark Dataset Specification

    definition

    DAVIS (Densely Annotated Video Segmentation) is a benchmark dataset designed specifically for binary video object segmentation, separating target foreground object(s) from the background.

    The dataset features:

    • Scale and Resolution: 50 video sequences comprising 3,455 total frames, recorded at 24 frames per second at Full HD 1080p spatial resolution.
    • Temporal Extent: Short sequences ranging from 2 to 4 seconds, designed to pack dense occurrences of realistic video segmentation challenges without excessive computational overhead.
    • Ground Truth: Pixel-accurate, manually annotated binary segmentation masks provided for every single frame.
    • Object Presence: Each clip contains either a single prominent foreground object or two spatially connected interacting objects, spanning four evenly distributed object categories: humans, animals, vehicles, and objects.
    • Annotation Attributes: Every video is annotated with a subset of 15 challenge attributes (including background clutter, deformation, fast motion, motion blur, occlusion, appearance changes, and shape complexity).
  2. Knowl 2 — Region Similarity and Contour Accuracy Metrics

    equation

    Video object segmentation quality per frame is evaluated through two complementary spatial metrics:

    1. Region Similarity (Jaccard Index J\mathcal{J}): Evaluates the spatial overlap between the estimated binary segmentation mask MM and the ground truth mask GG as the intersection-over-union:

    J=∣M∩G∣∣M∪G∣\mathcal{J} = \frac{|M \cap G|}{|M \cup G|}

    1. Contour Accuracy (Boundary F-measure F\mathcal{F}): Evaluates the spatial precision of the object boundary contours. Given closed contours c(M)c(M) and c(G)c(G) representing the boundaries of MM and GG respectively, contour-based precision PcP_c and recall RcR_c are calculated via bipartite graph matching between boundary points (approximated efficiently with morphological dilation operators):

    F=2PcRcPc+Rc\mathcal{F} = \frac{2 P_c R_c}{P_c + R_c}

  3. Knowl 3 — Temporal Boundary Instability Metric

    algorithm

    The temporal stability metric T\mathcal{T} penalizes boundary oscillations, jitter, and shape flickering across consecutive frames while remaining invariant to smooth deformation and rigid motion. Given consecutive segmentation masks MtM_t and Mt+1M_{t+1}:

    Input: Binary segmentation masks Mt,Mt+1M_t, M_{t+1} at consecutive frames t,t+1t, t+1
    Output: Temporal instability cost Tt→t+1\mathcal{T}_{t \to t+1}
    Extract boundary polygons P(Mt)={pt1,…,ptNt}P(M_t) = \{p_t^1, \dots, p_t^{N_t}\} and P(Mt+1)={pt+11,…,pt+1Nt+1}P(M_{t+1}) = \{p_{t+1}^1, \dots, p_{t+1}^{N_{t+1}}\}
    for each point pti∈P(Mt)p_t^i \in P(M_t) do
        Compute Shape Context Descriptor (SCD) stis_t^i
    for each point pt+1j∈P(Mt+1)p_{t+1}^j \in P(M_{t+1}) do
        Compute Shape Context Descriptor (SCD) st+1js_{t+1}^j
    Compute distance matrix Di,j=∥sti−st+1j∥2D_{i,j} = \|s_t^i - s_{t+1}^j\|_2
    Find cyclic point matching π:{1,…,Nt}→{1,…,Nt+1}\pi: \{1, \dots, N_t\} \to \{1, \dots, N_{t+1}\} minimizing matching cost using Dynamic Time Warping (DTW) while preserving boundary order
    Tt→t+1←\mathcal{T}_{t \to t+1} \leftarrow mean SCD matching distance per point under π\pi
    return Tt→t+1\mathcal{T}_{t \to t+1}

    Lower values of T\mathcal{T} indicate greater temporal stability. Because severe occlusions and rapid non-rigid deformations cause legitimate boundary discontinuities, T\mathcal{T} is evaluated on sequences without occlusions or extreme deformations. Ground-truth masks achieve a baseline stability of T=0.093\mathcal{T} = 0.093.

  4. Knowl 4 — Aggregate Evaluation Statistics: Mean, Temporal Decay, and Object Recall

    equation

    For any per-frame error metric C∈{J,F,T}C \in \{\mathcal{J}, \mathcal{F}, \mathcal{T}\} evaluated over a dataset of video sequences R={Si}i=1∣R∣R = \{S_i\}_{i=1}^{|R|}, where Cˉ(Si)\bar{C}(S_i) is the mean metric score on sequence SiS_i:

    1. Mean Performance MC(R)\mathcal{M}_C(R):

    MC(R)=1∣R∣∑Si∈RCˉ(Si)\mathcal{M}_C(R) = \frac{1}{|R|} \sum_{S_i \in R} \bar{C}(S_i)

    1. Temporal Decay DC(R)\mathcal{D}_C(R): Quantifies the performance degradation or drift over time by partitioning each sequence SiS_i into temporal quartiles Qi={Qi1,Qi2,Qi3,Qi4}Q_i = \{Q_i^1, Q_i^2, Q_i^3, Q_i^4\} and measuring the score drop from the first to the final quartile:

    DC(R)=1∣R∣∑Si∈R(Cˉ(Qi1)−Cˉ(Qi4))\mathcal{D}_C(R) = \frac{1}{|R|} \sum_{S_i \in R} \left(\bar{C}(Q_i^1) - \bar{C}(Q_i^4)\right)

    1. Object Recall OC(R)\mathcal{O}_C(R): Measures the fraction of video sequences whose average score exceeds a quality threshold τ\tau (set to τ=0.5\tau = 0.5):

    OC(R)=1∣R∣∑Si∈R1{Cˉ(Si)>τ}\mathcal{O}_C(R) = \frac{1}{|R|} \sum_{S_i \in R} \mathbf{1}_{\{\bar{C}(S_i) > \tau\}}

  5. Knowl 5 — Comparative Performance of Video Object Segmentation Methods on DAVIS

    data/table

    The benchmark evaluates 12 video segmentation methods and 3 baselines downsampled to 480p resolution. Semi-supervised methods receive the ground truth mask for the first frame.

    Algorithm Region Similarity J\mathcal{J} Contour Accuracy F\mathcal{F} Stability T\mathcal{T}
    Mean ↑\uparrow Recall ↑\uparrow Decay ↓\downarrow Mean ↑\uparrow Recall ↑\uparrow Decay ↓\downarrow Mean ↓\downarrow
    Preprocessing
    MCG 0.724 0.912 0.026 0.654 0.781 0.046 0.652
    SF-LAB 0.173 0.075 -0.020 0.218 0.052 -0.016 0.758
    SF-MOT 0.532 0.672 0.050 0.452 0.440 0.052 0.637
    Unsupervised
    NLC 0.641 0.731 0.086 0.593 0.658 0.086 0.356
    CVOS 0.514 0.581 0.127 0.490 0.578 0.138 0.243
    TRC 0.501 0.560 0.050 0.478 0.519 0.066 0.327
    MSG 0.543 0.636 0.028 0.525 0.613 0.057 0.250
    KEY 0.569 0.671 0.075 0.503 0.534 0.079 0.190
    SAL 0.426 0.386 0.084 0.383 0.264 0.072 0.600
    FST 0.575 0.652 0.044 0.536 0.579 0.065 0.276
    Semi-Supervised
    TSP 0.358 0.388 0.385 0.346 0.329 0.388 0.329
    SEA 0.556 0.606 0.355 0.533 0.559 0.339 0.137
    HVS 0.596 0.698 0.197 0.576 0.712 0.202 0.296
    JMP 0.607 0.693 0.372 0.586 0.656 0.373 0.131
    FCP 0.631 0.778 0.031 0.546 0.604 0.039 0.285

    Key takeaways:

    • NLC achieves the highest region similarity among unsupervised approaches (J=0.641\mathcal{J} = 0.641), and FCP leads semi-supervised approaches (J=0.631\mathcal{J} = 0.631) while exhibiting very low decay (DJ=0.031\mathcal{D}_\mathcal{J} = 0.031).
    • Sequential frame-propagation approaches (JMP, SEA, TSP) achieve the best temporal stability (lowest T\mathcal{T} around 0.1310.131--0.1370.137) but suffer substantial temporal drift (high decay DJ>0.35\mathcal{D}_\mathcal{J} > 0.35).
    • Motion saliency (SF-MOT, J=0.532\mathcal{J} = 0.532) substantially outperforms color saliency (SF-LAB, J=0.173\mathcal{J} = 0.173) for isolating foreground video objects.
  6. Knowl 6 — Video Challenge Attributes and Dependency Estimation via Stability Selection

    model/method

    To identify failure modes of segmentation algorithms, video clips are labeled with 15 binary challenge attributes:

    • Background Clutter (BC): Similar color distributions between background and foreground boundaries (χ2\chi^2 distance on histograms).
    • Deformation (DEF): Non-rigid object deformation.
    • Motion Blur (MB): Blurred/fuzzy object contours.
    • Fast Motion (FM): Centroid displacement exceeds 20 pixels/frame.
    • Low Resolution (LR): Target bounding box area is <10%< 10\% of frame area.
    • Occlusion (OCC): Object is partially or fully occluded.
    • Out-of-view (OV): Object touches/extends past image boundaries.
    • Scale Variation (SV): Min-to-max bounding box area ratio across frames is <0.5< 0.5.
    • Appearance Change (AC): Substantial change in lighting or 3D viewpoint.
    • Edge Ambiguity (EA): Average ground truth boundary probability is <0.5< 0.5.
    • Camera Shake (CS): Non-negligible video vibration.
    • Heterogeneous Object (HO): Object consists of multiple distinct color regions.
    • Interacting Objects (IO): Multiple spatially connected moving entities.
    • Dynamic Background (DB): Background contains motion or deformations (e.g. water, foliage).
    • Shape Complexity (SC): Intricate boundaries, holes, or thin structures.

    To decouple correlated attributes, attribute dependencies are modeled as a pairwise Markov Random Field over binary graph G=(V,E)G=(V, E) (∣V∣=16|V|=16) and estimated via ℓ1\ell_1-penalized logistic regression combined with stability selection across n/2n/2 subsamples. At selection probability threshold 0.6, strong mutual dependencies are uncovered between Fast Motion (FM) and Motion Blur (MB), as well as between Interacting Objects (IO) and Shape Complexity (SC).

  7. Knowl 7 — Attribute-Based Performance Sensitivity Analysis

    data/table

    The performance of segmentation methods varies strongly across challenging video attributes. The table below lists average region similarity J\mathcal{J} on sequences exhibiting a specific attribute, alongside the gain (++) or loss (−-) on sequences where the attribute is absent:

    Method AC DB FM MB OCC
    J\mathcal{J} Δ\Delta J\mathcal{J} Δ\Delta J\mathcal{J} Δ\Delta J\mathcal{J} Δ\Delta J\mathcal{J} Δ\Delta
    NLC 0.54 +0.13 0.53 +0.15 0.64 +0.00 0.61 +0.04 0.70 -0.09
    CVOS 0.42 +0.12 0.37 +0.18 0.37 +0.24 0.36 +0.23 0.43 +0.13
    TRC 0.37 +0.17 0.39 +0.15 0.41 +0.16 0.32 +0.27 0.44 +0.10
    MSG 0.48 +0.08 0.43 +0.15 0.46 +0.14 0.35 +0.29 0.48 +0.10
    KEY 0.42 +0.19 0.52 +0.07 0.50 +0.12 0.51 +0.08 0.52 +0.08
    SAL 0.33 +0.12 0.35 +0.10 0.35 +0.13 0.33 +0.15 0.44 -0.02
    FST 0.55 +0.04 0.53 +0.06 0.50 +0.12 0.48 +0.14 0.53 +0.07
    TSP 0.17 +0.23 0.40 -0.06 0.18 +0.31 0.15 +0.32 0.27 +0.14
    SEA 0.46 +0.12 0.58 -0.03 0.40 +0.28 0.39 +0.24 0.47 +0.13
    HVS 0.42 +0.23 0.60 -0.01 0.42 +0.31 0.44 +0.24 0.53 +0.11
    JMP 0.58 +0.03 0.60 +0.01 0.50 +0.18 0.51 +0.15 0.47 +0.21
    FCP 0.51 +0.16 0.62 +0.01 0.55 +0.13 0.53 +0.15 0.59 +0.07

    Key empirical findings:

    • Appearance Change (AC): Severely hurts approaches that update appearance with Gaussian processes or rely heavily on color consistency (e.g. TSP drops to J=0.17\mathcal{J}=0.17).
    • Dynamic Background (DB): Disproportionately impacts unsupervised motion-saliency techniques (NLC, SAL, MSG, TRC) due to invalid translational motion assumptions on non-rigid background regions.
    • Fast Motion (FM) & Motion Blur (MB): Degrade optical-flow estimation and boundary precision across nearly all algorithms, with TSP dropping by >0.31>0.31 IoU.
    • Occlusions (OCC): Cause sequential propagation models (JMP, SEA) to fail and drift, whereas globally connected models (NLC, FCP) maintain higher robustness.
  8. Knowl 8 — Complementarity and Independence of Spatial and Temporal Metrics

    empirical result

    Pairwise correlation analysis across frames on the DAVIS benchmark establishes the necessity of using region similarity J\mathcal{J}, contour accuracy F\mathcal{F}, and temporal instability T\mathcal{T} jointly:

    • J\mathcal{J} vs F\mathcal{F} Correlation: While region similarity J\mathcal{J} and contour accuracy F\mathcal{F} exhibit a moderate positive correlation, they diverge systematically in common failure regimes. Missing a thin or distant object appendage causes large region loss (low J\mathcal{J}) with minimal impact on contour recall (high F\mathcal{F}). Conversely, jagged, noisy, or vibrating contours with low area error preserve high J\mathcal{J} but yield poor F\mathcal{F}.
    • T\mathcal{T} Independence: Temporal instability T\mathcal{T} is virtually uncorrelated with both per-frame region overlap J\mathcal{J} and contour accuracy F\mathcal{F}. A method can produce high per-frame spatial accuracy while exhibiting high frame-to-frame boundary jitter, or conversely maintain smooth frame-to-frame temporal tracking that drifts away from ground truth.
  9. Knowl 9 — Computational and Resolution Scaling Bottlenecks in Video Object Segmentation

    limitation

    Existing video object segmentation methods face severe computational scalability limits:

    • Preprocessing Bottleneck: A predominant share of total runtime is consumed by dense low-level preprocessing, notably optical flow estimation, superpixel extraction, and object proposal generation.
    • Resolution Downsampling: Due to the memory and runtime complexity of global optimization routines (e.g. dense CRFs, large spatio-temporal graphs), state-of-the-art methods could not be evaluated directly at native 1080p resolution and required downsampling to 480p.
    • Impact on Fine Details: While 480p downsampling has modest impact on coarse region overlap (J\mathcal{J}), native Full HD resolution remains essential for segmenting thin object structures, complex silhouette boundaries, and fine topological details.

Coverage note — None was omitted; all key contributions including dataset design, evaluation metric formulations, aggregate benchmarks, attribute sensitivity analysis, and computational observations are covered.

References

  1. 1.V. Badrinarayanan, F. Galasso, and R. Cipolla. Label propagation in video sequences. In CVPR, 2010. 1, 2, 3
  2. 2.X. Bai, J. Wang, D. Simons, and G. Sapiro. Video snapcut: robust video object cutout using localized classifiers. ACM Trans. Graph., 28(3), 2009. 3
  3. 3.S. Belongie, J. Malik, and J. Puzicha. Shape matching and object recognition using shape contexts. TPAMI, 24(4), 2002. 4
  4. 4.G. J. Brostow, J. Fauqueur, and R. Cipolla. Semantic object classes in video: A high-definition ground truth database. Pattern Recognition Letters, 30(2), 2009. 1, 2
  5. 5.T. Brox and J. Malik. Object segmentation by long term analysis of point trajectories. In ECCV, 2010. 1, 2, 5, 6, 7
  6. 6.D. J. Butler, J. Wulff, G. B. Stanley, and M. J. Black. A naturalistic open source movie for optical flow evaluation. In ECCV, 2012. 1, 3
  7. 7.J. Chang, D. Wei, and J. W. F. III. A video representation using temporal superpixels. In CVPR, 2013. 2, 5, 6
  8. 8.A. Y. C. Chen and J. J. Corso. Propagating multi-class pixel labels throughout video frames. In WNYIPW, 2010. 2
  9. 9.R. Collins, X. Zhou, and S. K. Teh. An open source tracking testbed and evaluation web site. In PETS 2005, January 2005. 2
  10. 10.G. Csurka, D. Larlus, and F. Perronnin. What is a good evaluation measure for semantic segmentation? In BMVC, 2013. 4
  11. 11.P. Dollár and C. L. Zitnick. Structured forests for fast edge detection. In ICCV, 2013. 3
  12. 12.M. Everingham, L. J. V. Gool, C. K. I. Williams, J. M. Winn, and A. Zisserman. The pascal visual object classes (VOC) challenge. IJCV, 88(2), 2010. 1, 4
  13. 13.A. Faktor and M. Irani. Video segmentation by non-local consensus voting. In BMVC, 2014. 2, 5, 6
  14. 14.Q. Fan, F. Zhong, D. Lischinski, D. Cohen-Or, and B. Chen. Jumpcut: Non-successive mask transfer and interpolation for video cutout. ACM Trans. Graph., 34(6), 2015. 2, 3, 5, 6
  15. 15.A. Fathi, X. Ren, and J. M. Rehg. Learning to recognize objects in egocentric activities. In CVPR, 2011. 1, 2
  16. 16.R. B. Fisher. The pets04 surveillance ground-truth data sets. 2004. 2
  17. 17.K. Fragkiadaki and J. Shi. Detection free tracking: Exploiting motion and topology for segmenting and tracking under entanglement. In CVPR, 2011. 2
  18. 18.K. Fragkiadaki, G. Zhang, and J. Shi. Video segmentation by tracing discontinuities in a trajectory embedding. In CVPR, 2012. 2, 5, 6
  19. 19.F. Galasso, N. S. Nagaraja, T. J. Cardenas, T. Brox, and B. Schiele. A unified video segmentation benchmark: Annotation, metrics and analysis. In ICCV, 2013. 2
  20. 20.L. Gorelick, M. Blank, E. Shechtman, M. Irani, and R. Basri. Actions as space-time shapes. TPAMI, 29(12), 2007. 1, 2
  21. 21.M. Grundmann, V. Kwatra, M. Han, and I. A. Essa. Efficient hierarchical graph-based video segmentation. In CVPR, 2010. 1, 2, 5, 6
  22. 22.S. D. Jain and K. Grauman. Supervoxel-consistent foreground propagation in video. In ECCV, 2014. 3
  23. 23.P. Krähenbühl and V. Koltun. Geodesic object proposals. In ECCV, 2014. 7
  24. 24.Y. J. Lee, J. Kim, and K. Grauman. Key-segments for video object segmentation. In Proc. ICCV, 2011. 2, 5, 6
  25. 25.F. Li, T. Kim, A. Humayun, D. Tsai, and J. M. Rehg. Video segmentation by tracking many figure-ground segments. In ICCV, 2013. 1, 2
  26. 26.T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft COCO: common objects in context. In ECCV, 2014. 1
  27. 27.T. Liu, Z. Yuan, J. Sun, J. Wang, N. Zheng, X. Tang, and H. Shum. Learning to detect a salient object. TPAMI, 33(2), 2011. 3
  28. 28.D. Martin, C. Fowlkes, and J. Malik. Learning to detect natural image boundaries using local brightness, color, and texture cues. TPAMI, 26(5), 2004. 4
  29. 29.D. R. Martin, C. Fowlkes, D. Tal, and J. Malik. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In ICCV, 2001. 1
  30. 30.N. Meinshausen and P. Bühlmann. Stability selection. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 72(4):417–473, 2010. 7
  31. 31.N. Nicolas Märki, F. Perazzi, O. Wang, and A. Sorkine-Hornung. Bilateral space video segmentation. In CVPR, 2016. 3, 6
  32. 32.S. Oh, A. Hoogs, A. G. A. Perera, and M. Desai. A large-scale benchmark dataset for event recognition in surveillance video. In CVPR, 2011. 2
  33. 33.A. Papazoglou and V. Ferrari. Fast object segmentation in unconstrained video. In ICCV, 2013. 2, 5, 6
  34. 34.F. Perazzi, P. Krähenbühl, Y. Pritch, and A. Hornung. Saliency filters: Contrast based filtering for salient region detection. In CVPR, 2012. 2, 5, 6
  35. 35.F. Perazzi, O. Wang, M. Gross, and A. Sorkine-Hornung. Fully connected object proposals for video segmentation. In ICCV, 2015. 2, 3, 5, 6
  36. 36.J. Pont-Tuset, P. Arbelaez, J. T. Barron, F. Marques, and J. Malik. Multiscale combinatorial grouping for image segmentation and object proposal generation. TPAMI, 2016. 5, 6
  37. 37.J. Pont-Tuset and F. Marques. Supervised evaluation of image segmentation and object proposal techniques. TPAMI, 2015. 4
  38. 38.A. Prest, C. Leistner, J. Civera, C. Schmid, and V. Ferrari. Learning object class detectors from weakly annotated video. In CVPR, 2012. 1, 2
  39. 39.L. Rabiner and B.-H. Juang. Fundamentals of speech recognition. 1993. 4
  40. 40.S. A. Ramakanth and R. V. Babu. Seamseg: Video object segmentation using patch seams. In CVPR, 2014. 2, 3, 5, 6
  41. 41.X. Ren and M. Philipose. Egocentric recognition of handled objects: Benchmark and analysis. In CVPR Workshops, 2009. 1, 2
  42. 42.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and F. Li. Imagenet large scale visual recognition challenge. CoRR, abs/1409.0575, 2014. 1
  43. 43.J. Shen, W. Wenguan, and F. Porikli. Saliency-Aware geodesic video object segmentation. In CVPR, 2015. 2, 5, 6
  44. 44.P. Sundberg, T. Brox, M. Maire, P. Arbelaez, and J. Malik. Occlusion boundary detection and figure/ground assignment from optical flow. In CVPR, 2011. 1, 2
  45. 45.B. Taylor, V. Karasev, and S. Soatto. Causal video object segmentation from persistence of occlusions. In CVPR, 2015. 2, 5, 6
  46. 46.R. Tron and R. Vidal. A benchmark for the comparison of 3-d motion segmentation algorithms. In CVPR, 2007. 1, 2
  47. 47.D. Tsai, M. Flagg, and J. M. Rehg. Motion coherent tracking with multi-label MRF optimization. In BMVC, 2010. 1, 2
  48. 48.S. Vijayanarasimhan and K. Grauman. Active frame selection for label propagation in videos. In ECCV, 2012. 3
  49. 49.T. Wang, B. Han, and J. P. Collomosse. Touchcut: Fast image and video segmentation using single-touch interaction. Computer Vision and Image Understanding, 120, 2014. 3
  50. 50.Y. Wu, J. Lim, and M. Yang. Online object tracking: A benchmark. In CVPR, 2013. 2, 3
  51. 51.C. Xu and J. J. Corso. Evaluation of super-voxel methods for early video processing. In CVPR, 2012. 2
  52. 52.D. Zhang, O. Javed, and M. Shah. Video object segmentation through spatially accurate and temporally dense extraction of primary object regions. In CVPR, 2013. 2
  53. 53.F. Zhong, X. Qin, Q. Peng, and X. Meng. Discontinuity-aware video object cutout. ACM Trans. Graph., 31(6), 2012. 3

Citation

MLA
Perazzi, F., et al. “A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation”. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 724–32, https://doi.org/10.1109/CVPR.2016.85.
APA
Perazzi, F., Pont-Tuset, J., McWilliams, B., Van Gool, L., Gross, M., & Sorkine-Hornung, A. (2016). A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 724–732. https://doi.org/10.1109/CVPR.2016.85
Chicago
Perazzi, F., J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine-Hornung. 2016. “A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation”. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 724–32. https://doi.org/10.1109/CVPR.2016.85.
Harvard
Perazzi, F. et al. (2016) “A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation”, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 724–732. Available at: https://doi.org/10.1109/CVPR.2016.85.
Vancouver
1. Perazzi F, Pont-Tuset J, McWilliams B, Van Gool L, Gross M, Sorkine-Hornung A (2016) A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 724–732

BibTeX

@inproceedings{Perazzi_2016, title={A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation}, url={http://dx.doi.org/10.1109/CVPR.2016.85}, DOI={10.1109/cvpr.2016.85}, booktitle={2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Perazzi, F. and Pont-Tuset, J. and McWilliams, B. and Van Gool, L. and Gross, M. and Sorkine-Hornung, A.}, year={2016}, month=June, pages={724–732} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE