Simultaneous Detection and Segmentation

Bharath HariharanPablo ArbeláezRoss B. GirshickJitendra Malik

article2014ECCV1,348 citations

Introduces the simultaneous detection and segmentation framework, extending region-based convolutional networks to detect individual object instances and delineate their exact pixel boundaries.

Listen

Visual recognition systems have traditionally separated object detection, which provides coarse bounding boxes around individual objects, from semantic segmentation, which assigns category labels to pixels without distinguishing individual object instances. This division limits performance in applications requiring both exact pixel boundaries and instance counts, such as autonomous systems, robotics, and image editing. The article addresses this gap by defining and tackling Simultaneous Detection and Segmentation, a unified task aimed at identifying every instance of an object category within an image while accurately delineating the exact pixels belonging to each instance.

The primary objective of the article is to develop and evaluate a deep learning framework tailored for Simultaneous Detection and Segmentation, demonstrating that unified instance-level segmentation enhances both classical object detection and semantic segmentation benchmarks. The approach utilizes a multi-step pipeline: generating bottom-up region proposals, extracting visual representations using a dual-pathway convolutional neural network trained simultaneously on bounding boxes and masked foregrounds, classifying these proposals, and refining the final shapes using learned category-specific figure-ground predictions. The methodology was evaluated on standard benchmark datasets using region-based average precision metrics across varying overlap thresholds.

The findings show that this integrated pipeline substantially improves performance across all relevant tasks. On the primary instance segmentation metric, the full system achieved an average precision of 49.7%, delivering an approximate 16% relative improvement over baseline models and outperforming prior segmentation methods. In traditional bounding box object detection, performance increased to 53.0% average precision, outperforming standard single-frame detectors. When converted to standard semantic segmentation, the model achieved a 52.6% mean intersection-over-union score on the benchmark test set, advancing the state-of-the-art by about 10% relative. Diagnostic evaluations revealed that mislocalization constitutes the largest bottleneck, accounting for an estimated 16 percentage point setback, whereas confusion between different object categories or backgrounds was negligible.

These results imply that bounding box detection and semantic segmentation are mutually beneficial and should not be treated as separate engineering problems. In production and decision-making contexts, unifying these functions simplifies machine vision architectures into a single pipeline while improving spatial fidelity and instance awareness. The diagnostic analysis indicates that future technical development should concentrate on improving contour localization and boundary extraction rather than general category classifiers, as category recognition is already highly reliable.

The conclusions are supported by statistically significant results across standardized image benchmarks. However, the evaluation is bounded by existing proposal generation quality and standard training datasets with predefined 20-class object taxonomies. Future initiatives should pilot this simultaneous detection and segmentation approach on domain-specific datasets and explore end-to-end localization refinements to reduce remaining boundary errors.

  • Paper: Mask R-CNN, Kaiming He et al. (2017). Mask R-CNN extends and modernizes the simultaneous detection and segmentation paradigm into an efficient, unified end-to-end instance segmentation framework.
  • Paper: Fast R-CNN, Ross B. Girshick (2015). Fast R-CNN streamlines region-based deep learning by sharing convolutional computations across proposals and jointly optimizing multi-task losses.
  • Paper: Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks, Shaoqing Ren et al. (2015). Faster R-CNN replaces external bottom-up proposal algorithms with an internal Region Proposal Network, greatly advancing proposal-based object detection pipelines.
  • Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). Introduces fully convolutional architectures for dense pixel-level prediction, profoundly influencing subsequent top-down and bottom-up segmentation architectures.
  • Paper: Hybrid Task Cascade for Instance Segmentation, Kai Chen et al. (2019). Hybrid Task Cascade extends instance segmentation by interweaving multi-stage bounding box localization and mask prediction for reciprocal refinement.
  • Paper: YOLACT: Real-Time Instance Segmentation, Daniel Bolya et al. (2019). YOLACT explores real-time single-stage instance segmentation by decoupling the task into prototype mask generation and linear coefficient prediction.
  • Paper: Panoptic Segmentation, Alexander Kirillov et al. (2018). Panoptic Segmentation unifies instance segmentation with semantic background segmentation under a coherent global formulation.
  • Paper: Path Aggregation Network for Instance Segmentation, Shu Liu et al. (2018). Path Aggregation Network improves information flow and localization features within proposal-based instance segmentation networks.
  • Paper: Cascade R-CNN: High Quality Object Detection and Instance Segmentation, Zhaowei Cai et al. (2019). Cascade R-CNN demonstrates progressive multi-stage hypothesis refinement for high-quality object detection and instance segmentation.
  • Paper: Image Segmentation Using Deep Learning: A Survey, Shervin Minaee et al. (2020). Provides a comprehensive retrospective survey evaluating how deep learning architectures for semantic and instance segmentation evolved following early proposal-based methods.
Cover for Simultaneous Detection and Segmentation

Abstract

We aim to detect all instances of a category in an image and, for each instance, mark the pixels that belong to it. We call this task Simultaneous Detection and Segmentation (SDS). Unlike classical bounding box detection, SDS requires a segmentation and not just a box. Unlike classical semantic segmentation, we require individual object instances. We build on recent work that uses convolutional neural networks to classify category-independent region proposals (R-CNN [16]), introducing a novel architecture tailored for SDS. We then use category-specific, top- down figure-ground predictions to refine our bottom-up proposals. We show a 7 point boost (16% relative) over our baselines on SDS, a 5 point boost (10% relative) over state-of-the-art on semantic segmentation, and state-of-the-art performance in object detection. Finally, we provide diagnostic tools that unpack performance and provide directions for future work.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Our approach
  • 3.1 Proposal generation
  • 3.2 Feature extraction
  • 3.3 Region classification
  • 3.4 Region refinement
  • 4 Experiments and results
  • 4.1 Results on APr and APv​o​lr{}^{r}_{vol}
  • 4.2 Producing diagnostic information
  • 4.3 Results on APb and APv​o​lb{}^{b}_{vol}
  • 4.4 Results on pixel IU
  • References

Knowls

  1. Knowl 1 — Simultaneous Detection and Segmentation Task and Metrics

    definition

    Simultaneous Detection and Segmentation (SDS) is an object recognition task that requires detecting all individual object instances of predefined "thing" categories in an image and predicting a precise binary pixel segmentation mask for each detected instance.

    An SDS algorithm outputs a set of hypotheses H={(Mi,ci,si)}H = \{(M_i, c_i, s_i)\}, where MiM_i is a predicted binary pixel segmentation mask, cic_i is the predicted object category, and si∈Rs_i \in \mathbb{R} is the confidence score.

    Performance is evaluated using instance-level Average Precision (APrAP^r, where superscript rr denotes region segmentation):

    1. Region Average Precision (APrAP^r): A predicted instance (M,c,s)(M, c, s) is classified as a true positive if its intersection-over-union (IoU) overlap with an unassigned ground-truth instance mask GG of category cc exceeds a threshold of 0.500.50 (50%): IoU(M,G)=∣M∩G∣∣M∪G∣>0.50\text{IoU}(M, G) = \frac{|M \cap G|}{|M \cup G|} > 0.50 Duplicate detections matched to an already assigned ground-truth instance are treated as false positives. APrAP^r is computed as the area under the precision-recall (PR) curve for each category and averaged across categories.

    2. Volume-under-PR-Surface (APvolrAP^r_{vol}): To assess performance across varying levels of segmentation precision, APr(τ)AP^r(\tau) is computed across multiple overlap thresholds τ∈[0.1,0.9]\tau \in [0.1, 0.9]. APvolrAP^r_{vol} is evaluated as the numerical average of APrAP^r over 9 uniformly spaced thresholds τ∈{0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8,0.9}\tau \in \{0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9\}.

  2. Knowl 2 — Dual-Pathway Jointly Fine-Tuned CNN Architecture for Region Feature Extraction

    model/method

    The SDS feature extractor processes category-independent candidate region proposals generated by Multiscale Combinatorial Grouping (MCG). For each candidate region RR with tight bounding box BB, the system extracts representations using a two-pathway convolutional neural network (CNN):

    1. Box Pathway: Operates on the image patch extracted from the padded bounding box BB, warped to a fixed input size (227×227227 \times 227). It captures global appearance and scene context surrounding the object.

    2. Region Pathway: Operates on the image patch extracted from bounding box BB, warped to 227×227227 \times 227, with all pixels outside the candidate region mask RR masked out (replaced by the dataset mean RGB pixel values). It forces the network to focus strictly on the object's foreground appearance.

    Both pathways share the AlexNet-style convolutional architecture (five convolutional layers and two fully connected layers fc6,fc7fc_6, fc_7). The outputs of the penultimate layers (fc7fc_7) of both pathways are concatenated into a joint feature vector of dimension 4096+4096=81924096 + 4096 = 8192.

    Three training regimes are evaluated:

    • Feature Extractor A: A single network pre-trained on ImageNet is fine-tuned on PASCAL bounding boxes (using box overlap >50%> 50\% as positive) and applied to both pathways with identical weights.

    • Feature Extractor B: Two disjoint networks are fine-tuned separately: the box pathway is fine-tuned on bounding box overlap (>50%> 50\%), while the region pathway is fine-tuned on masked region proposals supervised by segmentation mask IoU overlap (>50%> 50\%).

    • Feature Extractor C: The entire two-pathway network is initialized from the weights of B and jointly fine-tuned end-to-end using a unified softmax classification layer placed on top of the concatenated fc7fc_7 activations, supervised by region mask overlap labels.

  3. Knowl 3 — Multiple Instance Learning Region Classification and Strict Mask NMS

    algorithm

    Region candidates scored with CNN features are classified using linear Support Vector Machines (SVMs) trained under a multiple instance learning (MIL) formulation, followed by strict mask non-maximum suppression (NMS).

    Input: Ground truth annotations G\mathcal{G}, candidate region proposals R\mathcal{R}, CNN feature extractor ϕ(⋅)\phi(\cdot)
    Output: Trained linear SVM weights wcw_c for each category cc, and filtered test detections Dc\mathcal{D}_c
    for each category cc do
        // Initial SVM training
        Positives P←{G∈G:class(G)=c}\mathcal{P} \leftarrow \{G \in \mathcal{G} : \text{class}(G) = c\}
        Negatives N←{R∈R:∀G∈G,  IoU(R,G)<0.20}\mathcal{N} \leftarrow \{R \in \mathcal{R} : \forall G \in \mathcal{G},\; \text{IoU}(R, G) < 0.20\}
        Train initial SVM weights wc(0)w_c^{(0)} on ϕ(P)\phi(\mathcal{P}) and ϕ(N)\phi(\mathcal{N})
        
        // Re-estimate positive set via Multiple Instance Learning
        PMIL←∅\mathcal{P}_{\text{MIL}} \leftarrow \emptyset
        for each ground truth instance G∈GG \in \mathcal{G} with class(G)=c\text{class}(G) = c do
            CG←{R∈R:IoU(R,G)≥0.50}\mathcal{C}_G \leftarrow \{R \in \mathcal{R} : \text{IoU}(R, G) \ge 0.50\}
            if CG≠∅\mathcal{C}_G \neq \emptyset then
                R∗←arg⁡max⁡R∈CG(wc(0)⋅ϕ(R))R^* \leftarrow \arg\max_{R \in \mathcal{C}_G} \left( w_c^{(0)} \cdot \phi(R) \right)
                PMIL←PMIL∪{R∗}\mathcal{P}_{\text{MIL}} \leftarrow \mathcal{P}_{\text{MIL}} \cup \{R^*\}
            end if
        end for
        
        // Retrain SVM with re-estimated positives
        Train final SVM weights wcw_c on ϕ(PMIL)\phi(\mathcal{P}_{\text{MIL}}) and ϕ(N)\phi(\mathcal{N})
    end for
    // Test inference and suppression per image
    for each candidate region R∈RR \in \mathcal{R} do
        Compute score sc(R)=wc⋅ϕ(R)s_c(R) = w_c \cdot \phi(R)
    end for
    Apply Greedy NMS on scored candidates using mask overlap:
    Discard candidate RjR_j if there exists RiR_i with sc(Ri)>sc(Rj)s_c(R_i) > s_c(R_j) and IoU(Ri,Rj)>0\text{IoU}(R_i, R_j) > 0
    Retain top 20,000 highest scoring detections per category across the full dataset
  4. Knowl 4 — Two-Stage Category-Specific Region Refinement

    model/method

    Because initial bottom-up region proposals from MCG are class-agnostic, they suffer from systematic overshooting (including background) or undershooting (omitting limbs/structures). SDS applies a two-stage top-down refinement process:

    1. Coarse Figure-Ground Mask Prediction: For each predicted detection box, the padded bounding box is discretized into a 10×1010 \times 10 spatial grid. For each of the 100 grid cells k∈{1,…,100}k \in \{1, \dots, 100\}, a category-specific logistic regression classifier is trained to output the probability pk∈[0,1]p_k \in [0, 1] that cell kk belongs to the object foreground. The classifier input is the concatenation of the 8192-dimensional CNN features and the binary candidate mask discretized into the same 10×1010 \times 10 grid. Training is performed on region candidates having IoU>0.70\text{IoU} > 0.70 with ground-truth instances.

    2. Superpixel Projection and Classification: The coarse 10×1010 \times 10 grid predictions are projected onto image superpixels by assigning each superpixel SS the average probability value pˉS\bar{p}_S of the grid cells overlapping SS. A secondary classifier then predicts the final foreground label for each superpixel using two input features:

    • The projected continuous probability value pˉS\bar{p}_S.
    • A binary flag bS∈{0,1}b_S \in \{0, 1\} indicating whether superpixel SS was part of the original bottom-up candidate region.

    Thresholding the resulting superpixel probabilities yields the final refined instance segmentation mask.

  5. Knowl 5 — Pixel Precision-Recall Guarantee and SDS Diagnostic Framework

    theoretical result

    To analyze whether region mislocalizations stem from overshooting (including background) or undershooting (missing parts of the object), the segmentation overlap is decomposed into pixel precision (PPPP) and pixel recall (PRPR): PP(M,G)=∣M∩G∣∣M∣,PR(M,G)=∣M∩G∣∣G∣PP(M, G) = \frac{|M \cap G|}{|M|}, \qquad PR(M, G) = \frac{|M \cap G|}{|G|} where MM is the predicted binary region mask and GG is the ground-truth instance mask.

    Theoretical Guarantee: If a predicted mask MM and ground truth mask GG satisfy PP(M,G)>23PP(M, G) > \frac{2}{3} and PR(M,G)>23PR(M, G) > \frac{2}{3}, then the Jaccard overlap IoU(M,G)\text{IoU}(M, G) is strictly greater than 50%50\% (0.500.50): IoU(M,G)=∣M∩G∣∣M∪G∣=∣M∩G∣∣M∣+∣G∣−∣M∩G∣>∣M∩G∣32∣M∩G∣+32∣M∩G∣−∣M∩G∣=∣M∩G∣2∣M∩G∣=0.50\text{IoU}(M, G) = \frac{|M \cap G|}{|M \cup G|} = \frac{|M \cap G|}{|M| + |G| - |M \cap G|} > \frac{|M \cap G|}{\frac{3}{2}|M \cap G| + \frac{3}{2}|M \cap G| - |M \cap G|} = \frac{|M \cap G|}{2|M \cap G|} = 0.50

    Diagnostic Tools:

    • Detections with 0.10≤IoU<0.500.10 \le \text{IoU} < 0.50 (or duplicates) are classified as mislocalizations (LL). In SDS, fixing mislocalizations accounts for ∼16\sim 16 percentage points of lost APrAP^r, whereas background (BB) and similar-class (SS) confusions are negligible.
    • By evaluating average precision under a strict 67%67\% threshold on pixel precision (APppAP^{pp}) versus pixel recall (APprAP^{pr}), the difference APpp−APprAP^{pp} - AP^{pr} indicates the failure mode: APpp−APpr<0AP^{pp} - AP^{pr} < 0 denotes overshooting (e.g., bicycle), whereas APpp−APpr>0AP^{pp} - AP^{pr} > 0 denotes undershooting (e.g., person, bird).
  6. Knowl 6 — Simultaneous Detection and Segmentation Performance on PASCAL VOC 2012

    data/table

    Performance of SDS variants on the PASCAL VOC 2012 validation set (trained on VOC 2012 train with SBD annotations) measured by APrAP^r (at 50%50\% IoU threshold) and APvolrAP^r_{vol} (averaged over 9 overlap thresholds from 0.10.1 to 0.90.9). All values are reported in percent (%).

    APrAP^r (%) APvolrAP^r_{vol} (%)
    Category O2P\text{O}_2\text{P} A B C C+ref O2P\text{O}_2\text{P} A B C C+ref
    aeroplane 56.5 61.8 65.7 67.4 68.4 46.8 48.3 51.1 53.2 52.3
    bicycle 19.0 43.4 49.6 49.6 49.4 21.2 39.8 42.1 42.1 42.6
    bird 23.0 46.6 47.2 49.1 52.1 22.1 39.2 40.8 42.1 42.2
    boat 12.2 27.2 30.0 29.9 32.8 13.0 25.1 27.5 27.1 28.6
    bottle 11.0 28.9 31.7 32.0 33.0 10.1 26.0 26.8 27.6 28.6
    bus 48.8 61.7 66.9 65.9 67.8 41.9 49.5 53.4 53.3 58.0
    car 26.0 46.9 50.9 51.4 53.6 24.0 39.5 42.6 42.7 45.4
    cat 43.3 58.4 69.2 70.6 73.9 39.2 50.7 56.3 57.3 58.9
    chair 4.7 17.8 19.6 20.2 19.9 6.7 17.6 18.5 19.3 19.7
    cow 15.6 38.8 42.7 42.7 43.7 14.6 32.5 36.0 36.3 37.1
    diningtable 7.8 18.6 22.8 22.9 25.7 9.9 18.5 20.6 21.4 22.8
    dog 24.2 52.6 56.2 58.7 60.6 24.0 46.8 48.9 49.0 49.5
    horse 27.5 44.3 51.9 54.4 55.9 24.4 37.7 41.9 43.6 42.9
    motorbike 32.3 50.2 52.6 53.5 58.9 28.6 41.1 43.2 43.5 45.9
    person 23.5 48.2 52.6 54.4 56.7 25.6 43.2 45.8 47.0 48.5
    pottedplant 4.6 23.8 25.7 24.9 28.5 7.0 23.4 24.8 24.4 25.5
    sheep 32.3 54.2 54.2 54.1 55.6 29.0 43.0 44.2 44.0 44.5
    sofa 20.7 26.0 32.2 31.4 32.1 18.8 26.2 29.7 29.9 30.2
    train 38.8 53.2 59.2 62.2 64.7 34.6 45.1 48.9 49.9 52.6
    tvmonitor 32.3 55.3 58.7 59.3 60.0 25.9 47.7 48.8 49.4 51.4
    Mean 25.2 42.9 47.0 47.7 49.7 23.4 37.0 39.6 40.2 41.4

    The naive baseline A improves by +17.7+17.7 points APrAP^r over O2P\text{O}_2\text{P}. Separate fine-tuning for foreground masks (B) yields a +4.1+4.1 point gain over A, joint two-pathway fine-tuning (C) adds +0.7+0.7 points, and top-down region refinement (C+ref) adds another +2.0+2.0 points, reaching 49.7%49.7\% APrAP^r (41.4%41.4\% APvolrAP^r_{vol}). All incremental steps are statistically significant at p<0.05p < 0.05 (paired sample t-test).

  7. Knowl 7 — Cross-Task Evaluation on Object Detection and Semantic Segmentation

    empirical result

    When adapted to standard bounding box detection and semantic segmentation benchmarks, the SDS pipeline establishes state-of-the-art results:

    1. Bounding Box Detection (APbAP^b and APvolbAP^b_{vol} on PASCAL VOC 2012):
    • On VOC 2012 validation, retraining the SVM classifiers on bounding box overlap yields mean APbAP^b scores: R-CNN with Selective Search = 51.0%51.0\%, R-CNN with MCG = 51.7%51.7\%, Model A = 51.9%51.9\%, Model B = 53.9%53.9\%, Model C = 53.0%53.0\%.
    • Mean APvolbAP^b_{vol} scores on VOC 2012 validation: R-CNN = 41.9%41.9\%, R-CNN-MCG = 42.4%42.4\%, Model A = 43.2%43.2\%, Model B = 44.6%44.6\%, Model C = 44.2%44.2\%.
    • On VOC 2012 test (without bounding box regression), Model C achieves 50.7%50.7\% mean APbAP^b, outperforming R-CNN (49.6%49.6\%) and SegDPM (40.7%40.7\%).
    1. Semantic Segmentation (Mean Pixel IU): Converting the output of C+ref into a dense pixel labeling using region pasting achieves:
    • PASCAL VOC 2011 test: 52.6%52.6\% mean pixel IU, surpassing O2P\text{O}_2\text{P} (47.6%47.6\%) and R-CNN (47.9%47.9\%) by ∼5\sim 5 points (10%10\% relative boost).
    • PASCAL VOC 2012 test: 51.6%51.6\% mean pixel IU, outperforming O2P\text{O}_2\text{P} (47.8%47.8\%).

Coverage note — None. All contributed components—including the SDS task definition, metrics, dual-pathway CNN architecture, MIL SVM training, two-stage region refinement, error diagnostics framework, and empirical evaluations across SDS, bounding box detection, and semantic segmentation—are fully covered.

References

  1. 1.Arbeláez, P., Pont-Tuset, J., Barron, J., Marques, F., Malik, J.: Multiscale combinatorial grouping. In: CVPR (2014)
  2. 2.Arbeláez, P., Hariharan, B., Gu, C., Gupta, S., Malik, J.: Semantic segmentation using regions and parts. In: CVPR (2012)
  3. 3.Boix, X., Gonfaus, J.M., van de Weijer, J., Bagdanov, A.D., Serrat, J., Gonzàlez, J.: Harmony potentials. IJCV 96(1) (2012)
  4. 4.Bourdev, L., Maji, S., Brox, T., Malik, J.: Detecting people using mutually consistent poselet activations. In: Daniilidis, K., Maragos, P., Paragios, N. (eds.) ECCV 2010, Part VI. LNCS, vol. 6316, pp. 168–181. Springer, Heidelberg (2010)
  5. 5.Carreira, J., Caseiro, R., Batista, J., Sminchisescu, C.: Semantic segmentation with second-order pooling. In: Fitzgibbon, A., Lazebnik, S., Perona, P., Sato, Y., Schmid, C. (eds.) ECCV 2012, Part VII. LNCS, vol. 7578, pp. 430–443. Springer, Heidelberg (2012)
  6. 6.Carreira, J., Sminchisescu, C.: Constrained parametric min-cuts for automatic object segmentation. In: CVPR (2010)
  7. 7.Dai, Q., Hoiem, D.: Learning to localize detected objects. In: CVPR (2012)
  8. 8.Dalal, N., Triggs, B.: Histograms of oriented gradients for human detection. In: CVPR (2005)
  9. 9.Deng, J., Berg, A., Satheesh, S., Su, H., Khosla, A., Fei-Fei, L.: ImageNet Large Scale Visual Recognition Competition 2012 (ILSVRC 2012) (2012), http://www.image-net.org/challenges/LSVRC/2012/
  10. 10.Donahue, J., Jia, Y., Vinyals, O., Hoffman, J., Zhang, N., Tzeng, E., Darrell, T.: Decaf: A deep convolutional activation feature for generic visual recognition. arXiv preprint arXiv:1310.1531 (2013)
  11. 11.Everingham, M., Van Gool, L., Williams, C.K.I., Winn, J., Zisserman, A.: The Pascal Visual Object Classes (VOC) Challenge. IJCV 88(2) (2010)
  12. 12.Farabet, C., Couprie, C., Najman, L., LeCun, Y.: Learning hierarchical features for scene labeling. TPAMI 35(8) (2013)
  13. 13.Felzenszwalb, P.F., Girshick, R.B., McAllester, D., Ramanan, D.: Object detection with discriminatively trained part-based models. TPAMI 32(9) (2010)
  14. 14.Fidler, S., Mottaghi, R., Yuille, A., Urtasun, R.: Bottom-up segmentation for top-down detection. In: CVPR (2013)
  15. 15.Fukushima, K.: Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biological Cybernetics 36(4) (1980)
  16. 16.Girshick, R., Donahue, J., Darrell, T., Malik, J.: Rich feature hierarchies for accurate object detection and semantic segmentation. In: CVPR (2014)
  17. 17.Hariharan, B., Arbelaez, P., Bourdev, L., Maji, S., Malik, J.: Semantic contours from inverse detectors. In: ICCV (2011)
  18. 18.Hoiem, D., Chodpathumwan, Y., Dai, Q.: Diagnosing error in object detectors. In: Fitzgibbon, A., Lazebnik, S., Perona, P., Sato, Y., Schmid, C. (eds.) ECCV 2012, Part III. LNCS, vol. 7574, pp. 340–353. Springer, Heidelberg (2012)
  19. 19.Jia, Y.: Caffe: An open source convolutional architecture for fast feature embedding (2013), http://caffe.berkeleyvision.org/
  20. 20.Kim, B.-S., Sun, M., Kohli, P., Savarese, S.: Relating things and stuff by high-order potential modeling. In: Fusiello, A., Murino, V., Cucchiara, R. (eds.) ECCV 2012 Ws/Demos, Part III. LNCS, vol. 7585, pp. 293–304. Springer, Heidelberg (2012)
  21. 21.Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. In: NIPS (2012)
  22. 22.Ladický, L., Sturgess, P., Alahari, K., Russell, C., Torr, P.H.S.: What, where and how many? Combining object detectors and CRFs. In: Daniilidis, K., Maragos, P., Paragios, N. (eds.) ECCV 2010, Part IV. LNCS, vol. 6314, pp. 424–437. Springer, Heidelberg (2010)
  23. 23.LeCun, Y., Boser, B., Denker, J.S., Henderson, D., Howard, R.E., Hubbard, W., Jackel, L.D.: Backpropagation applied to handwritten zip code recognition. Neural Computation 1(4) (1989)
  24. 24.Lowe, D.G.: Distinctive image features from scale-invariant keypoints. IJCV 60(2) (2004)
  25. 25.Mottaghi, R.: Augmenting deformable part models with irregular-shaped object patches. In: CVPR (2012)
  26. 26.Parkhi, O.M., Vedaldi, A., Jawahar, C., Zisserman, A.: The truth about cats and dogs. In: ICCV (2011)
  27. 27.van de Sande, K.E., Uijlings, J.R., Gevers, T., Smeulders, A.W.: Segmentation as selective search for object recognition. In: ICCV (2011)
  28. 28.Sermanet, P., Eigen, D., Zhang, X., Mathieu, M., Fergus, R., LeCun, Y.: Overfeat: Integrated recognition, localization and detection using convolutional networks. In: ICLR (2014)
  29. 29.Sermanet, P., Kavukcuoglu, K., Chintala, S., LeCun, Y.: Pedestrian detection with unsupervised multi-stage feature learning. In: CVPR (2013)
  30. 30.Shotton, J., Winn, J.M., Rother, C., Criminisi, A.: TextonBoost: Joint appearance, shape and context modeling for multi-class object recognition and segmentation. In: Leonardis, A., Bischof, H., Pinz, A. (eds.) ECCV 2006, Part I. LNCS, vol. 3951, pp. 1–15. Springer, Heidelberg (2006)
  31. 31.Tighe, J., Niethammer, M., Lazebnik, S.: Scene parsing with object instances and occlusion handling. In: ECCV (2010)
  32. 32.Yang, Y., Hallman, S., Ramanan, D., Fowlkes, C.C.: Layered object models for image segmentation. TPAMI 34(9) (2012)

Citation

MLA
Hariharan, B., et al. “Simultaneous Detection and Segmentation”. arXiv, 2014, http://arxiv.org/abs/1407.1808v1.
APA
Hariharan, B., Arbeláez, P., Girshick, R., & Malik, J. (2014). Simultaneous Detection and Segmentation. arXiv. http://arxiv.org/abs/1407.1808v1
Chicago
Hariharan, B., P. Arbeláez, R. Girshick, and J. Malik. 2014. “Simultaneous Detection and Segmentation”. arXiv. http://arxiv.org/abs/1407.1808v1.
Harvard
Hariharan, B. et al. (2014) “Simultaneous Detection and Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1407.1808v1.
Vancouver
1. Hariharan B, Arbeláez P, Girshick R, Malik J (2014) Simultaneous Detection and Segmentation. arXiv

BibTeX

@article{hariharan2014simultaneous,
  title = {Simultaneous Detection and Segmentation},
  author = {Hariharan, Bharath and Arbeláez, Pablo and Girshick, Ross and Malik, Jitendra},
  year = {2014},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1407.1808v1},
  eprint = {1407.1808}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF