Deformable ConvNets V2: More Deformable, Better Results

Xizhou ZhuHan HuStephen LinJifeng Dai

article2019CVPR2,702 citations

Proposes a reformulated Deformable Convolutional Network that integrates modulated sampling offsets and a feature mimicking mechanism to focus spatial support on relevant object regions, substantially improving object detection and instance segmentation accuracy on COCO.

Listen

The article addresses the challenge of geometric variations in objects, such as changes in scale, pose, and deformation, which complicate accurate object recognition and detection in computer vision. While the original Deformable Convolutional Networks (DCNv1) improved adaptation by learning offsets for sampling locations, analysis on the challenging COCO dataset revealed that its spatial support often extends beyond object boundaries, allowing irrelevant background content to influence features and reduce detection accuracy.

The work set out to create an enhanced version, Deformable ConvNets v2 (DCNv2), that better focuses on relevant image regions through greater modeling capacity and improved training. Researchers replaced more convolutional layers with deformable ones across stages conv3 to conv5, introduced modulated deformable modules that learn both offsets and feature amplitude modulations to control sample influence, and added an R-CNN feature mimicking loss during training to encourage features focused on object foregrounds.

Experiments integrated these changes into Faster R-CNN and Mask R-CNN systems and evaluated them on the COCO 2017 benchmark using multiple backbones. DCNv2 delivered clear gains, raising box average precision from 38.0% in the DCNv1 baseline to 41.7% on Faster R-CNN with ResNet-50, and similar improvements of 2-3 points on Mask R-CNN, with only modest increases in parameters and computation. Visualizations confirmed tighter alignment of support regions with objects, and the approach scaled effectively across ResNet and ResNeXt backbones.

These results matter because they advance detection and instance segmentation accuracy on a widely used benchmark without heavy computational cost, reducing errors from extraneous image content and supporting more reliable performance in varied real-world scenes. The gains hold across input resolutions and suggest broader applicability to tasks like classification when pretrained on ImageNet.

Next steps include releasing the code for wider use and exploring the modules in additional vision tasks. Limitations center on validation primarily with the COCO dataset, where larger training sets reduce the benefit of ImageNet pretraining for offsets; further testing on smaller or domain-specific datasets would strengthen confidence in generalization.

  • Paper: Deformable Convolutional Networks, Jifeng Dai et al. (2017). It introduces the fundamental concepts of deformable convolution and deformable RoI pooling that DCNv2 directly extends with modulation mechanisms and feature mimicking.
  • Paper: Spatial Transformer Networks, Max Jaderberg et al. (2015). It provides the foundational framework for learning differentiable geometric transformations and sampling grids inside neural networks.
  • Paper: Mask R-CNN, Kaiming He et al. (2017). It introduces the Mask R-CNN framework and RoIAlign mechanism that serve as core baseline architectures and evaluation benchmarks for DCNv2 modules.
  • Paper: Feature Pyramid Networks for Object Detection, Tsung-Yi Lin et al. (2017). It establishes the Feature Pyramid Network architecture commonly used in modern detectors to combine multi-scale features with deformable convolutional operators.
  • Paper: R-FCN: Object Detection via Region-based Fully Convolutional Networks, Jifeng Dai et al. (2016). It formulates position-sensitive score maps and pooling mechanisms that motivated deformable region-of-interest operations.
  • Paper: Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks, Shaoqing Ren et al. (2015). It details the standard Faster R-CNN two-stage detection pipeline that DCNv2 integrates into and enhances.
Cover for Deformable ConvNets V2: More Deformable, Better Results

Abstract

The superior performance of Deformable Convolutional Networks arises from its ability to adapt to the geometric variations of objects. Through an examination of its adaptive behavior, we observe that while the spatial support for its neural features conforms more closely than regular ConvNets to object structure, this support may nevertheless extend well beyond the region of interest, causing features to be influenced by irrelevant image content. To address this problem, we present a reformulation of Deformable ConvNets that improves its ability to focus on pertinent image regions, through increased modeling power and stronger training. The modeling power is enhanced through a more comprehensive integration of deformable convolution within the network, and by introducing a modulation mechanism that expands the scope of deformation modeling. To effectively harness this enriched modeling capability, we guide network training via a proposed feature mimicking scheme that helps the network to learn features that reflect the object focus and classification power of R-CNN features. With the proposed contributions, this new version of Deformable ConvNets yields significant performance gains over the original model and produces leading results on the COCO benchmark for object detection and instance segmentation.

Table of Contents

  • 1 Introduction
  • 2 Analysis of Deformable ConvNet Behavior
  • 2.1 Spatial Support Visualization
  • 2.2 Spatial Support of Deformable ConvNets
  • 3 More Deformable ConvNets
  • 3.1 Stacking More Deformable Conv Layers
  • 3.2 Modulated Deformable Modules
  • 3.3 R-CNN Feature Mimicking
  • 4 Related Work
  • 5 Experiments
  • 5.1 Experiment Settings
  • 5.2 Enriched Deformation Modeling
  • 5.3 R-CNN Feature Mimicking
  • 5.4 Application on Stronger Backbones
  • 6 Conclusion
  • A1 Error-bounded Image Saliency
  • A2 DCNv2 with Various Image Resolution
  • A3 ImageNet Pre-Trained DCNv2
  • References

Knowls

  1. Knowl 1 — Modulated Deformable Convolution

    model/method

    Modulated deformable convolution extends standard convolution by learning both 2D spatial sampling offsets and continuous modulation scalar amplitudes for each kernel location. Let x(p)x(p) and y(p)y(p) denote feature vectors at 2D coordinate pp on the input and output feature maps, respectively. Given a convolution kernel of KK sampling locations with pre-specified fixed offsets {pk}k=1K\{p_k\}_{k=1}^K and weights {wk}k=1K\{w_k\}_{k=1}^K (e.g., K=9K = 9 and pk∈{(−1,−1),(−1,0),…,(1,1)}p_k \in \{(-1, -1), (-1, 0), \dots, (1, 1)\} for a 3×33 \times 3 kernel with dilation 1), modulated deformable convolution is formulated as:

    y(p)=∑k=1Kwk⋅x(p+pk+Δpk)⋅Δmky(p) = \sum_{k=1}^K w_k \cdot x(p + p_k + \Delta p_k) \cdot \Delta m_k

    where Δpk∈R2\Delta p_k \in \mathbb{R}^2 is an unconstrained, learnable 2D displacement offset and Δmk∈[0,1]\Delta m_k \in [0, 1] is a learnable modulation scalar for the kk-th sampling point. Because the displaced sampling location p+pk+Δpkp + p_k + \Delta p_k is continuous, the feature value x(p+pk+Δpk)x(p + p_k + \Delta p_k) is computed using 2D bilinear interpolation over the grid points of xx.

    Both {Δpk}k=1K\{\Delta p_k\}_{k=1}^K and {Δmk}k=1K\{\Delta m_k\}_{k=1}^K are generated by applying a separate convolutional layer over the same input feature map xx. This auxiliary convolution matches the spatial resolution, kernel dimensions, and dilation of the parent convolution. It outputs 3K3K feature channels: the first 2K2K channels specify the 2D coordinate offsets {Δpk}k=1K\{\Delta p_k\}_{k=1}^K, and the remaining KK channels are passed through an element-wise sigmoid activation to generate the modulation scalars {Δmk}k=1K\{\Delta m_k\}_{k=1}^K. The weights of this auxiliary convolutional layer are initialized to zero, causing initial offsets Δpk=0\Delta p_k = 0 and initial modulation values Δmk=0.5\Delta m_k = 0.5. The learning rate for the added convolutional layer is set to 0.1×0.1\times that of the surrounding network layers.

  2. Knowl 2 — Modulated Deformable RoIPooling

    model/method

    Modulated deformable RoIPooling extends region of interest pooling by introducing learnable spatial offsets and bin modulation factors. Given an input Region of Interest (RoI), the RoI is partitioned into KK spatial bins (e.g., a 7×77 \times 7 grid where K=49K=49). Within the kk-th bin, an evenly spaced sampling grid containing nkn_k points {pkj}j=1nk\{p_{kj}\}_{j=1}^{n_k} is defined (e.g., 2×22 \times 2 grid cells, nk=4n_k=4). The pooled output feature vector y(k)y(k) for the kk-th bin is defined as:

    y(k)=∑j=1nkx(pkj+Δpk)⋅Δmknky(k) = \sum_{j=1}^{n_k} x(p_{kj} + \Delta p_k) \cdot \frac{\Delta m_k}{n_k}

    where xx represents the input feature maps, Δpk∈R2\Delta p_k \in \mathbb{R}^2 is a continuous learnable offset vector for the kk-th bin, and Δmk∈[0,1]\Delta m_k \in [0, 1] is a learnable modulation scalar. Features at fractional coordinates pkj+Δpkp_{kj} + \Delta p_k are evaluated via 2D bilinear interpolation.

    The offsets and modulation scalars are predicted by an auxiliary branch. The branch applies standard RoIPooling over the RoI on the input feature maps, followed by two 1024-dimensional fully connected (fcfc) layers (initialized from a Gaussian distribution with standard deviation 0.010.01). A subsequent fcfc layer (initialized with zero weights) produces 3K3K output channels. The first 2K2K channels encode normalized offsets, which are scaled via element-wise multiplication by the RoI's width and height to yield {Δpk}k=1K\{\Delta p_k\}_{k=1}^K. The remaining KK channels pass through a sigmoid activation to yield modulation scalars {Δmk}k=1K∈[0,1]\{\Delta m_k\}_{k=1}^K \in [0, 1].

  3. Knowl 3 — R-CNN Feature Mimicking for Deformable Faster R-CNN

    model/method

    To prevent per-RoI detection features from incorporating non-discriminative background context outside the region of interest, network training incorporates an auxiliary R-CNN teacher branch that guides feature learning via knowledge distillation.

    During stochastic gradient descent (SGD) training, 32 positive region proposals (denoted as Ω\Omega) that sufficiently overlap with ground-truth objects are sampled per image. For each RoI b∈Ωb \in \Omega, the corresponding image patch is cropped and resized to 224×224224 \times 224 pixels. An R-CNN branch processes this patch using a shared backbone network to output a 14×1414 \times 14 feature map. A (modulated) deformable RoIPooling layer covering the entire 224×224224 \times 224 region is applied, followed by two 1024-dimensional fcfc layers, yielding an R-CNN feature representation fRCNN(b)∈R1024f_{\text{RCNN}}(b) \in \mathbb{R}^{1024} and class prediction logits via a (C+1)(C+1)-way Softmax classifier (where CC is the number of foreground categories).

    The feature mimic loss LmimicL_{\text{mimic}} is defined as the cosine distance between fRCNN(b)f_{\text{RCNN}}(b) and the counterpart feature vector fFRCNN(b)∈R1024f_{\text{FRCNN}}(b) \in \mathbb{R}^{1024} produced by the Fast R-CNN head:

    Lmimic=∑b∈Ω[1−cos⁡(fRCNN(b),fFRCNN(b))]L_{\text{mimic}} = \sum_{b \in \Omega} [1 - \cos(f_{\text{RCNN}}(b), f_{\text{FRCNN}}(b))]

    Network training is jointly driven by LmimicL_{\text{mimic}}, the R-CNN classification cross-entropy loss over Ω\Omega, and the standard Faster R-CNN loss terms. Both newly introduced losses are assigned a loss weight of 0.1 relative to the original Faster R-CNN losses. Parameters for the backbone, deformable RoIPooling, and 2fc heads are shared between both branches (while the classification heads remain unshared). Feature mimicking is strictly omitted for negative (background) RoIs, as negative regions require broader context to prevent false positives. During inference, the auxiliary R-CNN branch is discarded, introducing zero computation or latency overhead.

  4. Knowl 4 — Error-Bounded Saliency Optimization for Visual Support Region Estimation

    algorithm

    To quantify the true spatial support of deep network nodes without heuristic regularization trade-offs, the visual support region is formulated as the smallest subset of image pixels whose masked content preserves the network output within a strict reconstruction error bound ϵ\epsilon:

    min⁡∥M∥1s.t.Lrec(N(I),N(I⊙M))<ϵ\min \|M\|_1 \quad \text{s.t.} \quad L_{\text{rec}}(\mathcal{N}(I), \mathcal{N}(I \odot M)) < \epsilon

    where II is the input image, N(I)\mathcal{N}(I) is the network feature or response on the unmasked image, M∈{0,1}H×WM \in \{0, 1\}^{H \times W} is a binary spatial mask, I⊙MI \odot M denotes pixel-wise masking (setting unselected pixels to zero), and LrecL_{\text{rec}} is a reconstruction loss function (defined as 1−cos⁡(N(I),N(I⊙M))1 - \cos(\mathcal{N}(I), \mathcal{N}(I \odot M))). To solve this constrained optimization problem, a two-step greedy heuristic search is employed.

    Input: Input image II, trained network N\mathcal{N}, target node response function N(⋅)\mathcal{N}(\cdot), error bound threshold ϵ=0.1\epsilon = 0.1, target node spatial coordinate p0p_0
    Output: Binary visual support mask MM
    Step 1: Rectangle Bound Search
    Initialize rectangular region RR centered at p0p_0 with area 0
    loop
        Define mask MRM_R such that MR(p)=1M_R(p) = 1 if p∈Rp \in R else 00
        if Lrec(N(I),N(I⊙MR))<ϵL_{\text{rec}}(\mathcal{N}(I), \mathcal{N}(I \odot M_R)) < \epsilon then
            break loop
        else
            Enlarge RR by uniform area increments (square shape for full-image feature nodes, or matching RoI aspect ratio for RoI nodes)
        end if
    end loop
    Step 2: Super-Pixel Greedy Pruning
    Segment image content inside RR into super-pixels S={s1,s2,…,sN}\mathcal{S} = \{s_1, s_2, \dots, s_N\} via SLIC algorithm
    Initialize active super-pixel set A←S\mathcal{A} \leftarrow \mathcal{S}
    Initialize mask MM such that M(p)=1M(p) = 1 for all p∈⋃s∈Asp \in \bigcup_{s \in \mathcal{A}} s, and 00 elsewhere
    loop
        Identify s∗=arg⁡min⁡s∈ALrec(N(I),N(I⊙MA∖{s}))s^* = \arg\min_{s \in \mathcal{A}} L_{\text{rec}}(\mathcal{N}(I), \mathcal{N}(I \odot M_{\mathcal{A} \setminus \{s\}}))
        if Lrec(N(I),N(I⊙MA∖{s∗}))<ϵL_{\text{rec}}(\mathcal{N}(I), \mathcal{N}(I \odot M_{\mathcal{A} \setminus \{s^*\}})) < \epsilon then
            A←A∖{s∗}\mathcal{A} \leftarrow \mathcal{A} \setminus \{s^*\}
            Update MM by setting M(p)←0M(p) \leftarrow 0 for all p∈s∗p \in s^*
        else
            break loop
        end if
    end loop
    return MM
  5. Knowl 5 — Ablation of Deformation Stacking and Modulation on COCO

    data/table

    The effectiveness of progressively replacing standard 3×33 \times 3 convolutions with deformable counterparts across network stages (conv3 to conv5) and adding feature amplitude modulation is evaluated on the COCO 2017 validation set using ResNet-50 backbones at an input resolution of shorter side 1,000 pixels. In the table, dconv and dpool denote deformable convolution and deformable RoIPooling from DCNv1, while mdconv and mdpool denote modulated deformable convolution and modulated deformable RoIPooling.

    Method Setting Faster R-CNN Params / FLOP Mask R-CNN
    APbbox\text{AP}^{\text{bbox}} APSbbox\text{AP}^{\text{bbox}}_{S} APMbbox\text{AP}^{\text{bbox}}_{M} APLbbox\text{AP}^{\text{bbox}}_{L} Param FLOP APbbox\text{AP}^{\text{bbox}} APmask\text{AP}^{\text{mask}}
    Baseline regular (RoIPooling) 32.1 14.9 37.5 44.4 51.3M 326.7G - -
    regular (aligned RoIPooling) 34.7 19.3 39.5 45.3 51.3M 326.7G 36.6 32.2
    dconv@c5 + dpool (DCNv1) 38.0 20.7 41.8 52.2 52.7M 328.2G 40.4 35.3
    Enriched
    Deformation
    dconv@c5 37.4 20.0 40.9 51.0 51.5M 327.1G 40.2 35.1
    dconv@c4∼\simc5 40.0 21.4 43.8 55.3 51.7M 328.6G 41.8 36.8
    dconv@c3∼\simc5 40.4 21.6 44.2 56.2 51.8M 330.6G 42.2 37.0
    dconv@c3∼\simc5 + dpool 41.0 22.0 45.1 56.6 53.0M 331.8G 42.4 37.0
    mdconv@c3∼\simc5 + mdpool 41.7 22.2 45.8 58.7 65.5M 346.2G 43.1 37.3

    Progressively stacking deformable layers from conv5 across conv4 and conv3 yields steady accuracy gains of 2.0%2.0\% to 3.0%3.0\% in APbbox\text{AP}^{\text{bbox}} and APmask\text{AP}^{\text{mask}}. Upgrading standard deformable modules to modulated deformable modules provides an additional 0.3%0.3\% to 0.7%0.7\% improvement with modest computational overhead.

  6. Knowl 6 — COCO 2017 Test-Dev Benchmark Across Stronger Backbones

    data/table

    Object detection and instance segmentation performance of regular ConvNets, DCNv1, and DCNv2 are compared on the COCO 2017 test-dev benchmark across ResNet-50, ResNet-101, and ResNeXt-101 backbones. For DCNv1, deformable convolutions replace 3×33 \times 3 conv layers in conv5, and deformable RoIPooling is used. For DCNv2, all 3×33 \times 3 conv layers in stages conv3 through conv5 are replaced by modulated deformable convolutions, modulated deformable RoIPooling is used, and training incorporates R-CNN feature mimicking.

    Backbone Method Faster R-CNN Mask R-CNN
    APbbox\text{AP}^{\text{bbox}} (%) APbbox\text{AP}^{\text{bbox}} (%) APmask\text{AP}^{\text{mask}} (%)
    ResNet-50 regular 35.1 37.0 32.4
    DCNv1 38.4 40.7 35.5
    DCNv2 43.3 44.5 38.4
    ResNet-101 regular 39.2 40.9 35.3
    DCNv1 41.4 42.9 37.1
    DCNv2 44.8 45.8 39.7
    ResNeXt-101 regular 40.1 41.7 36.2
    DCNv1 41.7 43.4 37.7
    DCNv2 45.3 46.7 40.5

    DCNv2 consistently outperforms regular ConvNets by 4.9%4.9\% to 8.2%8.2\% in APbbox\text{AP}^{\text{bbox}} and outperforms DCNv1 by 3.4%3.4\% to 4.9%4.9\% across all backbone architectures.

  7. Knowl 7 — Impact of Region Selection for R-CNN Feature Mimicking

    data/table

    The choice of RoI regions supervised by the R-CNN feature mimicking loss was evaluated on the COCO 2017 validation set for both modulated deformable models (mdconv3~5 + mdpool) and regular ConvNets using ResNet-50 backbones. 'FG' denotes foreground (positive) RoIs, and 'BG' denotes background (negative) RoIs.

    Network Setting Regions to Mimic Faster R-CNN Mask R-CNN
    APbbox\text{AP}^{\text{bbox}} (%) APbbox\text{AP}^{\text{bbox}} (%) APmask\text{AP}^{\text{mask}} (%)
    mdconv3∼\sim5 + mdpool None 41.7 43.1 37.3
    FG BG 42.1 43.4 37.6
    BG Only 41.7 43.3 37.5
    FG Only 43.1 44.3 38.3
    regular None 34.7 36.6 32.2
    FG Only 35.0 36.8 32.3

    Enforcing feature mimicking on foreground-only RoIs provides a gain of 1.4%1.4\% APbbox\text{AP}^{\text{bbox}} on Faster R-CNN and 1.2%1.2\% APbbox\text{AP}^{\text{bbox}} / 1.0%1.0\% APmask\text{AP}^{\text{mask}} on Mask R-CNN. Mimicking negative background RoIs fails to improve accuracy because background regions require broader contextual features rather than focused local features. Applying feature mimicking to regular ConvNets yields virtually no gain (+0.3%+0.3\% APbbox\text{AP}^{\text{bbox}}) due to regular convolution's inability to dynamically reshape spatial support.

  8. Knowl 8 — ImageNet Pretraining and Downstream Transfer of Deformable Parameters

    data/table

    Pretraining the learnable offsets and modulation scalars of DCNv2 directly on ImageNet-1K classification is evaluated alongside downstream transfer performance on PASCAL VOC object detection, PASCAL VOC semantic segmentation, ImageNet VID video object detection, and COCO object detection using a ResNet-101 backbone.

    Backbone Method Top-1 Acc (%) Top-5 Acc (%) Param FLOP
    ResNet-50 regular 76.5 93.1 26.6M 4.1G
    DCNv1 76.6 93.2 26.8M 4.1G
    DCNv2 78.2 94.0 27.4M 4.3G
    ResNet-101 regular 78.4 94.2 45.5M 7.8G
    DCNv1 78.4 94.2 45.8M 7.8G
    DCNv2 79.2 94.6 47.4M 8.2G
    ResNeXt-101 regular 78.8 94.4 45.1M 8.0G
    DCNv1 78.9 94.4 45.6M 8.0G
    DCNv2 79.8 94.8 49.0M 8.7G
    Method Offset Mod. VOC Detection VOC Seg. ImageNet VID COCO Det.
    Pretraining AP50bbox\text{AP}^{\text{bbox}}_{50} AP70bbox\text{AP}^{\text{bbox}}_{70} mIoU (%) APbbox\text{AP}^{\text{bbox}} (%) APbbox\text{AP}^{\text{bbox}} (%)
    regular None 81.9 68.2 72.0 74.9 39.2
    DCNv2 None (scratch) 83.7 72.4 76.1 79.2 44.8
    DCNv2 ImageNet 84.9 73.5 78.3 80.7 44.9

    Pretraining deformable parameters on ImageNet yields substantial improvements on smaller target datasets (VOC detection +1.2%+1.2\%, VOC segmentation +2.2%+2.2\%, ImageNet VID +1.5%+1.5\%). On COCO, pretraining shows negligible benefit (+0.1%+0.1\%) because COCO is sufficiently large to learn offsets and modulation parameters from scratch.

  9. Knowl 9 — Resolution Robustness and Multi-Scale Evaluation in DCNv2

    empirical result

    When evaluated on Faster R-CNN (ResNet-50 and ResNet-101) across input image resolutions with shorter sides ranging over {400,600,800,1000,1200,1400}\{400, 600, 800, 1000, 1200, 1400\} pixels on COCO 2017 test-dev:

    1. Both regular ConvNets and DCNv2 achieve their highest single-scale detection scores at a shorter side of 1,000 pixels (e.g., DCNv2 ResNet-101 achieves 44.8%44.8\% APbbox\text{AP}^{\text{bbox}} at 1,000 pixels vs 44.0%44.0\% at 800 pixels).
    2. As the shorter side increases beyond 1,000 pixels (up to 1,400 pixels), the APbbox\text{AP}^{\text{bbox}} of regular ConvNets drops significantly (especially for medium and large objects, APMbbox\text{AP}^{\text{bbox}}_M and APLbbox\text{AP}^{\text{bbox}}_L), because standard fixed receptive fields cover only a small fraction of enlarged objects. In contrast, DCNv2 maintains nearly constant performance across high resolutions due to its ability to dynamically adapt its spatial support.
    3. Multi-scale testing (evaluating image shorter sides from 400 to 1400 pixels with step size 200) further boosts DCNv2 ResNet-101 performance from 44.8%44.8\% to 46.0%46.0\% APbbox\text{AP}^{\text{bbox}} (+1.2%+1.2\% gain).
  10. Knowl 10 — Spatial Support Visualization Modalities for Analyzing Deformable Networks

    definition

    To analyze how neural network feature units respond to underlying image regions, three distinct and complementary visualization modalities are defined:

    1. Effective Receptive Field (ERF): Measures the relative contribution of individual image pixels to a network node's activation, computed as the gradient of the node's scalar response with respect to intensity perturbations at each input pixel.
    2. Effective Sampling / Bin Locations: Measures the contribution strength of individual learned sampling locations (in deformable convolutions) or pooling bins (in deformable RoIPooling), computed as the gradient of the network node response with respect to the sampling/bin coordinates.
    3. Error-Bounded Saliency Region: Identifies the minimal structured visual image region required to reconstruct the node's original full-image activation within a fixed relative error bound ϵ\epsilon.

    Evaluating sampling coordinates in isolation without gradient-weighted influence can lead to misleading conclusions, as regular networks with fixed grid coordinates adapt spatial support through weights, while deformable network predictions are jointly determined by learned offsets and convolutional weights.

Coverage note — All substantial contributions—including modulated deformable convolution, modulated deformable RoIPooling, R-CNN feature mimicking, the error-bounded saliency algorithm, spatial support analysis modalities, ImageNet pre-training, and multi-scale/resolution benchmark results across COCO, VOC, and ImageNet VID—have been captured as self-contained knowls.

References

  1. 1.R. Achanta, A. Shaji, K. Smith, A. Lucchi, P. Fua, S. Süsstrunk, et al. Slic superpixels compared to state-of-the-art superpixel methods. IEEE transactions on pattern analysis and machine intelligence, 34(11):2274–2282, 2012. 9
  2. 2.J. Ba and R. Caruana. Do deep nets really need to be deep? In NIPS, 2014. 2, 5, 7
  3. 3.P. Battaglia, R. Pascanu, M. Lai, D. J. Rezende, et al. Interaction networks for learning about objects, relations and physics. In NIPS, 2016. 6
  4. 4.D. Britz, A. Goldie, M.-T. Luong, and Q. Le. Massive exploration of neural machine translation architectures. In EMNLP, 2017. 6
  5. 5.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. arXiv preprint arXiv:1606.00915, 2016. 7, 11
  6. 6.B. Cheng, Y. Wei, H. Shi, R. Feris, J. Xiong, and T. Huang. Revisiting rcnn: On awakening the classification power of faster rcnn. In ECCV, 2018. 3, 5
  7. 7.P. Dabkowski and Y. Gal. Real time image saliency for black box classifiers. In NIPS, 2017. 2, 7, 9
  8. 8.J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei. Deformable convolutional networks. In ICCV, 2017. 1, 2, 3, 4, 5, 8, 11
  9. 9.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009. 7, 11
  10. 10.M. Denil, S. G. Colmenarejo, S. Cabi, D. Saxton, and N. de Freitas. Programmable agents. arXiv preprint arXiv:1706.06383, 2017. 6
  11. 11.M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman. The PASCAL Visual Object Classes (VOC) Challenge. IJCV, 2010. 1
  12. 12.P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan. Object detection with discriminatively trained part-based models. TPAMI, 2010. 6
  13. 13.R. C. Fong and A. Vedaldi. Interpretable explanations of black boxes by meaningful perturbation. In ICCV, 2017. 2, 7, 9
  14. 14.J. Gehring, M. Auli, D. Grangier, and Y. N. Dauphin. A convolutional encoder model for neural machine translation. In ACL, 2017. 6
  15. 15.J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin. Convolutional sequence to sequence learning. arXiv preprint arXiv:1705.03122, 2017. 6
  16. 16.R. Girshick. Fast R-CNN. In ICCV, 2015. 1, 2
  17. 17.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, 2014. 2
  18. 18.R. Girshick, I. Radosavovic, G. Gkioxari, P. Dollár, and K. He. Detectron. https://github.com/facebookresearch/detectron, 2018. 7, 11
  19. 19.J. Gu, H. Hu, L. Wang, Y. Wei, and J. Dai. Learning region features for object detection. In ECCV, 2018. 6
  20. 20.K. He, G. Gkioxari, P. Dollár, and R. Girshick. Mask r-cnn. In ICCV, 2017. 2, 7
  21. 21.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016. 2, 10
  22. 22.G. Hinton, O. Vinyals, and J. Dean. Distilling the knowledge in a neural network. STAT, 2015. 2, 5, 7
  23. 23.Y. Hoshen. Vain: Attentional multi-agent predictive modeling. In NIPS, 2017. 6
  24. 24.H. Hu, J. Gu, Z. Zhang, J. Dai, and Y. Wei. Relation networks for object detection. In CVPR, 2018. 6
  25. 25.M. Jaderberg, K. Simonyan, A. Zisserman, and K. Kavukcuoglu. Spatial transformer networks. In NIPS, 2015. 6
  26. 26.Y. Jeon and J. Kim. Active convolution: Learning the shape of convolution for image classification. In CVPR, 2017. 7
  27. 27.B. Lee, E. Erdenee, S. Jin, M. Y. Nam, Y. G. Jung, and P. K. Rhee. Multi-class multi-object tracking using changing point detection. In ECCV, 2016. 11
  28. 28.Q. Li, S. Jin, and J. Yan. Mimicking very efficient network for object detection. In CVPR, 2017. 5, 7
  29. 29.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 1
  30. 30.D. G. Lowe. Object recognition from local scale-invariant features. In ICCV, 1999. 6
  31. 31.W. Luo, Y. Li, R. Urtasun, and R. Zemel. Understanding the effective receptive field in deep convolutional neural networks. arXiv preprint arXiv:1701.04128, 2017. 2, 7
  32. 32.D. Raposo, A. Santoro, D. Barrett, R. Pascanu, T. Lillicrap, and P. Battaglia. Discovering objects and their relations from entangled scene representations. In ICLR, 2017. 6
  33. 33.S. Ren, K. He, R. Girshick, and J. Sun. Faster R-CNN: Towards real-time object detection with region proposal networks. In NIPS, 2015. 2, 10
  34. 34.E. Rublee, V. Rabaud, K. Konolige, and G. Bradski. Orb: an efficient alternative to sift or surf. In ICCV, 2011. 6
  35. 35.A. Santoro, D. Raposo, D. G. Barrett, M. Malinowski, R. Pascanu, P. Battaglia, and T. Lillicrap. A simple neural network module for relational reasoning. In NIPS, 2017. 6
  36. 36.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. Attention is all you need. In NIPS, 2017. 6
  37. 37.X. Wang, R. Girshick, A. Gupta, and K. He. Non-local neural networks. In CVPR, 2018. 6
  38. 38.N. Watters, D. Zoran, T. Weber, P. Battaglia, R. Pascanu, and A. Tacchetti. Visual interaction networks. In NIPS, 2017. 6
  39. 39.S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He. Aggregated residual transformations for deep neural networks. In CVPR, 2017. 9, 11
  40. 40.S. Zagoruyko, A. Lerer, T.-Y. Lin, P. H. Pinheiro, S. Gross, S. Chintala, and P. Dollar. A multipath network for object detection. In BMVC, 2016. 7
  41. 41.B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba. Learning deep features for discriminative localization. In CVPR, 2016. 2, 7, 9
  42. 42.X. Zhu, Y. Wang, J. Dai, L. Yuan, and Y. Wei. Flow-guided feature aggregation for video object detection. In ICCV, 2017. 11
  43. 43.X. Zhu, Y. Xiong, J. Dai, L. Yuan, and Y. Wei. Deep feature flow for video recognition. In CVPR, 2017. 11
  44. 44.L. M. Zintgraf, T. S. Cohen, T. Adel, and M. Welling. Visualizing deep neural network decisions: Prediction difference analysis. In ICLR, 2017. 2, 7, 9

Citation

MLA
Zhu, X., et al. “Deformable ConvNets V2: More Deformable, Better Results”. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 9300–08, https://doi.org/10.1109/CVPR.2019.00953.
APA
Zhu, X., Hu, H., Lin, S., & Dai, J. (2019). Deformable ConvNets V2: More Deformable, Better Results. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9300–9308. https://doi.org/10.1109/CVPR.2019.00953
Chicago
Zhu, X., H. Hu, S. Lin, and J. Dai. 2019. “Deformable ConvNets V2: More Deformable, Better Results”. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9300–9308. https://doi.org/10.1109/CVPR.2019.00953.
Harvard
Zhu, X. et al. (2019) “Deformable ConvNets V2: More Deformable, Better Results”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 9300–9308. Available at: https://doi.org/10.1109/CVPR.2019.00953.
Vancouver
1. Zhu X, Hu H, Lin S, Dai J (2019) Deformable ConvNets V2: More Deformable, Better Results. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 9300–9308

BibTeX

@inproceedings{Zhu_2019, title={Deformable ConvNets V2: More Deformable, Better Results}, url={http://dx.doi.org/10.1109/CVPR.2019.00953}, DOI={10.1109/cvpr.2019.00953}, booktitle={2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Zhu, Xizhou and Hu, Han and Lin, Stephen and Dai, Jifeng}, year={2019}, month=June, pages={9300–9308} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE