Dilated Residual Networks

Fisher YuVladlen KoltunThomas Funkhouser

article2017CVPR1,835 citations

Proposes dilated residual networks that retain high spatial feature resolution without increasing model complexity, introducing a degridding technique that improves performance across image classification, object localization, and semantic segmentation.

Listen

Standard computer vision architectures typically downsample images until critical spatial details are lost, reducing their feature maps to a tiny fraction of the original image area. While this aggressive reduction simplifies basic classification, it impairs the recognition of small or thin objects and makes models difficult to transfer to complex tasks requiring precise spatial awareness, such as scene segmentation and object localization. This article evaluates whether preserving higher spatial resolution throughout deep neural networks improves both standard image classification and downstream spatial reasoning tasks.

To test this, the researchers modified standard residual neural networks by replacing interior downsampling steps with dilated convolutions—a technique that expands the spatial resolution of output feature layers without reducing the receptive field of individual neurons. The resulting dilated residual network architecture limits downsampling to an 8-fold reduction rather than the traditional 32-fold reduction, producing output feature maps with 16 times greater spatial resolution. The authors also developed an architectural refinement to eliminate artificial grid-like distortions caused by dilation, replacing early pooling operations with convolutional filters and appending layers with decreasing dilation. The evaluation compared standard residual networks against the dilated designs across the ImageNet benchmark for classification and localization, as well as the Cityscapes dataset for urban scene semantic segmentation.

Key findings show that dilated networks consistently outperform standard networks across all evaluated benchmarks. In standard image classification on ImageNet, dilated models reduced error rates without adding depth or complexity; for example, an 18-layer dilated network reduced top-1 single-crop classification error from 30.43% to 28.00%, while a degridded 42-layer model achieved an error rate of 22.94%, rivaling a standard 101-layer network that is over twice as deep. In weakly-supervised object localization, the high-resolution activation maps enabled direct localization without any fine-tuning, with a 26-layer degridded dilated model achieving a top-1 error of 52.3%, significantly outperforming the standard 101-layer model's 54.6% error. In semantic segmentation on the Cityscapes dataset, a 42-layer degridded dilated network attained an overall mean score of 70.9%, outperforming the standard 101-layer baseline of 66.6% by more than four percentage points.

These results demonstrate that deep networks do not need to discard spatial details to achieve high classification accuracy. Preserving spatial acuity allows organizations and engineering teams to deploy shallower, less complex models that simultaneously achieve superior accuracy on classification, localization, and segmentation tasks. By eliminating the need for post-hoc recovery mechanisms like complex skip connections or up-convolutions, dilated networks streamline the development lifecycle and reduce structural overhead in vision pipelines.

Organizations developing computer vision systems should adopt dilated residual architectures as a baseline framework when building applications involving complex scene understanding. Teams should implement the degridded architecture rather than basic dilation to prevent output artifacts and maximize segmentation quality. Future engineering work should evaluate memory optimization techniques, as maintaining higher-resolution feature maps increases memory consumption during processing. Overall, the extensive experimental validation on standard benchmark datasets provides high confidence in the performance benefits of dilated residual networks.

Cover for Dilated Residual Networks

Abstract

Convolutional networks for image classification progressively reduce resolution until the image is represented by tiny feature maps in which the spatial structure of the scene is no longer discernible. Such loss of spatial acuity can limit image classification accuracy and complicate the transfer of the model to downstream applications that require detailed scene understanding. These problems can be alleviated by dilation, which increases the resolution of output feature maps without reducing the receptive field of individual neurons. We show that dilated residual networks (DRNs) outperform their non-dilated counterparts in image classification without increasing the model's depth or complexity. We then study gridding artifacts introduced by dilation, develop an approach to removing these artifacts (`degridding'), and show that this further increases the performance of DRNs. In addition, we show that the accuracy advantage of DRNs is further magnified in downstream applications such as object localization and semantic segmentation.

Table of Contents

  • 1 Introduction
  • 2 Dilated Residual Networks
  • 3 Localization
  • 4 Degridding
  • 5 Experiments
  • 5.1 Image Classification
  • 5.2 Object Localization
  • 5.3 Semantic Segmentation
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Dilated Residual Network Architecture (DRN-A)

    model/method

    Standard convolutional residual networks (ResNets) progressively downsample internal representations across five layer groups (G1,…,G5G^1, \dots, G^5) by a factor of 32 in each dimension using strided convolutions at the first layer of each group. Dilated Residual Networks (DRN-A) preserve spatial resolution without decreasing neuron receptive fields by substituting striding in the final two groups with dilated convolutions.

    Let GiℓG_i^\ell denote the ii-th convolutional layer in group ℓ∈{1,…,5}\ell \in \{1, \dots, 5\}. The DRN-A construction modifies groups G4G^4 and G5G^5 as follows:

    1. Removing Subsampling: Striding is eliminated in the first layer of group 4 (G14G_1^4) and group 5 (G15G_1^5). This retains the spatial resolution of group 3 (G3G^3) throughout the remaining network (reducing the overall downsampling factor from 32 to 8, yielding an output feature map of resolution 28×2828 \times 28 for a standard 224×224224 \times 224 input).

    2. Dilation in Group 4: To compensate for the factor-of-2 receptive field reduction caused by removing striding in G14G_1^4, all subsequent convolutions in group 4 (Gi4G_i^4 for i≥2i \ge 2) and the first layer of group 5 (G15G_1^5) are replaced by 2-dilated convolutions (d=2d = 2).

    3. Dilation in Group 5: Because subsequent layers in group 5 follow two eliminated striding operations, their receptive fields would otherwise decrease by a factor of 4. Therefore, all layers Gi5G_i^5 for i≥2i \ge 2 are replaced by 4-dilated convolutions (d=4d = 4).

    Following group G5G^5, global average pooling reduces the 28×2828 \times 28 feature map to a vector before a final 1×11 \times 1 convolution computes prediction scores across classes. DRN-A retains the exact parameter count and depth of the corresponding baseline ResNet while preserving 16 times higher spatial area in output feature maps.

  2. Knowl 2 — Degridding Techniques for Dilated Residual Networks (DRN-B and DRN-C)

    model/method

    Dilated convolutions can produce high-frequency gridding artifacts when a feature map contains frequency content higher than the sampling rate of the dilated filter. To eliminate these artifacts, two successive structural transformations are applied to DRN-A:

    1. DRN-B (Intermediate Architecture):

    • Removal of Early Max Pooling: Standard ResNets apply a max pooling layer (stride 2) following the initial 7×77 \times 7 convolution, which introduces high-amplitude, high-frequency activations that propagate and exacerbate gridding in deeper layers. DRN-B replaces this max pooling layer with two residual blocks (levels 1 and 2): level 1 comprises two 3×33 \times 3 convolutions with 16 channels, and level 2 comprises two 3×33 \times 3 convolutions with 32 channels and stride 2.
    • Addition of Post-Processing Residual Layers: After the 4-dilated layers in level 6, DRN-B adds two residual blocks with decreasing dilation: level 7 consists of two 2-dilated 3×33 \times 3 convolutions (512 channels), and level 8 consists of two 1-dilated 3×33 \times 3 convolutions (512 channels), functioning as anti-aliasing filters.

    2. DRN-C (Final Degridded Architecture):

    • Removal of Residual Connections in Levels 7 and 8: While DRN-B introduces filtering layers at levels 7 and 8, the identity residual skip connections in those levels bypass the filters and propagate gridding artifacts from level 6 directly to the output. DRN-C removes the residual skip connections in levels 7 and 8, converting them into plain convolutional blocks. This forces representations through the decreasing-dilation convolutions, eliminating gridding artifacts from the final activations.
  3. Knowl 3 — Dense Class Activation Mapping for Weakly-Supervised Localization

    model/method

    A Dilated Residual Network trained strictly for whole-image classification can produce dense, high-resolution pixel-level class activation maps for object localization without retraining, fine-tuning, or structural parameter modifications.

    The localization network is constructed directly from the classification model:

    1. The global average pooling layer following the final convolutional group (G5G^5 in DRN-A or level 8 in DRN-C) is removed.

    2. The classification 1×11 \times 1 convolutional operator KK, which maps cc feature channels to nn target categories, is connected directly to the output of the final convolutional group.

    3. Applying KK directly to the spatial feature map of dimension W×HW \times H (e.g., 28×2828 \times 28) yields an unnormalized score volume of dimension n×W×Hn \times W \times H.

    4. A softmax operation is evaluated across the nn class channels independently at every spatial coordinate (w,h)(w, h). For each category yy, the resulting activation map provides the probability f(y,w,h)f(y, w, h) that the object observed at pixel (w,h)(w, h) belongs to category yy.

  4. Knowl 4 — Minimal Bounding Box Extraction for Weakly-Supervised Object Localization

    algorithm

    Given dense class response maps f∈RC×W×Hf \in \mathbb{R}^{C \times W \times H} produced by a classification-trained DRN, where f(c,w,h)f(c, w, h) is the response for category cc at pixel (w,h)(w, h), bounding boxes for a predicted class cic_i are extracted without training a separate bounding box regressor.

    At each spatial position (w,h)(w, h), the dominant class index is identified by: g(w,h)=arg⁡max⁡1≤c≤Cf(c,w,h)g(w, h) = \arg\max_{1 \le c \le C} f(c, w, h)

    For a target class cic_i and an activation threshold tt, the set of valid bounding boxes BiB_i comprises all rectangular coordinate intervals [w1,w2]×[h1,h2][w_1, w_2] \times [h_1, h_2] where every contained pixel is dominated by class cic_i and exceeds threshold tt: Bi={((w1,h1),(w2,h2))∣∀(w,h)∈[w1,w2]×[h1,h2],  g(w,h)=ci and f(ci,w,h)>t}B_i = \left\{ ((w_1, h_1), (w_2, h_2)) \mid \forall (w, h) \in [w_1, w_2] \times [h_1, h_2],\; g(w, h) = c_i \text{ and } f(c_i, w, h) > t \right\}

    The predicted minimal bounding box bib_i is chosen as the box in BiB_i that minimizes enclosed area: bi=arg⁡min⁡((w1,h1),(w2,h2))∈Bi(w2−w1)(h2−h1)b_i = \arg\min_{((w_1, h_1), (w_2, h_2)) \in B_i} (w_2 - w_1)(h_2 - h_1)

    Input: Dense class response volume f∈RC×W×Hf \in \mathbb{R}^{C \times W \times H}, target class index cic_i, activation threshold tt
    Output: Predicted bounding box coordinates bi=((w1,h1),(w2,h2))b_i = ((w_1, h_1), (w_2, h_2))
    for w=1w = 1 to WW do
        for h=1h = 1 to HH do
            g(w,h)←arg⁡max⁡1≤c≤Cf(c,w,h)g(w, h) \leftarrow \arg\max_{1 \le c \le C} f(c, w, h)
        end for
    end for
    Bi←∅B_i \leftarrow \emptyset
    for each w1∈{1,…,W}w_1 \in \{1, \dots, W\} and w2∈{w1,…,W}w_2 \in \{w_1, \dots, W\} do
        for each h1∈{1,…,H}h_1 \in \{1, \dots, H\} and h2∈{h1,…,H}h_2 \in \{h_1, \dots, H\} do
            is_valid←Trueis\_valid \leftarrow \text{True}
            for w=w1w = w_1 to w2w_2 do
                for h=h1h = h_1 to h2h_2 do
                    if g(w,h)≠cig(w, h) \ne c_i or f(ci,w,h)≤tf(c_i, w, h) \le t then
                        is_valid←Falseis\_valid \leftarrow \text{False}
                        break
                    end if
                end for
                if not is_validis\_valid then
                    break
                end if
            end for
            if is_validis\_valid then
                Bi←Bi∪{((w1,h1),(w2,h2))}B_i \leftarrow B_i \cup \{((w_1, h_1), (w_2, h_2))\}
            end if
        end for
    end for
    bi←arg⁡min⁡((w1,h1),(w2,h2))∈Bi(w2−w1)(h2−h1)b_i \leftarrow \arg\min_{((w_1, h_1), (w_2, h_2)) \in B_i} (w_2 - w_1)(h_2 - h_1)
    return bib_i
  5. Knowl 5 — ImageNet 2012 Classification Performance Across Network Configurations

    data/table

    The table below reports error rates on the ImageNet 2012 validation set under 1-crop (224×224224 \times 224 center crop) and 10-crop evaluation protocols. PP denotes the number of parameters.

    Model 1 crop 10 crops PP
    top-1 top-5 top-1 top-5
    ResNet-18 30.43 10.76 28.22 9.42 11.7M
    DRN-A-18 28.00 9.50 25.75 8.25 11.7M
    DRN-B-26 25.19 7.91 23.33 6.69 21.1M
    DRN-C-26 24.86 7.55 22.93 6.39 21.1M
    ResNet-34 27.73 8.74 24.76 7.35 21.8M
    DRN-A-34 24.81 7.54 22.64 6.34 21.8M
    DRN-C-42 22.94 6.57 21.20 5.60 31.2M
    ResNet-50 24.01 7.02 22.24 6.08 25.6M
    DRN-A-50 22.94 6.57 21.34 5.74 25.6M
    ResNet-101 22.44 6.21 21.08 5.35 44.5M

    These results demonstrate two primary findings:

    1. Direct dilation (DRN-A) improves classification over standard ResNets without adding any parameters or depth (e.g., DRN-A-18 reduces 1-crop top-1 error from 30.43% to 28.00%, and DRN-A-34 reduces top-1 error from 27.73% to 24.81%).
    2. Degridded architectures (DRN-C) yield further accuracy gains: DRN-C-26 matches the accuracy of the deeper DRN-A-34, and DRN-C-42 (22.94% top-1 error) matches DRN-A-50 while closely approaching ResNet-101 (22.44%) despite having 2.4 times fewer layers.
  6. Knowl 6 — Weakly-Supervised Object Localization Error Rates on ImageNet

    data/table

    The table below presents weakly-supervised object localization error rates (top-1 and top-5) on the ImageNet 2012 validation set. Localization bounding boxes are computed directly from classification-trained networks without bounding box annotations, extra parameters, or fine-tuning. A prediction is deemed correct when the classification prediction is correct and the predicted bounding box achieves an Intersection-over-Union (IoU) ≥0.5\ge 0.5 with the ground truth.

    Model top-1 top-5
    ResNet-18 61.5 59.3
    DRN-A-18 54.6 48.2
    DRN-B-26 53.8 49.3
    DRN-C-26 52.3 47.7
    ResNet-34 58.7 56.4
    DRN-A-34 55.5 50.7
    DRN-C-42 50.7 46.8
    ResNet-50 55.7 52.8
    DRN-A-50 54.0 48.4
    ResNet-101 54.6 51.9

    The evaluation demonstrates that preserving spatial resolution significantly improves localization accuracy: DRN-A-18 improves top-1 localization error over ResNet-18 by 6.9 percentage points (54.6% vs 61.5%). The degridded DRN-C-26 achieves a top-1 error of 52.3%, outperforming both DRN-A-50 (54.0%) and ResNet-101 (54.6%) despite its significantly shallower depth.

  7. Knowl 7 — Semantic Segmentation Performance of Dilated Residual Networks on Cityscapes

    data/table

    The table below reports per-class IoU and mean Intersection-over-Union (mIoU) percentages on the Cityscapes validation set (19 semantic classes). Classification-pretrained models are transferred directly to semantic segmentation without appending context aggregation modules, and predictions are upsampled via bilinear interpolation.

    Model Road Sidewalk Building Wall Fence Pole Light Sign Vegetation Terrain
    DRN-A-50 96.9 77.4 90.3 35.8 42.8 59.0 66.8 74.5 91.6 57.0
    DRN-C-26 97.4 80.7 90.4 36.1 47.0 56.9 63.8 73.0 91.2 57.9
    DRN-C-42 97.7 82.2 91.2 40.5 52.6 59.2 66.7 74.6 91.7 57.7
    Model Sky Person Rider Car Truck Bus Train Motorcycle Bicycle mean IoU
    DRN-A-50 93.4 78.7 55.3 92.1 43.2 59.5 36.2 52.0 75.2 67.3
    DRN-C-26 93.4 77.3 53.8 92.7 45.0 70.5 48.4 44.2 72.8 68.0
    DRN-C-42 94.1 79.1 56.0 93.6 56.0 74.3 54.7 50.9 74.1 70.9

    For reference, a comparable baseline setup of ResNet-101 achieves an mIoU of 66.6% on Cityscapes validation. DRN-C-26 achieves 68.0% mIoU, outperforming both the ResNet-101 baseline and the deeper DRN-A-50 (67.3% mIoU) while having 4 times fewer layers than ResNet-101. DRN-C-42 attains 70.9% mIoU, surpassing the ResNet-101 baseline by 4.3 percentage points.

  8. Knowl 8 — Definition of Dilated Discrete Convolution

    definition

    Let GG be a discrete two-dimensional feature map and let ff be a discrete convolutional filter kernel. For a dilation factor d∈N+d \in \mathbb{N}^+, the dd-dilated convolution operator (denoted ∗d\ast_d) at spatial coordinate p=(x,y)p = (x, y) is defined as:

    (G∗df)(p)=∑a+d⋅b=pG(a)f(b)(G \ast_d f)(p) = \sum_{a + d \cdot b = p} G(a) f(b)

    where aa ranges over valid spatial coordinates of the feature map GG, and bb indexes spatial coordinates within the filter ff.

    When d=1d = 1, this reduces to standard discrete convolution. When d>1d > 1, filter weights are applied at a sampling stride of dd, expanding the effective receptive field of a k×kk \times k kernel to (d(k−1)+1)×(d(k−1)+1)(d(k-1) + 1) \times (d(k-1) + 1) without increasing the number of learnable parameters or shrinking output spatial resolution.

  9. Knowl 9 — ImageNet Classification Training and Evaluation Protocol for DRNs

    experimental setup

    ImageNet classification models are trained on the 1.28-million image ImageNet 2012 training dataset (1,000 object categories) using the following parameters:

    • Optimization: Stochastic gradient descent (SGD) with momentum 0.9 and weight decay 10−410^{-4}.
    • Learning Rate Schedule: Initial learning rate of 10−110^{-1} (0.10.1), decreased by a factor of 10 every 30 epochs, trained for 120 epochs total.
    • Data Augmentation: Scale and aspect ratio augmentation combined with photometric color perturbation.
    • Validation Preprocessing: Images are resized so that the shorter dimension is 256 pixels.
    • Evaluation Modes:
      • 1-crop: Evaluated on the central 224×224224 \times 224 crop.
      • 10-crop: Evaluated on 10 crops per image (the central crop, four corner crops, and horizontal reflections of all five crops), averaging predicted class probabilities across crops.
  10. Knowl 10 — Direct Fully-Convolutional Transfer of DRNs to Semantic Segmentation

    model/method

    Traditional image classification networks downsample inputs by a factor of 32, necessitating post-hoc architectural modifications (such as learned deconvolution layers, skip connections, or auxiliary multi-scale branches) to recover resolution for semantic segmentation.

    Because classification-trained Dilated Residual Networks maintain high internal feature map resolution (downsampling factor of 8 throughout groups G3,G4,G5G^3, G^4, G^5), they can be transferred directly to dense semantic segmentation without any additional structural parameters:

    1. The global average pooling layer is removed.
    2. The network operates fully-convolutionally on arbitrary-sized input images.
    3. The classification layer outputs class score maps at 1/81/8 of the input resolution.
    4. Predictions are upsampled to the original input resolution via parameter-free bilinear interpolation.

Coverage note — None was omitted; all key contributions including DRN-A, degridding architectures (DRN-B, DRN-C), localization formulation and algorithm, semantic segmentation transfer, and empirical benchmarks on ImageNet and Cityscapes are covered.

References

  1. 1.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. arXiv:1606.00915, 2016.
  2. 2.M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele. The Cityscapes dataset for semantic urban scene understanding. In CVPR, 2016.
  3. 3.C. Galleguillos and S. J. Belongie. Context based object categorization: A critical survey. Computer Vision and Image Understanding, 114(6), 2010.
  4. 4.R. B. Girshick, J. Donahue, T. Darrell, and J. Malik. Region-based convolutional networks for accurate object detection and segmentation. PAMI, 38(1), 2016.
  5. 5.B. Hariharan, P. A. Arbeláez, R. B. Girshick, and J. Malik. Hypercolumns for object segmentation and fine-grained localization. In CVPR, 2015.
  6. 6.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016.
  7. 7.A. G. Howard. Some improvements on deep convolutional neural network based image classification. arXiv:1312.5402, 2013.
  8. 8.A. Krizhevsky, I. Sutskever, and G. E. Hinton. ImageNet classification with deep convolutional neural networks. In NIPS, 2012.
  9. 9.Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel. Backpropagation applied to handwritten zip code recognition. Neural Computation, 1(4), 1989.
  10. 10.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015.
  11. 11.H. Noh, S. Hong, and B. Han. Learning deconvolution network for semantic segmentation. In ICCV, 2015.
  12. 12.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and F. Li. ImageNet large scale visual recognition challenge. IJCV, 115(3), 2015.
  13. 13.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015.
  14. 14.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. E. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In CVPR, 2015.
  15. 15.A. Torralba, R. Fergus, and W. T. Freeman. 80 million tiny images: A large data set for nonparametric object and scene recognition. PAMI, 30(11), 2008.
  16. 16.B. Triggs. Empirical filter estimation for subpixel interpolation and matching. In ICCV, 2001.
  17. 17.P. Wang, P. Chen, Y. Yuan, D. Liu, Z. Huang, X. Hou, and G. Cottrell. Understanding convolution for semantic segmentation. arXiv:1702.08502, 2017.
  18. 18.F. Yu and V. Koltun. Multi-scale context aggregation by dilated convolutions. In ICLR, 2016.
  19. 19.B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba. Learning deep features for discriminative localization. In CVPR, 2016.

Citation

MLA
Yu, F., et al. “Dilated Residual Networks”. arXiv, 2017, http://arxiv.org/abs/1705.09914v1.
APA
Yu, F., Koltun, V., & Funkhouser, T. (2017). Dilated Residual Networks. arXiv. http://arxiv.org/abs/1705.09914v1
Chicago
Yu, F., V. Koltun, and T. Funkhouser. 2017. “Dilated Residual Networks”. arXiv. http://arxiv.org/abs/1705.09914v1.
Harvard
Yu, F., Koltun, V. and Funkhouser, T. (2017) “Dilated Residual Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1705.09914v1.
Vancouver
1. Yu F, Koltun V, Funkhouser T (2017) Dilated Residual Networks. arXiv

BibTeX

@article{yu2017dilated,
  title = {Dilated Residual Networks},
  author = {Yu, Fisher and Koltun, Vladlen and Funkhouser, Thomas},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1705.09914v1},
  eprint = {1705.09914}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE