Dilated Residual Networks
Fisher YuVladlen KoltunThomas Funkhouser
Proposes dilated residual networks that retain high spatial feature resolution without increasing model complexity, introducing a degridding technique that improves performance across image classification, object localization, and semantic segmentation.
Standard computer vision architectures typically downsample images until critical spatial details are lost, reducing their feature maps to a tiny fraction of the original image area. While this aggressive reduction simplifies basic classification, it impairs the recognition of small or thin objects and makes models difficult to transfer to complex tasks requiring precise spatial awareness, such as scene segmentation and object localization. This article evaluates whether preserving higher spatial resolution throughout deep neural networks improves both standard image classification and downstream spatial reasoning tasks.
To test this, the researchers modified standard residual neural networks by replacing interior downsampling steps with dilated convolutions—a technique that expands the spatial resolution of output feature layers without reducing the receptive field of individual neurons. The resulting dilated residual network architecture limits downsampling to an 8-fold reduction rather than the traditional 32-fold reduction, producing output feature maps with 16 times greater spatial resolution. The authors also developed an architectural refinement to eliminate artificial grid-like distortions caused by dilation, replacing early pooling operations with convolutional filters and appending layers with decreasing dilation. The evaluation compared standard residual networks against the dilated designs across the ImageNet benchmark for classification and localization, as well as the Cityscapes dataset for urban scene semantic segmentation.
Key findings show that dilated networks consistently outperform standard networks across all evaluated benchmarks. In standard image classification on ImageNet, dilated models reduced error rates without adding depth or complexity; for example, an 18-layer dilated network reduced top-1 single-crop classification error from 30.43% to 28.00%, while a degridded 42-layer model achieved an error rate of 22.94%, rivaling a standard 101-layer network that is over twice as deep. In weakly-supervised object localization, the high-resolution activation maps enabled direct localization without any fine-tuning, with a 26-layer degridded dilated model achieving a top-1 error of 52.3%, significantly outperforming the standard 101-layer model's 54.6% error. In semantic segmentation on the Cityscapes dataset, a 42-layer degridded dilated network attained an overall mean score of 70.9%, outperforming the standard 101-layer baseline of 66.6% by more than four percentage points.
These results demonstrate that deep networks do not need to discard spatial details to achieve high classification accuracy. Preserving spatial acuity allows organizations and engineering teams to deploy shallower, less complex models that simultaneously achieve superior accuracy on classification, localization, and segmentation tasks. By eliminating the need for post-hoc recovery mechanisms like complex skip connections or up-convolutions, dilated networks streamline the development lifecycle and reduce structural overhead in vision pipelines.
Organizations developing computer vision systems should adopt dilated residual architectures as a baseline framework when building applications involving complex scene understanding. Teams should implement the degridded architecture rather than basic dilation to prevent output artifacts and maximize segmentation quality. Future engineering work should evaluate memory optimization techniques, as maintaining higher-resolution feature maps increases memory consumption during processing. Overall, the extensive experimental validation on standard benchmark datasets provides high confidence in the performance benefits of dilated residual networks.
- Paper: Deep Residual Learning for Image Recognition, Kaiming He et al. (2016). Read the original ResNet paper first to understand the residual backbone that Dilated Residual Networks modifies.
- Paper: Multi-Scale Context Aggregation by Dilated Convolutions, Fisher Yu et al. (2016). This earlier work introduces dilated convolutions for preserving resolution while enlarging receptive fields—the central operation the source adapts to residual networks.
- Paper: DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs, Liang-Chieh Chen et al. (2016). DeepLab establishes how atrous convolution adapts classification networks for dense segmentation, setting up the source’s investigation of dilation in residual architectures.
- Paper: Understanding Convolution for Semantic Segmentation, Panqu Wang et al. (2017). Its hybrid dilated convolutions directly pursue the source’s goal of reducing gridding artifacts while preserving spatial detail in segmentation.
