Road Extraction by Deep Residual U-Net

Zhengxin ZhangQingjie LiuYunhong Wang

article2017IEEE Geoscience and Remote Sensing Letters2,795 citations

Proposes a deep residual U-Net architecture that integrates residual learning into semantic segmentation to extract roads from aerial imagery with fewer parameters and higher accuracy than standard U-Net models.

Listen

Road extraction from high-resolution aerial images supports applications such as navigation, urban planning, and geographic information updates, yet remains difficult because of noise, occlusions, and varied backgrounds in remote sensing data. Traditional approaches and earlier deep-learning models have produced useful but still limited results on this task.

The article set out to develop and evaluate a neural network that improves road-area segmentation accuracy while using fewer parameters than existing models. The authors created a deep residual U-Net by replacing the plain convolutional blocks in a U-Net architecture with residual units that include identity mappings and batch normalization. They trained the network on 30,000 randomly cropped 224-by-224 patches from the 1,108 training images of the Massachusetts roads dataset and tested it on the 49 held-out images, comparing performance against Mnih-CNN, Saito-CNN, and the original U-Net using relaxed precision-recall metrics with a three-pixel tolerance.

The proposed ResUnet reached a break-even point of 0.9187, exceeding U-Net (0.9053), Saito-CNN (0.9047), and the best Mnih variant (0.9006). It achieved these gains with roughly one-quarter the parameters of U-Net (7.8 million versus 30.6 million) and produced visibly cleaner segmentations that better handled two-lane roads, intersections, tree occlusions, and contextually similar features such as airport runways.

These results indicate that residual connections combined with U-Net-style skip paths can simultaneously ease training and improve information flow, yielding higher accuracy at lower computational cost for large-scale remote-sensing tasks. The approach therefore offers a practical route to more reliable automated road mapping that could reduce manual editing and accelerate updates to geographic databases.

The authors note that the model still misses some roads in parking lots when such areas are unlabeled in the training data. Additional labeled examples covering these edge cases, together with tests on imagery from different sensors or regions, would strengthen before operational deployment.

Cover for Road Extraction by Deep Residual U-Net

Abstract

Road extraction from aerial images has been a hot research topic in the field of remote sensing image analysis. In this letter, a semantic segmentation neural network which combines the strengths of residual learning and U-Net is proposed for road area extraction. The network is built with residual units and has similar architecture to that of U-Net. The benefits of this model is two-fold: first, residual units ease training of deep networks. Second, the rich skip connections within the network could facilitate information propagation, allowing us to design networks with fewer parameters however better performance. We test our network on a public road dataset and compare it with U-Net and other two state of the art deep learning based road extraction methods. The proposed approach outperforms all the comparing methods, which demonstrates its superiority over recently developed state of the arts.

Table of Contents

  • I Introduction
  • II Methodology
  • II-A Deep ResUnet
  • II-A1 U-Net
  • II-A2 Residual unit
  • II-A3 Deep ResUnet
  • II-B Loss function
  • II-C Result refinement
  • III Experiments
  • III-A Dataset
  • III-B Implementation details
  • III-C Evaluation metrics
  • III-D Comparisons
  • IV Conclusion
  • References

Knowls

  1. Knowl 1 — Deep Residual U-Net (ResUnet) Architecture

    model/method

    The Deep Residual U-Net (ResUnet) is an encoder-bridge-decoder neural network designed for semantic segmentation of roads from aerial imagery. It adapts the U-Net architecture by replacing standard convolutional units with full pre-activation residual units and removing cropping operations, enabling equal input and output spatial resolutions with significantly fewer parameters (7.8 million vs. 30.6 million in standard U-Net).

    The network comprises three parts:

    • Encoding Path: Contains three residual units. Instead of pooling operations, downsampling is performed by setting the convolution stride to 2 in the first convolution block of each residual unit after the input level.
    • Bridge: A single residual unit connecting the encoder and decoder paths, also initialized with a stride-2 convolution to downsample feature maps.
    • Decoding Path: Contains three residual units. Prior to each unit, feature maps from the lower level undergo upsampling and are concatenated with the skip-connected feature maps of the corresponding encoding level.

    All residual units consist of two 3×33 \times 3 convolutional blocks with identity mappings. Each block applies Batch Normalization (BN), ReLU activation, and a 3×33 \times 3 convolution. The final layer consists of a 1×11 \times 1 convolution followed by a sigmoid activation function to output pixel-wise road probability maps across 15 total convolutional layers.

  2. Knowl 2 — Layer Specifications of the ResUnet Model

    data/table

    The detailed layer configuration, kernel sizes, strides, and output tensor dimensions of the 7-level ResUnet for an input image of size 224×224×3224 \times 224 \times 3 are specified as follows:

    Path / Level Unit Level Conv Layer Filter (Size/Channels) Stride Output Size
    Input - - - - 224×224×3224 \times 224 \times 3
    Encoding Level 1 Conv 1 3×33 \times 3 / 64 1 224×224×64224 \times 224 \times 64
    Conv 2 3×33 \times 3 / 64 1 224×224×64224 \times 224 \times 64
    Level 2 Conv 3 3×33 \times 3 / 128 2 112×112×128112 \times 112 \times 128
    Conv 4 3×33 \times 3 / 128 1 112×112×128112 \times 112 \times 128
    Level 3 Conv 5 3×33 \times 3 / 256 2 56×56×25656 \times 56 \times 256
    Conv 6 3×33 \times 3 / 256 1 56×56×25656 \times 56 \times 256
    Bridge Level 4 Conv 7 3×33 \times 3 / 512 2 28×28×51228 \times 28 \times 512
    Conv 8 3×33 \times 3 / 512 1 28×28×51228 \times 28 \times 512
    Decoding Level 5 Conv 9 3×33 \times 3 / 256 1 56×56×25656 \times 56 \times 256
    Conv 10 3×33 \times 3 / 256 1 56×56×25656 \times 56 \times 256
    Level 6 Conv 11 3×33 \times 3 / 128 1 112×112×128112 \times 112 \times 128
    Conv 12 3×33 \times 3 / 128 1 112×112×128112 \times 112 \times 128
    Level 7 Conv 13 3×33 \times 3 / 64 1 224×224×64224 \times 224 \times 64
    Conv 14 3×33 \times 3 / 64 1 224×224×64224 \times 224 \times 64
    Output - Conv 15 1×11 \times 1 / 1 1 224×224×1224 \times 224 \times 1

    Feature maps decrease in spatial resolution via strided convolutions from 224×224224 \times 224 down to 28×2828 \times 28 at Level 4, while channel depth increases from 64 to 512. In the decoding path, upsampled feature maps are concatenated with the corresponding encoder level's features before passing through each level's residual block.

  3. Knowl 3 — Pre-Activation Residual Unit Formulation

    equation

    The ResUnet architecture constructs its encoder, bridge, and decoder stages using full pre-activation residual units. For the ll-th residual unit with input tensor xlx_l and learnable parameters Wl\mathcal{W}_l, the unit transformations are governed by:

    yl=h(xl)+F(xl,Wl)y_l = h(x_l) + \mathcal{F}(x_l, \mathcal{W}_l)

    xl+1=f(yl)x_{l+1} = f(y_l)

    where h(xl)=xlh(x_l) = x_l denotes the identity mapping shortcut, F()\mathcal{F}(\cdot) denotes the residual mapping comprising two sequential operations of Batch Normalization \rightarrow ReLU activation 3×3\rightarrow 3 \times 3 Convolution, and f(yl)=ylf(y_l) = y_l is the activation function (identity when using pre-activation). This formulation facilitates forward and backward signal propagation, avoiding gradient degradation in deep architectures.

  4. Knowl 4 — ResUnet Training Loss Function

    equation

    The network parameters W\mathcal{W} of the ResUnet are estimated by minimizing the Mean Squared Error (MSE) loss between the predicted segmentation maps and the corresponding ground-truth binary masks:

    L(W)=1Ni=1NNet(Ii;W)si2\mathcal{L}(\mathcal{W}) = \frac{1}{N} \sum_{i=1}^{N} \|\text{Net}(I_i; \mathcal{W}) - s_i\|^2

    where NN is the total number of training image patches, IiI_i is the ii-th input aerial image patch, sis_i is the binary ground-truth segmentation label where road pixels are 1 and non-road pixels are 0, and Net(Ii;W)\text{Net}(I_i; \mathcal{W}) denotes the sigmoid-activated pixel probability map predicted by the network.

  5. Knowl 5 — Overlapping Tiled Inference for Boundary Artifact Refinement

    model/method

    Due to zero-padding in the convolutional layers of the ResUnet, pixels near the spatial boundaries of output patches exhibit lower prediction accuracy than central pixels. To segment large aerial images (e.g., 1500×15001500 \times 1500 pixels) while mitigating boundary artifacts:

    1. Sub-images of size 224×224224 \times 224 pixels are cropped from the large input image using a sliding window with a spatial overlap of oo pixels (set to o=14o = 14).
    2. Each sub-image is independently processed by the network to produce a 224×224224 \times 224 segmentation prediction.
    3. The full-image segmentation is reconstructed by stitching the sub-segmentations together, taking the average of overlapping predicted pixel values across overlapping regions.
  6. Knowl 6 — Relaxed Precision, Relaxed Recall, and Break-Even Point Metrics

    definition

    Due to small labeling inaccuracies in remote sensing ground-truth annotations, road extraction performance is evaluated using relaxed precision and relaxed recall with a slack parameter ρ\rho (set to ρ=3\rho = 3 pixels):

    • Relaxed Precision (Correctness): The fraction of predicted road pixels that fall within a Euclidean distance of ρ\rho pixels from any ground-truth labeled road pixel.
    • Relaxed Recall (Completeness): The fraction of ground-truth labeled road pixels that fall within a Euclidean distance of ρ\rho pixels from any predicted road pixel.
    • Break-Even Point: The point on the relaxed precision-recall curve where precision equals recall, representing the intersection of the precision-recall curve with the identity line y=xy = x.
  7. Knowl 7 — Road Extraction Performance on the Massachusetts Roads Dataset

    data/table

    The ResUnet was evaluated against state-of-the-art road extraction architectures on the 49 test images of the Massachusetts roads dataset using the break-even point metric:

    Model Breakeven point
    Mnih-CNN 0.8873
    Mnih-CNN + CRF 0.8904
    Mnih-CNN + Post-Processing 0.9006
    Saito-CNN 0.9047
    U-Net 0.9053
    ResUnet (Proposed) 0.9187

    ResUnet achieves the highest break-even point score (0.9187), outperforming U-Net (0.9053) and CNN-based baselines while utilizing only 7.8 million parameters (approximately 1/4 of U-Net's 30.6 million parameters).

  8. Knowl 8 — Massachusetts Roads Dataset Training Setup

    experimental setup

    The Massachusetts roads dataset consists of 1171 aerial images covering approximately 500 km2500\text{ km}^2 (urban, suburban, and rural areas) at 1.2 m/pixel1.2\text{ m/pixel} spatial resolution, with an image size of 1500×15001500 \times 1500 pixels. The dataset is partitioned into 1108 training images, 14 validation images, and 49 test images.

    Training details:

    • Input Data: 30,000 image patches of size 224×224×3224 \times 224 \times 3 randomly sampled from the 1108 training images without any data augmentation.
    • Hardware & Framework: Implemented in Keras and trained on a single NVIDIA Titan 1080 GPU.
    • Optimization: Stochastic Gradient Descent (SGD) with a mini-batch size of 8.
    • Learning Rate Schedule: Initial learning rate of 0.0010.001, reduced by a factor of 0.10.1 every 20 epochs.
    • Convergence: The network converges within 50 epochs.
  9. Knowl 9 — Contextual Separation Capabilities and Parking Lot Limitation

    empirical result

    ResUnet leverages context information along with rich multi-scale skip connections to produce distinct segmentation behaviors:

    • Multi-lane Separation: The model successfully segments adjacent individual lanes with clear, sharp boundaries without fusing them, whereas comparison models confuse adjacent lanes.
    • Occlusion Robustness: The model recovers road segments partially occluded by tree canopies.
    • Class Disambiguation: The model successfully differentiates roads from visually similar structural objects such as airport runways.
    • Limitation on Parking Lots: The model fails to segment roads inside parking lots. Because most parking lot paths are unlabeled in the dataset ground truth, the network's context-learning mechanism learns to treat parking lot lanes as background despite their visual similarity to ordinary roads.

Coverage note — None was omitted; all key architectural components, experimental configurations, equations, evaluation metrics, empirical results, and failure modes are fully covered.

References

  1. 1.X. Huang and L. Zhang, “Road centreline extraction from highresolution imagery based on multiscale structural features and support vector machines,” IJRS, vol. 30, no. 8, pp. 1977–1987, 2009.
  2. 2.V. Mnih and G. Hinton, “Learning to detect roads in high-resolution aerial images,” ECCV, pp. 210–223, 2010.
  3. 3.C. Unsalan and B. Sirmacek, “Road network detection using probabilistic and graph theoretical methods,” TGRS, vol. 50, no. 11, pp. 4441–4453, 2012.
  4. 4.G. Cheng, Y. Wang, Y. Gong, F. Zhu, and C. Pan, “Urban road extraction via graph cuts based probability propagation,” in ICIP, 2015, pp. 5072–5076.
  5. 5.S. Saito, T. Yamashita, and Y. Aoki, “Multiple object extraction from aerial imagery with convolutional neural networks,” J. ELECTRON IMAGING, vol. 2016, no. 10, pp. 1–9, 2016.
  6. 6.R. Alshehhi and P. R. Marpu, “Hierarchical graph-based segmentation for extracting road networks from high-resolution satellite images,” P&RS, vol. 126, pp. 245–260, 2017.
  7. 7.B. Liu, H. Wu, Y. Wang, and W. Liu, “Main road extraction from ZY-3 grayscale imagery based on directional mathematical morphology and VGI prior knowledge in urban areas,” PLOS ONE, vol. 10, no. 9, p. e0138071, 2015.
  8. 8.C. Sujatha and D. Selvathi, “Connected component-based technique for automatic extraction of road centerline in high resolution satellite images,” J. Image Video Process., vol. 2015, no. 1, p. 8, 2015.
  9. 9.G. Cheng, Y. Wang, S. Xu, H. Wang, S. Xiang, and C. Pan, “Automatic road detection and centerline extraction via cascaded end-to-end convolutional neural network,” TGRS, vol. 55, no. 6, pp. 3322–3337, 2017.
  10. 10.G. Cheng, F. Zhu, S. Xiang, and C. Pan, “Road centerline extraction via semisupervised segmentation and multidirection nonmaximum suppression,” GRSL, vol. 13, no. 4, pp. 545–549, 2016.
  11. 11.M. Song and D. Civco, “Road extraction using SVM and image segmentation,” PE&RS, vol. 70, no. 12, pp. 1365–1371, 2004.
  12. 12.S. Das, T. T. Mirnalinee, and K. Varghese, “Use of salient features for the design of a multistage framework to extract roads from high-resolution multispectral satellite images,” TGRS, vol. 49, no. 10, pp. 3906–3931, 2011.
  13. 13.B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva, “Learning deep features for scene recognition using places database,” in NIPS, 2014, pp. 487–495.
  14. 14.S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” TPAMI, vol. 39, no. 6, p. 1137, 2017.
  15. 15.V. Mnih and G. E. Hinton, “Learning to label aerial images from noisy data,” in ICML, 2012, pp. 567–574.
  16. 16.Q. Zhang, Y. Wang, Q. Liu, X. Liu, and W. Wang, “CNN based suburban building detection using monocular high resolution google earth images,” in IGARSS, 2016, pp. 661–664.
  17. 17.L. Zhang, L. Zhang, and B. Du, “Deep learning for remote sensing data: A technical tutorial on the state of the art,” Geosci. Remote Sens. Mag., vol. 4, no. 2, pp. 22–40, 2016.
  18. 18.Z. Zhang, Y. Wang, Q. Liu, L. Li, and P. Wang, “A CNN based functional zone classification method for aerial images,” in IGARSS, 2016, pp. 5449–5452.
  19. 19.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in CVPR, 2015, pp. 1–9.
  20. 20.K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv:1409.1556, 2014.
  21. 21.K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR, 2016, pp. 770–778.
  22. 22.——, “Identity mappings in deep residual networks,” in ECCV, 2016, pp. 630–645.
  23. 23.J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in CVPR, 2015, pp. 3431–3440.
  24. 24.O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in MICCAI, 2015, pp. 234–241.
  25. 25.F. Chollet et al., “Keras,” https://github.com/fchollet/keras, 2015.
  26. 26.M. Ehrig and J. Euzenat, “Relaxed precision and recall for ontology matching,” in Workshop on Integrating ontology, 2005, pp. 25–32.

Citation

MLA
Zhang, Z., et al. “Road Extraction by Deep Residual U-Net”. IEEE Geoscience and Remote Sensing Letters, vol. 15, no. 5, 2018, pp. 749–53, https://doi.org/10.1109/LGRS.2018.2802944.
APA
Zhang, Z., Liu, Q., & Wang, Y. (2018). Road Extraction by Deep Residual U-Net. IEEE Geoscience and Remote Sensing Letters, 15(5), 749–753. https://doi.org/10.1109/LGRS.2018.2802944
Chicago
Zhang, Z., Q. Liu, and Y. Wang. 2018. “Road Extraction by Deep Residual U-Net”. IEEE Geoscience and Remote Sensing Letters 15 (5): 749–53. https://doi.org/10.1109/LGRS.2018.2802944.
Harvard
Zhang, Z., Liu, Q. and Wang, Y. (2018) “Road Extraction by Deep Residual U-Net”, IEEE Geoscience and Remote Sensing Letters, 15(5), pp. 749–753. Available at: https://doi.org/10.1109/LGRS.2018.2802944.
Vancouver
1. Zhang Z, Liu Q, Wang Y (2018) Road Extraction by Deep Residual U-Net. IEEE Geoscience and Remote Sensing Letters 15:749–753

BibTeX

@article{Zhang_2018, title={Road Extraction by Deep Residual U-Net}, volume={15}, ISSN={1558-0571}, url={http://dx.doi.org/10.1109/LGRS.2018.2802944}, DOI={10.1109/lgrs.2018.2802944}, number={5}, journal={IEEE Geoscience and Remote Sensing Letters}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Zhang, Zhengxin and Liu, Qingjie and Wang, Yunhong}, year={2018}, month=May, pages={749–753} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/