Fully Convolutional Siamese Networks for Change Detection

Rodrigo Caye DaudtBertrand Le SauxAlexandre Boulch

article2018International Conference on Information Photonics1,778 citations

Proposes fully convolutional Siamese architectures for remote sensing change detection that train from scratch on coregistered image pairs to deliver superior accuracy at over 500 times the speed of previous methods.

Listen

Earth observation programs generate vast streams of satellite and aerial imagery that are vital for tracking urban expansion, deforestation, and environmental changes. However, conventional automated systems often rely on slow, patch-based image comparisons or require complex pre-training steps on unrelated datasets. This bottleneck limits the ability of organizations to efficiently process massive incoming imagery data in near-real-time.

The article demonstrates and evaluates three end-to-end deep learning models designed specifically for automated change detection between pairs of aligned images. It tests whether fully convolutional neural network architectures—which output dense, pixel-level predictions directly—can be trained from scratch on existing change detection datasets to simultaneously improve mapping accuracy and operational speed.

The researchers developed three distinct models based on encoder-decoder architectures with shortcut connections that preserve fine spatial details. The first merges both images at the input stage, while the other two use twin network branches (a Siamese setup) that process each image separately before combining them through concatenation or absolute feature differences. The team evaluated these models on two open datasets: the Onera Satellite Change Detection dataset, containing multispectral satellite images, and the Air Change dataset, containing standard aerial photography. The models were evaluated using precision, recall, and balanced overall performance scores against established industry benchmarks.

The evaluation produced several critical findings. First, the proposed architectures processed image pairs in under 0.1 seconds, achieving an inference speedup of at least 500 times compared to existing methods that took 50 seconds to several minutes. Second, the Siamese network using feature differencing and the unified input network achieved the strongest balanced accuracy across tests. On the satellite dataset, they boosted balanced performance scores from previous baselines of approximately 38–42% up to nearly 58%. Third, the models successfully learned directly from raw training data without requiring pre-training on outside datasets, functioning effectively on both standard color channels and 13-band multispectral data.

These results show that large-scale Earth observation monitoring can be automated at a fraction of current computational times without sacrificing accuracy. For operational programs such as Copernicus and Landsat, this speedup lowers computing infrastructure costs, eliminates the operational risk of processing backlogs, and allows near-instantaneous global land-use monitoring. The Siamese architecture that calculates feature differences is especially effective because its design directly mirrors the core task of isolating change.

Organizations handling high-volume geospatial analytics should consider adopting fully convolutional and Siamese difference architectures for core image-comparison workflows. However, decision-makers should note that the current models assign binary change labels rather than categorizing specific types of semantic change. Before broad deployment across diverse platforms, technical teams should conduct pilot tests to evaluate the architectures on broader image sequences, radar data, and larger geographic areas.

Cover for Fully Convolutional Siamese Networks for Change Detection

Abstract

This paper presents three fully convolutional neural network architectures which perform change detection using a pair of coregistered images. Most notably, we propose two Siamese extensions of fully convolutional networks which use heuristics about the current problem to achieve the best results in our tests on two open change detection datasets, using both RGB and multispectral images. We show that our system is able to learn from scratch using annotated change detection images. Our architectures achieve better performance than previously proposed methods, while being at least 500 times faster than related systems. This work is a step towards efficient processing of data from large scale Earth observation systems such as Copernicus or Landsat.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Proposed approach
  • 4 Experiments
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Fully Convolutional Siamese-Difference Architecture

    model/method

    The Fully Convolutional Siamese - Difference (FC-Siam-diff) architecture performs pixel-level change detection on a pair of coregistered images, I1,I2∈RH×W×C\mathbf{I}_1, \mathbf{I}_2 \in \mathbb{R}^{H \times W \times C}, where HH and WW denote spatial dimensions and CC is the number of spectral bands (e.g., C=3C=3 for RGB or C=13C=13 for multispectral Sentinel-2).

    The network employs a two-stream Siamese encoder with shared weights followed by a single decoder with skip connections based on feature differences:

    1. Shared Encoder Streams: Each image I1\mathbf{I}_1 and I2\mathbf{I}_2 is processed independently through identical encoder branches with shared parameters. Each branch consists of four convolutional stages separated by 2×22 \times 2 max-pooling operations (stride 2):

      • Block 1: C→16→16C \to 16 \to 16 convolution channels, followed by 2×22\times 2 max pooling.
      • Block 2: 16→32→3216 \to 32 \to 32 convolution channels, followed by 2×22\times 2 max pooling.
      • Block 3: 32→64→64→6432 \to 64 \to 64 \to 64 convolution channels, followed by 2×22\times 2 max pooling.
      • Block 4: 64→128→128→12864 \to 128 \to 128 \to 128 convolution channels, followed by 2×22\times 2 max pooling.
    2. Difference Skip Connections: To explicitly guide the network toward change representations while preserving spatial boundary localization, skip connections at each encoder stage l∈{1,2,3,4}l \in \{1, 2, 3, 4\} compute the element-wise absolute difference of the feature activations from both streams: Dl=∣Fl(1)−Fl(2)∣\mathbf{D}_l = |\mathbf{F}_l^{(1)} - \mathbf{F}_l^{(2)}| where Fl(1)\mathbf{F}_l^{(1)} and Fl(2)\mathbf{F}_l^{(2)} are the feature representations produced by encoder block ll for I1\mathbf{I}_1 and I2\mathbf{I}_2, respectively. Because Dl\mathbf{D}_l retains the single-branch channel width (128,64,32,16128, 64, 32, 16), it maintains a compact channel profile.

    3. Decoder: The decoder upsamples features via 2×22\times 2 transpose convolutions and concatenates them with Dl\mathbf{D}_l at matching spatial resolutions:

      • Stage 4: Transpose convolution (128→128128 \to 128, up 2) concatenated with D4\mathbf{D}_4 (128 channels) →\to convolution block (256→128→128→64256 \to 128 \to 128 \to 64).
      • Stage 3: Transpose convolution (64→6464 \to 64, up 2) concatenated with D3\mathbf{D}_3 (64 channels) →\to convolution block (128→64→64→32128 \to 64 \to 64 \to 32).
      • Stage 2: Transpose convolution (32→3232 \to 32, up 2) concatenated with D2\mathbf{D}_2 (32 channels) →\to convolution block (64→32→1664 \to 32 \to 16).
      • Stage 1: Transpose convolution (16→1616 \to 16, up 2) concatenated with D1\mathbf{D}_1 (16 channels) →\to convolution block (32→16→232 \to 16 \to 2).

    The final output is a 2-channel dense logit map corresponding to change and no-change classes for every pixel in the input image pair.

  2. Knowl 2 — Fully Convolutional Early Fusion Architecture

    model/method

    The Fully Convolutional Early Fusion (FC-EF) network performs dense change detection on a coregistered multi-temporal image pair by concatenating the two images along their channel dimension at the network input.

    For two images I1,I2∈RH×W×C\mathbf{I}_1, \mathbf{I}_2 \in \mathbb{R}^{H \times W \times C}, the input tensor X=[I1,I2]∈RH×W×2C\mathbf{X} = [\mathbf{I}_1, \mathbf{I}_2] \in \mathbb{R}^{H \times W \times 2C} is passed through a single U-Net-style encoder-decoder network. To prevent overfitting on small change detection datasets, the network is shallower and narrower than the standard U-Net, comprising four downsampling and four upsampling stages:

    1. Encoder:

      • Block 1: 2C→16→162C \to 16 \to 16 convolution channels, followed by 2×22\times 2 max pooling.
      • Block 2: 16→32→3216 \to 32 \to 32 convolution channels, followed by 2×22\times 2 max pooling.
      • Block 3: 32→64→64→6432 \to 64 \to 64 \to 64 convolution channels, followed by 2×22\times 2 max pooling.
      • Block 4: 64→128→128→12864 \to 128 \to 128 \to 128 convolution channels, followed by 2×22\times 2 max pooling.
    2. Skip Connections & Decoder: Features from each encoder block are concatenated directly across matching spatial resolutions to the upsampled decoder features:

      • Stage 4: Transpose convolution (128→128128 \to 128, up 2) concatenated with Block 4 skip (128 channels) →\to convolution block (256→128→128→64256 \to 128 \to 128 \to 64).
      • Stage 3: Transpose convolution (64→6464 \to 64, up 2) concatenated with Block 3 skip (64 channels) →\to convolution block (128→64→64→32128 \to 64 \to 64 \to 32).
      • Stage 2: Transpose convolution (32→3232 \to 32, up 2) concatenated with Block 2 skip (32 channels) →\to convolution block (64→32→1664 \to 32 \to 16).
      • Stage 1: Transpose convolution (16→1616 \to 16, up 2) concatenated with Block 1 skip (16 channels) →\to convolution block (32→16→232 \to 16 \to 2).

    The output is a 2-channel dense prediction map of change vs. no-change probabilities across arbitrary input spatial dimensions.

  3. Knowl 3 — Fully Convolutional Siamese-Concatenation Architecture

    model/method

    The Fully Convolutional Siamese - Concatenation (FC-Siam-conc) architecture applies a Siamese encoder paired with direct feature concatenation across skip connections for dense change detection on an image pair (I1,I2)(\mathbf{I}_1, \mathbf{I}_2).

    1. Shared Siamese Encoder: The two images are processed separately by identical weight-shared branches. For an input of CC channels:

      • Block 1: C→16→16C \to 16 \to 16 channels, followed by 2×22\times 2 max pooling.
      • Block 2: 16→32→3216 \to 32 \to 32 channels, followed by 2×22\times 2 max pooling.
      • Block 3: 32→64→64→6432 \to 64 \to 64 \to 64 channels, followed by 2×22\times 2 max pooling.
      • Block 4: 64→128→128→12864 \to 128 \to 128 \to 128 channels, followed by 2×22\times 2 max pooling.
    2. Concatenation Skip Connections & Decoder: During decoding, skip connections from both encoder streams (denoted Fl(1)\mathbf{F}_l^{(1)} and Fl(2)\mathbf{F}_l^{(2)}) are simultaneously concatenated with the upsampled decoder feature map, doubling the skip channel volume compared to single-stream models:

      • Stage 4: Transpose convolution (128→128128 \to 128, up 2) concatenated with F4(1)\mathbf{F}_4^{(1)} (128 channels) and F4(2)\mathbf{F}_4^{(2)} (128 channels), yielding 384384 channels →\to convolution block (384→128→128→64384 \to 128 \to 128 \to 64).
      • Stage 3: Transpose convolution (64→6464 \to 64, up 2) concatenated with F3(1)\mathbf{F}_3^{(1)} (64 channels) and F3(2)\mathbf{F}_3^{(2)} (64 channels), yielding 192192 channels →\to convolution block (192→64→64→32192 \to 64 \to 64 \to 32).
      • Stage 2: Transpose convolution (32→3232 \to 32, up 2) concatenated with F2(1)\mathbf{F}_2^{(1)} (32 channels) and F2(2)\mathbf{F}_2^{(2)} (32 channels), yielding 9696 channels →\to convolution block (96→32→1696 \to 32 \to 16).
      • Stage 1: Transpose convolution (16→1616 \to 16, up 2) concatenated with F1(1)\mathbf{F}_1^{(1)} (16 channels) and F1(2)\mathbf{F}_1^{(2)} (16 channels), yielding 4848 channels →\to convolution block (48→16→248 \to 16 \to 2).

    The output produces binary change classification logits per pixel.

  4. Knowl 4 — Quantitative Evaluation on OSCD and Air Change Datasets

    data/table

    The proposed fully convolutional architectures (FC-EF, FC-Siam-conc, and FC-Siam-diff) were benchmarked against previous patch-based convolutional neural networks (Early Fusion and Siamese) and existing change detection methods (DSCN, CXM, and SCCN) across the Onera Satellite Change Detection (OSCD) dataset (evaluated on both 3-channel RGB and 13-channel multispectral Sentinel-2 imagery) and the Air Change (AC) dataset (Szada/1 and Tiszadob/3 test subsets). Performance is reported as percentages for Precision, Recall, and F1 score on the "change" class, along with Global Pixel Accuracy.

    Data Network Prec. (%) Recall (%) Global (%) F1 (%)
    OSCD-3 ch. Siam. 21.57 79.40 76.76 33.85
    EF 21.56 82.14 83.63 34.15
    FC-EF 44.72 53.92 94.23 48.89
    FC-Siam-conc 42.89 47.77 94.07 45.20
    FC-Siam-diff 49.81 47.94 94.86 48.86
    OSCD-13 ch. Siam. 24.16 85.63 85.37 37.69
    EF 28.35 84.69 88.15 42.48
    FC-EF 64.42 50.97 96.05 56.91
    FC-Siam-conc 42.39 65.15 93.68 51.36
    FC-Siam-diff 57.84 57.99 95.68 57.92
    Szada/1 DSCN 41.2 57.4 NA 47.9
    CXM 36.5 58.4 NA 44.9
    SCCN 24.4 34.7 NA 28.7
    FC-EF 43.57 62.65 93.08 51.40
    FC-Siam-conc 40.93 65.61 92.46 50.41
    FC-Siam-diff 41.38 72.38 92.40 52.66
    Tiszadob/3 DSCN 88.3 85.1 NA 86.7
    CXM 61.7 93.4 NA 74.3
    SCCN 92.7 79.8 NA 85.8
    FC-EF 90.28 96.74 97.66 93.40
    FC-Siam-conc 72.07 96.87 93.04 82.65
    FC-Siam-diff 69.51 88.29 91.37 77.78

    The results show that:

    • On OSCD (both 3-channel and 13-channel), the fully convolutional networks outperform previous patch-based CNNs by over 14 percentage points in F1 score (reaching 57.92%57.92\% with FC-Siam-diff and 56.91%56.91\% with FC-EF on 13-channel imagery). Patch-based methods suffer from very low precision (21--28%).
    • Utilizing all 13 multispectral bands significantly improves precision and F1 over 3-channel RGB across all models (e.g., FC-Siam-diff F1 increases from 48.86%48.86\% to 57.92%57.92\%).
    • On the Air Change dataset, FC-Siam-diff achieves the top F1 score on Szada/1 (52.66%52.66\%) and FC-EF achieves the top F1 score on Tiszadob/3 (93.40%93.40\%), outperforming previous methods without requiring pre-training or transfer learning.
  5. Knowl 5 — Supervised End-to-End Training Setup for Change Detection

    experimental setup

    The change detection models (FC-EF, FC-Siam-conc, and FC-Siam-diff) are trained supervised and end-to-end from scratch without pre-training or transfer learning from external datasets.

    1. Class Balancing: To handle extreme class imbalance between the positive ("change") and negative ("no-change") classes, the cross-entropy loss assigns weights to the two classes that are inversely proportional to their respective pixel frequencies in the training set.

    2. Data Augmentation: The training patches are augmented using all 8 dihedral transformations (all possible horizontal/vertical flips and rotations by multiples of 90∘90^\circ).

    3. Regularization and Optimization: Dropout is incorporated during training to prevent overfitting. Experiments were implemented in PyTorch on an Nvidia GTX 1070 GPU.

    4. Data Splits:

      • Onera Satellite Change Detection (OSCD): 14 image pairs used for training and 10 image pairs used for testing.
      • Air Change (AC) Dataset: Szada and Tiszadob locations are trained and evaluated independently. For Szada, the top-left 748×448748 \times 448 rectangle of image Szada-1 serves as the test set and the remaining 6 images are used for training. For Tiszadob, the top-left 748×448748 \times 448 rectangle of image Tiszadob-3 serves as the test set and the remaining 4 images are used for training. The "Archieve" set is omitted as it contains only a single image pair.
  6. Knowl 6 — Inference Latency and Computational Efficiency of Fully Convolutional Change Detection

    empirical result

    The fully convolutional architectures (FC-EF, FC-Siam-conc, and FC-Siam-diff) achieve an inference latency of under 0.1 s0.1\text{ s} per image pair when tested on an Nvidia GTX 1070 GPU.

    In comparison:

    • Previous patch-based voting CNN approaches required several minutes per image pair to generate a full change map.
    • SCCN (Symmetric Convolutional Coupling Network) requires approximately 50 s50\text{ s} per image on a comparable hardware setup.

    The proposed fully convolutional models provide a speedup of over 500×500\times relative to SCCN and patch-based methods while simultaneously improving detection accuracy, enabling real-time processing of large-scale Earth observation streams such as Copernicus and Landsat.

Coverage note — Deliberately omitted qualitative visual comparisons (Figures 2 and 3) as their empirical findings are fully captured by the quantitative metrics and architectural descriptions.

References

  1. 1.Ashbindu Singh, ‘‘Review article digital change detection techniques using remotely-sensed data,’’ International journal of remote sensing, vol. 10, no. 6, pp. 989– 1003, 1989.
  2. 2.Masroor Hussain, Dongmei Chen, Angela Cheng, Hui Wei, and David Stanley, ‘‘Change detection from remotely sensed images: From pixel-based to object-based approaches,’’ ISPRS Journal of Photogrammetry and Remote Sensing, vol. 80, pp. 91–106, 2013.
  3. 3.Rodrigo Caye Daudt, Bertrand Le Saux, Alexandre Boulch, and Yann Gousseau, ‘‘Urban change detection for multispectral earth observation using convolutional neural networks,’’ in International Geoscience and Remote Sensing Symposium (IGARSS). IEEE, 2018 (To Appear).
  4. 4.Csaba Benedek and Tam'as Szir'anyi, ‘‘Change detection in optical aerial images by a multilayer conditional mixed markov model,’’ IEEE Transactions on Geoscience and Remote Sensing, vol. 47, no. 10, pp. 3416– 3430, 2009.
  5. 5.Olaf Ronneberger, Philipp Fischer, and Thomas Brox, ‘‘U-net: Convolutional networks for biomedical image segmentation,’’ in International Conference on Medical image computing and computer-assisted intervention. Springer, 2015, pp. 234–241.
  6. 6.Bertrand Le Saux and Hicham Randrianarivo, ‘‘Urban change detection in sar images by interactive learning,’’ in Geoscience and Remote Sensing Symposium (IGARSS), 2013 IEEE International. IEEE, 2013, pp. 3990–3993.
  7. 7.Simon Stent, Riccardo Gherardi, Bj"orn Stenger, and Roberto Cipolla, ‘‘Detecting change for multi-view, long-term surface inspection.,’’ in BMVC, 2015, pp. 127–1.
  8. 8.Jia Liu, Maoguo Gong, Kai Qin, and Puzhao Zhang, ‘‘A deep convolutional coupling network for change detection based on heterogeneous optical and radar images,’’ IEEE transactions on neural networks and learning systems, 2016.
  9. 9.Maoguo Gong, Jiaojiao Zhao, Jia Liu, Qiguang Miao, and Licheng Jiao, ‘‘Change detection in synthetic aperture radar images based on deep neural networks,’’ IEEE transactions on neural networks and learning systems, vol. 27, no. 1, pp. 125–138, 2016.
  10. 10.Arabi Mohammed El Amin, Qingjie Liu, and Yunhong Wang, ‘‘Zoom out cnns features for optical remote sensing change detection,’’ in Image, Vision and Computing (ICIVC), 2017 2nd International Conference on. IEEE, 2017, pp. 812–817.
  11. 11.Yang Zhan, Kun Fu, Menglong Yan, Xian Sun, Hongqi Wang, and Xiaosong Qiu, ‘‘Change detection based on deep siamese convolutional network for optical aerial images,’’ IEEE Geoscience and Remote Sensing Letters, vol. 14, no. 10, pp. 1845–1849, 2017.
  12. 12.L. Mou, X. Zhu, M. Vakalopoulou, K. Karantzalos, N. Paragios, B. Le Saux, G. Moser, and D. Tuia, ‘‘Multitemporal Very High Resolution From Space: Outcome of the 2016 IEEE GRSS Data Fusion Contest,’’ IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 10, no. 8, pp. 3435– 3447, June 2017.
  13. 13.Sumit Chopra, Raia Hadsell, and Yann LeCun, ‘‘Learning a similarity metric discriminatively, with application to face verification,’’ in Computer Vision and Pattern Recognition, 2005. CVPR 2005. IEEE Computer Society Conference on. IEEE, 2005, vol. 1, pp. 539–546.
  14. 14.Sergey Zagoruyko and Nikos Komodakis, ‘‘Learning to compare image patches via convolutional neural networks,’’ in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 4353–4361.
  15. 15.Jonathan Long, Evan Shelhamer, and Trevor Darrell, ‘‘Fully convolutional networks for semantic segmentation,’’ in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 3431– 3440.
  16. 16.Luca Bertinetto, Jack Valmadre, Joao F Henriques, Andrea Vedaldi, and Philip HS Torr, ‘‘Fully-convolutional siamese networks for object tracking,’’ in European conference on computer vision. Springer, 2016, pp. 850– 865.
  17. 17.Nicolas Audebert, Bertrand Le Saux, and S'ebastien Lef`evre, ‘‘Beyond rgb: Very high resolution urban remote sensing with multimodal deep networks,’’ ISPRS Journal of Photogrammetry and Remote Sensing, 2017.
  18. 18.Nicolas Audebert, Bertrand Le Saux, and S'ebastien Lef`evre, ‘‘Segment-before-detect: Vehicle detection and classification through semantic segmentation of aerial images,’’ Remote Sensing, vol. 9, no. 4, pp. 368, 2017.
  19. 19.Jane Bromley, Isabelle Guyon, Yann LeCun, Eduard S"ackinger, and Roopak Shah, ‘‘Signature verification using a’’ siamese’’ time delay neural network,’’ in Advances in Neural Information Processing Systems, 1994, pp. 737–744.

Citation

MLA
Daudt, R. C., et al. “Fully Convolutional Siamese Networks for Change Detection”. arXiv, 2018, http://arxiv.org/abs/1810.08462v1.
APA
Daudt, R. C., Saux, B. L., & Boulch, A. (2018). Fully Convolutional Siamese Networks for Change Detection. arXiv. http://arxiv.org/abs/1810.08462v1
Chicago
Daudt, R. C., B. L. Saux, and A. Boulch. 2018. “Fully Convolutional Siamese Networks for Change Detection”. arXiv. http://arxiv.org/abs/1810.08462v1.
Harvard
Daudt, R.C., Saux, B.L. and Boulch, A. (2018) “Fully Convolutional Siamese Networks for Change Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1810.08462v1.
Vancouver
1. Daudt RC, Saux BL, Boulch A (2018) Fully Convolutional Siamese Networks for Change Detection. arXiv

BibTeX

@article{daudt2018fully,
  title = {Fully Convolutional Siamese Networks for Change Detection},
  author = {Daudt, Rodrigo Caye and Saux, Bertrand Le and Boulch, Alexandre},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1810.08462v1},
  eprint = {1810.08462}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF