Fully Convolutional Siamese Networks for Change Detection
Rodrigo Caye DaudtBertrand Le SauxAlexandre Boulch
Proposes fully convolutional Siamese architectures for remote sensing change detection that train from scratch on coregistered image pairs to deliver superior accuracy at over 500 times the speed of previous methods.
Earth observation programs generate vast streams of satellite and aerial imagery that are vital for tracking urban expansion, deforestation, and environmental changes. However, conventional automated systems often rely on slow, patch-based image comparisons or require complex pre-training steps on unrelated datasets. This bottleneck limits the ability of organizations to efficiently process massive incoming imagery data in near-real-time.
The article demonstrates and evaluates three end-to-end deep learning models designed specifically for automated change detection between pairs of aligned images. It tests whether fully convolutional neural network architectures—which output dense, pixel-level predictions directly—can be trained from scratch on existing change detection datasets to simultaneously improve mapping accuracy and operational speed.
The researchers developed three distinct models based on encoder-decoder architectures with shortcut connections that preserve fine spatial details. The first merges both images at the input stage, while the other two use twin network branches (a Siamese setup) that process each image separately before combining them through concatenation or absolute feature differences. The team evaluated these models on two open datasets: the Onera Satellite Change Detection dataset, containing multispectral satellite images, and the Air Change dataset, containing standard aerial photography. The models were evaluated using precision, recall, and balanced overall performance scores against established industry benchmarks.
The evaluation produced several critical findings. First, the proposed architectures processed image pairs in under 0.1 seconds, achieving an inference speedup of at least 500 times compared to existing methods that took 50 seconds to several minutes. Second, the Siamese network using feature differencing and the unified input network achieved the strongest balanced accuracy across tests. On the satellite dataset, they boosted balanced performance scores from previous baselines of approximately 38–42% up to nearly 58%. Third, the models successfully learned directly from raw training data without requiring pre-training on outside datasets, functioning effectively on both standard color channels and 13-band multispectral data.
These results show that large-scale Earth observation monitoring can be automated at a fraction of current computational times without sacrificing accuracy. For operational programs such as Copernicus and Landsat, this speedup lowers computing infrastructure costs, eliminates the operational risk of processing backlogs, and allows near-instantaneous global land-use monitoring. The Siamese architecture that calculates feature differences is especially effective because its design directly mirrors the core task of isolating change.
Organizations handling high-volume geospatial analytics should consider adopting fully convolutional and Siamese difference architectures for core image-comparison workflows. However, decision-makers should note that the current models assign binary change labels rather than categorizing specific types of semantic change. Before broad deployment across diverse platforms, technical teams should conduct pilot tests to evaluate the architectures on broader image sequences, radar data, and larger geographic areas.
- Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). It introduces the foundational fully convolutional network (FCN) architecture with skip connections for dense pixel-level prediction upon which the source's change detection networks are built.
- Paper: Fully-Convolutional Siamese Networks for Object Tracking, Luca Bertinetto et al. (2016). It establishes the fully-convolutional Siamese network paradigm for pairwise feature comparison that the source adapts directly to image pairs for change detection.
- Paper: FlowNet: Learning Optical Flow with Convolutional Networks, Philipp Fischer et al. (2015). It presents foundational architectures (FlowNetSimple and FlowNetCorr) for comparing dual-image inputs via stacking and feature correlation in fully convolutional pipelines.
- Paper: Deep Metric Learning Using Triplet Network, Elad Hoffer et al. (2014). It provides foundational insights into deep metric learning and multi-branch feature representations for pairwise and triplet distance modeling.
- Paper: Object Detection in Optical Remote Sensing Images: A Survey and A New Benchmark, Ke Li et al. (2019). It explores large-scale benchmark evaluation of deep learning vision architectures specifically tailored to optical Earth observation and remote sensing imagery.
- Paper: U2Fusion: A Unified Unsupervised Image Fusion Network, Han Xu et al. (2020). It builds on multi-image convolutional processing by providing a unified network to fuse complementary information across multi-modal and multi-temporal image pairs.
