GEO-Bench: Toward Foundation Models for Earth Monitoring

Alexandre LacosteNils LehmannPau RodríguezEvan D. SherwinHannah KernerBjörn LütjensJeremy IrvinDavid DaoHamed AlemohammadAlexandre Drouin

article2023NeurIPS175 citations

Introduces GEO-Bench, a standardized benchmark of twelve classification and segmentation tasks alongside twenty baseline evaluations to accelerate the development and reliable assessment of foundation models for satellite Earth monitoring.

Listen

Earth monitoring via satellite imagery is critical for climate change mitigation, disaster response, and environmental protection. However, current artificial intelligence models built for standard photography perform poorly on remote sensing tasks due to overhead viewpoints, multispectral sensor bands, and temporal variations. In addition, existing research evaluates Earth observation models inconsistently across narrow, non-standardized datasets, preventing objective comparisons and slowing practical progress.

The article introduces GEO-Bench, an open-source evaluation benchmark designed to standardize and accelerate the development of machine learning foundation models for Earth monitoring across diverse real-world tasks.

To construct the benchmark, domain experts assembled twelve distinct geospatial datasets spanning global geographies—six image classification and six semantic segmentation tasks covering multispectral, radar, and aerial imagery. The datasets were standardized into accessible sizes, balanced across classes, and formatted to run on single graphics processing units to reduce compute costs and carbon emissions. The authors established an evaluation protocol that benchmarks twenty baseline models across varying training data sizes using ten random seeds, robust statistical aggregation via interquartile means, and bootstrapped confidence intervals.

The findings show that standard computer vision pre-training significantly outperforms models trained from scratch, with modern architectures such as ConvNeXt and Swin Transformer systematically achieving top performance. Remarkably, pre-trained ConvNeXt models matched the accuracy of baseline models trained from scratch while requiring only 2% of the training data, demonstrating a fifty-fold improvement in data efficiency. Conversely, existing remote sensing models pre-trained specifically on satellite data did not demonstrate significant gains over models pre-trained on standard visual datasets, showing only modest gains when incorporating multispectral bands in convolutional architectures and performance declines in vision transformers.

These results highlight substantial opportunities to reduce labeling costs and compute timelines by adopting modern architectures and pre-trained weights for environmental monitoring. The lack of distinct outperformance from current satellite-specific pre-training suggests that existing geospatial foundation models require substantial architectural improvements to properly exploit multimodal remote sensing data.

Organizations developing Earth observation systems should adopt standardized evaluation workflows and prioritize modern convolutional and transformer backbones over older architectures. Future research should develop novel multimodal pre-training frameworks capable of effectively fusing temporal revisits, weather data, and multispectral information. While the benchmark provides strong coverage, its current limitations include data gaps in regions such as South America and parts of Asia, as well as a focus on single-image inputs rather than time-series sequences.

  • Paper: DINOv3, Oriane Siméoni et al. (2025). Develops a next-generation visual foundation model family that directly tackles the dense prediction and remote sensing performance challenges highlighted in GEO-Bench.
  • Paper: Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data, Lihe Yang et al. (2024). Demonstrates how large-scale unlabeled visual data pre-training can be scaled and transferred effectively to downstream dense prediction tasks.
Cover for GEO-Bench: Toward Foundation Models for Earth Monitoring

Abstract

Recent progress in self-supervision has shown that pre-training large neural networks on vast amounts of unsupervised data can lead to substantial increases in generalization to downstream tasks. Such models, recently coined foundation models, have been transformational to the field of natural language processing. Variants have also been proposed for image data, but their applicability to remote sensing tasks is limited. To stimulate the development of foundation models for Earth monitoring, we propose a benchmark comprised of six classification and six segmentation tasks, which were carefully curated and adapted to be both relevant to the field and well-suited for model evaluation. We accompany this benchmark with a robust methodology for evaluating models and reporting aggregated results to enable a reliable assessment of progress. Finally, we report results for 20 baselines to gain information about the performance of existing models. We believe that this benchmark will be a driver of progress across a variety of Earth monitoring tasks.

Table of Contents

  • 1 Introduction
  • 2 Remote sensing data for self-supervision
  • 3 GEO-Bench
  • 3.1 Design Principles
  • 3.2 Dataset Transformations
  • 4 Using The Benchmark
  • 4.1 Reporting Results
  • 5 Related Works
  • 6 Experiments
  • 6.1 Protocol
  • 6.2 Classification
  • 6.2.1 Baselines Naming Schema
  • 6.2.2 Comparing Baselines on RGB only
  • 6.2.3 Accuracy vs training set size
  • 6.2.4 Leveraging Multispectral Information
  • 6.3 Segmentation
  • 6.4 Resource Usage
  • 7 Conclusion
  • References
  • A Extended Results
  • A.1 Benchmark Coverage
  • A.2 Classification
  • A.3 Segmentation
  • A.3.1 Baselines
  • A.3.2 Comparing Baselines on RGB only
  • A.4 Convergence Time
  • A.5 Discriminativity of Datasets
  • A.6 Resources Usage
  • B Remote Sensing Data Schema
  • C Societal Impact of Foundation Models for Earth Monitoring
  • C.1 Climate mitigation and adaptation
  • C.2 Increased accessibility
  • C.3 Emissions of large pre-trained models
  • C.4 Fairness and biases

Knowls

  1. Knowl 1 — GEO-Bench Benchmark Composition and Task Specifications

    data/table

    GEO-Bench provides a standardized suite of twelve Earth observation downstream datasets, divided equally into six image classification tasks and six semantic segmentation tasks. Each dataset was adapted from open-source sources (indicated by the prefix m-) with permissive licenses, spatial resolutions spanning from 0.1 m/pixel0.1\text{ m/pixel} (aerial) to 30 m/pixel30\text{ m/pixel} (satellite), and spectral modalities including RGB, RGBN, multi-spectral (Sentinel-2, Landsat-8), synthetic aperture radar (SAR Sentinel-1), hyperspectral, and terrain elevation.

    Task Name Image Size # Classes Train Val Test # Bands Res (m) Sensors / Modalities
    Classification Datasets
    m-bigearthnet 120×120120 \times 120 43 20,000 1,000 1,000 12 10.0 Sentinel-2
    m-so2sat 32×3232 \times 32 17 19,992 986 986 18 10.0 Sentinel-2 + Sentinel-1
    m-brick-kiln 64×6464 \times 64 2 15,063 999 999 13 10.0 Sentinel-2
    m-forestnet 332×332332 \times 332 12 6,464 989 993 6 15.0 Landsat-8
    m-eurosat 64×6464 \times 64 10 2,000 1,000 1,000 13 10.0 Sentinel-2
    m-pv4ger 320×320320 \times 320 2 11,814 999 999 3 0.1 Aerial RGB
    Segmentation Datasets
    m-pv4ger-seg 320×320320 \times 320 2 3,000 403 403 3 0.1 Aerial RGB
    m-chesapeake-landcover 256×256256 \times 256 7 3,000 1,000 1,000 4 1.0 Aerial RGBN
    m-cashew-plantation 256×256256 \times 256 7 1,350 400 50 13 10.0 Sentinel-2
    m-SA-crop-type 256×256256 \times 256 10 3,000 1,000 1,000 13 10.0 Sentinel-2
    m-nz-cattle 500×500500 \times 500 2 524 66 65 3 0.1 Aerial RGB
    m-NeonTree 400×400400 \times 400 2 270 94 93 5 0.1 RGB + Hyperspectral + Elevation

    Object detection and object counting tasks were reformulated as semantic segmentation (e.g., m-nz-cattle and m-NeonTree) to maintain structural uniformity across dense prediction evaluations. All splits enforce zero spatial overlap between training, validation, and test partitions.

  2. Knowl 2 — GEO-Bench Model Evaluation and Aggregation Protocol

    experimental setup

    To ensure statistical reliability and comparability when evaluating models on GEO-Bench, the following reporting protocol is specified:

    1. Seed Repetitions: For any chosen hyperparameter configuration, models must be trained across at least 1010 independent random seeds.
    2. Hyperparameter Tuning Budget: Hyperparameter searches (e.g., learning rates for the backbone and new classification/segmentation heads) are restricted to a maximum of 1616 trials per task, with early stopping governed by validation accuracy or Intersection over Union (IoU).
    3. Data Augmentations: Allowed geometric transformations during fine-tuning are strictly constrained to 90∘90^\circ rotations and horizontal/vertical flips. Random crops and resizing are excluded to preserve native sensor ground sampling distance (spatial resolution).
    4. Interquartile Mean (IQM): For task-level point estimates across random seeds, IQM trims the bottom 25%25\% and top 25%25\% of observed test scores and computes the arithmetic mean of the remaining central 50%50\%, reducing variance relative to the mean and bias relative to the median.
    5. Task Score Normalization: Scores SlS_l on task ll are linearly scaled to Snorml∈[0,1]S_{\text{norm}}^l \in [0, 1] relative to fixed reference performance bounds (the minimum Smin⁡lS_{\min}^l and maximum Smax⁡lS_{\max}^l achieved across strong baselines): Snorml=Sl−Smin⁡lSmax⁡l−Smin⁡lS_{\text{norm}}^l = \frac{S_l - S_{\min}^l}{S_{\max}^l - S_{\min}^l}
    6. Stratified Bootstrap Aggregation: Aggregated benchmark scores are computed by taking the IQM across all normalized dataset outcomes. Confidence intervals (80%80\% or 95%95\%) are generated via stratified bootstrapping with N=1,000N = 1,000 resamples with replacement per dataset.
  3. Knowl 3 — Dataset Discriminativity Metric Formulation

    equation

    To quantify the statistical capability of a benchmark dataset ll to differentiate performance between two models ii and jj, pairwise discriminativity dijld_{ij}^l is defined using the binary entropy function H2H_2:

    dijl:=1−H2[p(Ail>Ajl)]d_{ij}^l := 1 - H_2\left[p\left(A_i^l > A_j^l\right)\right]

    where AilA_i^l and AjlA_j^l are random variables representing the validation performance metrics (e.g., accuracy or IoU) of algorithms ii and jj on task ll across random seeds, and p(Ail>Ajl)p(A_i^l > A_j^l) is the empirical probability that model ii outperforms model jj. Binary entropy in bits is:

    H2(p)=−plog⁡2(p)−(1−p)log⁡2(1−p)H_2(p) = -p \log_2(p) - (1-p) \log_2(1-p)

    If one algorithm always outperforms the other (p=0p = 0 or p=1p = 1), entropy is 00 and dijl=1d_{ij}^l = 1. If both algorithms perform identically (p=0.5p = 0.5), entropy is 11 and dijl=0d_{ij}^l = 0.

    The overall discriminativity DlD_l of dataset ll across a collection of mm candidate models is the average over all pairs:

    Dl:=1m2∑i=1m∑j=1mdijlD_l := \frac{1}{m^2} \sum_{i=1}^m \sum_{j=1}^m d_{ij}^l

  4. Knowl 4 — Empirical Performance of Vision Architectures on Optical Remote Sensing Classification

    empirical result

    Evaluating pre-trained convolutional and vision transformer architectures on the RGB channels of the GEO-Bench classification suite demonstrates the following performance dynamics:

    1. Modern Convolutional and Hierarchical Transformer Dominance: ConvNeXt-B and SwinV2-T (pre-trained on ImageNet via timm) substantially outperform standard ResNet architectures (ResNet-18, ResNet-50) and plain Vision Transformers (ViT-T, ViT-S) on aggregate normalized accuracy across all classification datasets.
    2. Supervision vs. Random Initialization: Pre-training provides massive data efficiency gains; a ResNet-18 initialized from scratch (random weights) severely underperforms ImageNet-pretrained ResNet-18 across all tasks. Furthermore, ConvNeXt-B trained on only 2%2\% of the downstream training data achieves aggregate performance equal to or greater than a ResNet-18 trained from scratch on 100%100\% of the data (a 50×50\times data efficiency factor).
    3. Domain-Specific Pre-training Parity on RGB: Self-supervised backbones pre-trained specifically on remote sensing imagery (such as ResNet-50 trained with SeCo or MoCo-S2 on Sentinel-2) do not exhibit significant gains over standard ImageNet-pretrained weights (timm) when evaluated exclusively on RGB channels.
  5. Knowl 5 — Fractional Training Set Scaling and Computational Convergence Rate

    empirical result

    GEO-Bench defines standardized fractional partitions of downstream training sets at training ratios r∈{0.01,0.02,0.05,0.10,0.20,0.50,1.00}r \in \{0.01, 0.02, 0.05, 0.10, 0.20, 0.50, 1.00\}.

    Empirical evaluation of the convergence time τi,j,kr\tau_{i,j,k}^r (defined as the number of gradient update steps required to reach peak validation performance for model ii on dataset jj during trial kk) demonstrates that convergence time scales approximately linearly with the training set size ratio rr.

    Because computation scales proportionally with rr, evaluating a model across all seven training size ratios requires a cumulative computational cost factor of:

    ∑rr=1.0+0.5+0.2+0.1+0.05+0.02+0.01=1.88\sum_{r} r = 1.0 + 0.5 + 0.2 + 0.1 + 0.05 + 0.02 + 0.01 = 1.88

    Evaluating the full spectrum of sample-efficiency curves requires only 1.88×1.88\times the compute of training on the 100%100\% dataset alone, rather than a naive 7×7\times increase.

  6. Knowl 6 — Impact of Multi-Spectral Channels and Pre-training Schemes

    empirical result

    Experiments analyzing the integration of multi-spectral sensor bands on Sentinel-2 tasks in GEO-Bench (m-bigearthnet, m-brick-kiln, m-eurosat, and m-so2sat) reveal distinct behaviors across architectures:

    1. Random Channel Initialization Penalty (+R-Multi): Adapting an RGB-pretrained backbone (ImageNet timm) by randomly initializing weights for additional non-RGB input channels in the first convolutional layer does not produce consistent accuracy gains and substantially increases fine-tuning time due to the delayed convergence of the newly initialized input channels.
    2. Domain-Specific Multi-Spectral SSL: Pre-training ResNet-50 directly on Sentinel-2 multi-spectral data using self-supervised objectives (MoCo-S2 or DINO-S2) yields modest performance gains on multi-spectral downstream tasks over RGB-only baselines.
    3. Transformer Degradation with Multi-Spectral Inputs: For Vision Transformers (ViT-S), incorporating multi-spectral bands consistently leads to a decrease in downstream fine-tuning classification performance compared to using RGB channels alone.
  7. Knowl 7 — Semantic Segmentation Baseline Performance in GEO-Bench

    empirical result

    Benchmarking six segmentation configurations formed by pairing convolutional backbones (ResNet-18, ResNet-50, ResNet-101 with ImageNet weights) with decoder architectures (U-Net and DeepLabV3) across the six GEO-Bench segmentation datasets indicates:

    1. Decoder Architecture Impact: Models utilizing U-Net decoders systematically outperform equivalent backbones equipped with DeepLabV3 decoders across diverse spatial resolutions (0.1 m0.1\text{ m} to 10.0 m10.0\text{ m}) and modalities.
    2. Backbone Depth and Stability: ResNet-50 with U-Net achieves the most consistent performance. While ResNet-101 with DeepLabV3 achieves competitive mean scores on certain datasets, it exhibits higher performance variance across random seeds and underperforms ResNet-50 on multiple downstream tasks.
  8. Knowl 8 — Object-Oriented Remote Sensing Data Schema and Task Specifications

    model/method

    GEO-Bench defines an object-oriented Python data schema to harmonize heterogeneous sensor formats and enable automated architecture configuration without loading large image datasets into memory:

    • Band: The primitive data structure representing a multi-dimensional spatial array paired with a BandInfo descriptor.
    • BandInfo: Metadata class specifying spectral band names, spatial resolution (meters/pixel), wavelength ranges, and sensor provenance (subclasses include SpectralBand for Sentinel-2/Landsat, Radar for Sentinel-1, Elevation, Hyperspectral, and MultiBand).
    • Sample: A unified collection associating input Band objects with a corresponding target label (which can also be a dense Band for segmentation tasks).
    • TaskSpecifications: An introspection object containing the dataset name, target label specification, band shapes, and BandInfo metadata. This enables models (such as Transformers with custom band encodings) to procedurally generate their input/output layers prior to data loading.
    • Band Statistics: Precomputed per-band distribution statistics (minimum, maximum, mean, variance, and percentiles) provided to facilitate input normalization.
  9. Knowl 9 — Dataset Curation and Balancing Principles for Downstream Benchmarking

    model/method

    GEO-Bench applies explicit filtering and transformation rules to standard Earth observation datasets to optimize them for model benchmarking:

    1. Subsampling Large Datasets: Datasets exceeding 20,00020,000 training samples are randomly subsampled. This aligns the benchmark with typical downstream remote sensing scenarios where ground-truth labels are scarce, enhances the benchmark's discriminativity between models of similar capacity, and allows complete multi-seed replication on single 32 GB32\text{ GB} GPUs.
    2. Mitigating Class Imbalance: Majority classes are randomly downsampled to achieve near-uniform class frequency distributions. This prevents models from inflating benchmark scores via class-imbalance heuristics rather than superior visual representations.
    3. Spatial Disjointness: When original benchmark splits are absent, validation and test splits are extracted such that no spatial tile overlap exists between training and evaluation partitions, avoiding spatial autocorrelation bias.
  10. Knowl 10 — Limitations of GEO-Bench for Foundation Model Assessment

    limitation

    The GEO-Bench suite exhibits specific structural limitations:

    1. Temporal and Non-Spatial Modalities: The benchmark evaluates static spatial image inputs and does not assess model capacity to process multi-temporal time-series or fuse spatial imagery with non-image data (e.g., weather metrics, text annotations, or climate reanalysis).
    2. Geographic Coverage Gaps: While spanning six continents, the geographic coverage omits Antarctica and contains sparse representation across major biomes in Russia, China, South America, and North Africa.
    3. Aleatoric Performance Ceilings: As foundation models approach optimal performance, test metrics approach the intrinsic aleatoric uncertainty of remote sensing labels (such as noisy human annotations or low-resolution mixed pixels), leading to overlapping confidence intervals when distinguishing highly performant models.

Coverage note — None was omitted; all core benchmarking specifications, evaluation methodologies, mathematical discriminativity formulations, data schemas, empirical baseline findings, and stated limitations were fully incorporated.

References

  1. 1.Diab Abuaiadah and Alexander Switzer. Remote sensing dataset for detecting cows from high resolution aerial images. 2022.
  2. 2.Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare. Deep reinforcement learning at the edge of the statistical precipice. Advances in neural information processing systems, 34:29304–29320, 2021.
  3. 3.Hamed Alemohammad. The case for open-access ML-ready geospatial training data. In International Geoscience and Remote Sensing Symposium. IEEE, 2021.
  4. 4.Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling. The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research, 47:253–279, 2013.
  5. 5.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623, 2021.
  6. 6.Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  7. 7.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  8. 8.Marshall Burke, Anne Driscoll, David B Lobell, and Stefano Ermon. Using satellite imagery to understand and promote sustainable development. Science, 371(6535), 2021.
  9. 9.Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. arXiv preprint arXiv:2104.14294, 2021.
  10. 10.Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam. Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706.05587, 2017.
  11. 11.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020.
  12. 12.Elijah Cole, Benjamin Deneu, Titouan Lorieul, Maximilien Servajean, Christophe Botella, Dan Morris, Nebojsa Jojic, Pierre Bonnet, and Alexis Joly. The geolifeclef 2020 dataset. arXiv preprint arXiv:2004.04192, 2020.
  13. 13.Yezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu, Erik Rozi, Yutong He, Marshall Burke, David B Lobell, and Stefano Ermon. Satmae: Pre-training transformers for temporal and multi-spectral satellite imagery. arXiv preprint arXiv:2207.08051, 2022.
  14. 14.Walter T Dado, Jillian M Deines, Rinkal Patel, Sang-Zi Liang, and David B Lobell. High-resolution soybean yield mapping across the us midwest using subfield harvester data. Remote Sensing, 12(21):3471, 2020.
  15. 15.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  16. 16.Sonu Dileep, Daniel Zimmerle, J Ross Beveridge, and Timothy Vaughn. Automated identification of oil field features using cnns. 2020.
  17. 17.Ivica Dimitrovski, Ivan Kitanovski, Dragi Kocev, and Nikola Simidjievski. Current trends in deep learning for earth observation: An open-source benchmark arena for image classification. ISPRS Journal of Photogrammetry and Remote Sensing, 197:18–35, 2023.
  18. 18.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arxiv 2020. arXiv preprint arXiv:2010.11929, 2010.
  19. 19.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  20. 20.Matthias Drusch, Umberto Del Bello, Sébastien Carlier, Olivier Colin, Veronica Fernandez, Ferran Gascon, Bianca Hoersch, Claudia Isola, Paolo Laberinti, Philippe Martimort, et al. Sentinel-2: Esa’s optical high-resolution mission for gmes operational services. Remote sensing of Environment, 120:25–36, 2012.
  21. 21.B Eforn. Bootstrap methods: another look at the jackknife. The Annals of Statistics, 7:1–26, 1979.
  22. 22.EPA. Greenhouse Gas Emissions: Understanding Global Warming Potentials. Technical report, US Environmental Protection Agency, February 2017. URL https://www.epa.gov/ghgemissions/understanding-global-warming-potentials.
  23. 23.ESA. Sentinel-2. Technical report, European Space Agency, Paris, France, 2021. URL https://sentinel.esa.int/web/sentinel/missions/sentinel-2.
  24. 24.William Falcon and The PyTorch Lightning team. PyTorch Lightning, 3 2019. URL https://github.com/Lightning-AI/lightning.
  25. 25.Pedro F Felzenszwalb, Ross B Girshick, David McAllester, and Deva Ramanan. Object detection with discriminatively trained part-based models. IEEE transactions on pattern analysis and machine intelligence, 32(9):1627–1645, 2010.
  26. 26.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  27. 27.Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth. Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2019.
  28. 28.Dan Hendrycks, Mantas Mazeika, Saurav Kadavath, and Dawn Song. Using self-supervised learning can improve model robustness and uncertainty. arXiv preprint arXiv:1906.12340, 2019.
  29. 29.Jeremy Irvin, Hao Sheng, Neel Ramachandran, Sonja Johnson-Yu, Sharon Zhou, Kyle Story, Rose Rustowicz, Cooper Elsworth, Kemen Austin, and Andrew Y Ng. Forestnet: Classifying drivers of deforestation in indonesia using deep learning on satellite imagery. arXiv preprint arXiv:2011.05479, 2020.
  30. 30.Neal Jean, Sherrie Wang, Anshul Samar, George Azzari, David Lobell, and Stefano Ermon. Tile2vec: Unsupervised representation learning for spatially distributed data. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 3967–3974, 2019.
  31. 31.Forrest Johnson, Andrew Wlazlo, Ryan Keys, Viren Desai, Erin Wetherley, Ryan Calvert, and Elena Berman. Airborne methane surveys pay for themselves: An economic case study of increased revenue from emissions control. preprint, Environmental Monitoring, July 2021. URL http://eartharxiv.org/repository/view/2532/.
  32. 32.Siraput Jongaramrungruang, Christian Frankenberg, Andrew K. Thorpe, and Georgios Matheou. Methanet - an ai-driven approach to quantifying methane point-source emission from high-resolution 2-d plume imagery. ICML Workshop on Tackling Climate Change with AI, 2021.
  33. 33.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  34. 34.Hannah Kerner, Gabriel Tseng, Inbal Becker-Reshef, Catherine Nakalembe, Brian Barker, Blake Munshell, Madhava Paliyam, and Mehdi Hosseini. Rapid response crop maps in data sparse regions. arXiv preprint arXiv:2006.16866, 2020.
  35. 35.Alexandre Lacoste, Alexandra Luccioni, Victor Schmidt, and Thomas Dandres. Quantifying the carbon emissions of machine learning. arXiv preprint arXiv:1910.09700, 2019.
  36. 36.Issam Laradji, Pau Rodriguez, Freddie Kalaitzis, David Vazquez, Ross Young, Ed Davey, and Alexandre Lacoste. Counting cows: Tracking illegal cattle ranching from high-resolution satellite imagery. arXiv preprint arXiv:2011.07369, 2020.
  37. 37.Jihyeon Lee, Nina R. Brooks, Fahim Tajwar, Marshall Burke, Stefano Ermon, David B. Lobell, Debashish Biswas, and Stephen P. Luby. Scalable deep learning to identify brick kilns and aid regulatory capacity. Proceedings of the National Academy of Sciences, 118(17), 2021. ISSN 0027-8424. doi: 10.1073/pnas.2018863118. URL https://www.pnas.org/content/118/17/e2018863118.
  38. 38.Victor Lempitsky and Andrew Zisserman. Learning to count objects in images. Advances in neural information processing systems, 23, 2010.
  39. 39.Haifeng Li, Xin Dou, Chao Tao, Zhixiang Wu, Jie Chen, Jian Peng, Min Deng, and Ling Zhao. Rsi-cb: A large-scale remote sensing image classification benchmark using crowdsourced data. Sensors, 20(6):1594, 2020.
  40. 40.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12009–12019, 2022a.
  41. 41.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11976–11986, 2022b.
  42. 42.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  43. 43.Björn Lütjens, Lucas Liebenwein, and Katharina Kramer. Machine learning-based estimation of forest carbon stocks to increase transparency of forest preservation efforts. 2019 NeurIPS Workshop on Tackling Climate Change with AI (CCAI), 2019.
  44. 44.Björn Lütjens, Brandon Leshchinskiy, Christian Requena-Mesa, Farrukh Chishtie, Natalia Díaz-Rodríguez, Océane Boulais, Aruna Sankaranarayanan, Aaron Pina, Yarin Gal, Chedy Raissi, Alexander Lavin, and Dava Newman. Physically-consistent generative adversarial networks for coastal flood visualization. ICML Workshop on AI for Modeling Oceans and Climate Change (AIMOCC), 2021.
  45. 45.Lei Ma, Yu Liu, Xueliang Zhang, Yuanxin Ye, Gaofei Yin, and Brian Alan Johnson. Deep learning in remote sensing applications: A meta-analysis and review. ISPRS journal of photogrammetry and remote sensing, 152:166–177, 2019.
  46. 46.Oscar Manas, Alexandre Lacoste, Xavier Giro-i Nieto, David Vazquez, and Pau Rodriguez. Seasonal contrast: Unsupervised pre-training from uncurated remote sensing data. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9414–9423, 2021.
  47. 47.M Maskey, H Alemohammad, KJ Murphy, and R Ramachandran. Advancing ai for earth science: A data systems perspective. Eos, 101, 2020.
  48. 48.Kevin Mayer, Benjamin Rausch, Marie-Louise Arlt, Gunther Gust, Zhecheng Wang, Dirk Neumann, and Ram Rajagopal. 3d-pv-locator: Large-scale detection of rooftop-mounted photovoltaic systems in 3d. Applied Energy, 310:118469, 2022. ISSN 0306-2619. doi: https://doi.org/10.1016/j.apenergy.2021.118469. URL https://www.sciencedirect.com/science/article/pii/S0306261921016937.
  49. 49.Amy McGovern, Kimberly L. Elmore, David John Gagne, Sue Ellen Haupt, Christopher D. Karstens, Ryan Lagerquist, Travis Smith, and John K. Williams. Using artificial intelligence to improve real-time decision-making for high-impact weather. Bulletin of the American Meteorological Society, 98(10), 2017.
  50. 50.T. Mitchell, W. Cohen, E. Hruschka, P. Talukdar, B. Yang, J. Betteridge, A. Carlson, B. Dalvi, M. Gardner, B. Kisiel, J. Krishnamurthy, N. Lao, K. Mazaitis, T. Mohamed, N. Nakashole, E. Platanios, A. Ritter, M. Samadi, B. Settles, R. Wang, D. Wijaya, A. Gupta, X. Chen, A. Saparov, M. Greaves, and J. Welling. Never-ending learning. Communications of the ACM, 61(5):103–115, April 2018. ISSN 0001-0782, 1557-7317. doi: 10.1145/3191513. URL https://dl.acm.org/doi/10.1145/3191513.
  51. 51.Cassandra Pallai and Kathryn Wesson. Chesapeake bay program partnership high-resolution land cover classification accuracy assessment methodology, 2017. URL https://chesapeakeconservancy.org/wp-content/uploads/2017/01/Chesapeake_Conservancy_Accuracy_Assessment_Methodology.pdf.
  52. 52.David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350, 2021.
  53. 53.Otávio AB Penatti, Keiller Nogueira, and Jefersson A Dos Santos. Do deep features generalize from everyday objects to remote sensing and aerial scenes domains? In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pages 44–51, 2015.
  54. 54.Karissa Pepin, Howard A. Zebker, and William Ellsworth. High-Pass Filters to Reduce the Effects of Broad Atmospheric Contributions in Sbas Inversions: A Case Study in the Delaware Basin. In IGARSS 2020 - 2020 IEEE International Geoscience and Remote Sensing Symposium, pages 1030–1033, Waikoloa, HI, USA, September 2020. IEEE. ISBN 978-1-72816-374-1. doi: 10.1109/IGARSS39084.2020.9324656. URL https://ieeexplore.ieee.org/document/9324656/.
  55. 55.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. arXiv preprint arXiv:2103.00020, 2021.
  56. 56.Caleb Robinson, Le Hou, Kolya Malkin, Rachel Soobitsky, Jacob Czawlytko, Bistra Dilkina, and Nebojsa Jojic. Large scale high-resolution land cover mapping with multi-resolution data. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 12726–12735, 2019.
  57. 57.David Rolnick, Priya L Donti, Lynn H Kaack, Kelly Kochanski, Alexandre Lacoste, Kris Sankaran, Andrew Slavin Ross, Nikola Milojevic-Dupont, Natasha Jaques, Anna Waldman-Brown, et al. Tackling climate change with machine learning. arXiv preprint arXiv:1906.05433, 2019.
  58. 58.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015.
  59. 59.Victor Schmidt, Kamal Goyal, Aditya Joshi, Boris Feld, Liam Conell, Nikolas Laskaris, Doug Blank, Jonathan Wilson, Sorelle Friedler, and Sasha Luccioni. CodeCarbon: Estimate and Track Carbon Emissions from Machine Learning Computing. 2021. doi: 10.5281/zenodo.4658424.
  60. 60.Roy Schwartz, Jesse Dodge, Noah A Smith, and Oren Etzioni. Green ai. Communications of the ACM, 63(12):54–63, 2020.
  61. 61.Hao Sheng, Jeremy Irvin, Sasankh Munukutla, Shawn Zhang, Christopher Cross, Kyle Story, Rose Rustowicz, Cooper Elsworth, Zutao Yang, Mark Omara, et al. Ognet: Towards a global oil and gas infrastructure database using deep learning on remotely sensed imagery. arXiv preprint arXiv:2011.07227, 2020.
  62. 62.Adam J Stewart, Caleb Robinson, Isaac A Corley, Anthony Ortiz, Juan M Lavista Ferres, and Arindam Banerjee. Torchgeo: deep learning with geospatial data. arXiv preprint arXiv:2111.08872, 2021.
  63. 63.Emma Strubell, Ananya Ganesh, and Andrew McCallum. Energy and policy considerations for deep learning in nlp. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3645–3650, 2019.
  64. 64.Gencer Sumbul, Arne De Wall, Tristan Kreuziger, Filipe Marcelino, Hugo Costa, Pedro Benevides, Mario Caetano, Begüm Demir, and Volker Markl. Bigearthnet-mm: A large-scale, multimodal, multilabel benchmark archive for remote sensing image classification and retrieval [software and data sets]. IEEE Geoscience and Remote Sensing Magazine, 9(3):174–180, 2021.
  65. 65.Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan, Cordelia Schmid, and Phillip Isola. What makes for good views for contrastive learning? Advances in Neural Information Processing Systems, 33:6827–6839, 2020.
  66. 66.USGS. Landsat 8. Technical report, United States Geological Survey, Reston, Virginia, USA, 2021. URL https://www.usgs.gov/core-science-systems/nli/landsat/landsat-8?qt-science_support_page_related_con=0#qt-science_support_page_related_con.
  67. 67.Burak Uzkent, Evan Sheehan, Chenlin Meng, Zhongyi Tang, Marshall Burke, David Lobell, and Stefano Ermon. Learning to interpret satellite images using wikipedia. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, 2019.
  68. 68.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  69. 69.Di Wang, Jing Zhang, Bo Du, Gui-Song Xia, and Dacheng Tao. An empirical study of remote sensing pretraining. IEEE Transactions on Geoscience and Remote Sensing, pages 1–1, 2022. doi: 10.1109/TGRS.2022.3176603.
  70. 70.Yi Wang, Nassim Ait Ali Braham, Zhitong Xiong, Chenying Liu, Conrad M Albrecht, and Xiao Xiang Zhu. Ssl4eo-s12: A large-scale multi-modal, multi-temporal dataset for self-supervised learning in earth observation.
  71. 71.Ben G Weinstein, Sarah J Graves, Sergio Marconi, Aditya Singh, Alina Zare, Dylan Stewart, Stephanie A Bohlman, and Ethan P White. A benchmark dataset for canopy crown detection and delineation in co-registered airborne rgb, lidar and hyperspectral imagery from the national ecological observation network. PLoS computational biology, 17(7):e1009180, 2021.
  72. 72.Zhitong Xiong, Fahong Zhang, Yi Wang, Yilei Shi, and Xiao Xiang Zhu. Earthnets: Empowering ai in earth observation. arXiv preprint arXiv:2210.04936, 2022.
  73. 73.Christopher Yeh, Chenlin Meng, Sherrie Wang, Anne Driscoll, Erik Rozi, Patrick Liu, Jihyeon Lee, Marshall Burke, David B Lobell, and Stefano Ermon. Sustainbench: Benchmarks for monitoring the sustainable development goals with machine learning. arXiv preprint arXiv:2111.04724, 2021.
  74. 74.Jin Z., Lin C., Weigl C., Obarowski J., and Hale D. Smallholder cashew plantations in benin, 2021.
  75. 75.Valentina Zantedeschi, Fabrizio Falasca, Alyson Douglas, Richard Strange, Matt J Kusner, and Duncan Watson-Parris. Cumulo: A dataset for learning cloud classes. arXiv preprint arXiv:1911.04227, 2019.
  76. 76.Xiao Xiang Zhu, Devis Tuia, Lichao Mou, Gui-Song Xia, Liangpei Zhang, Feng Xu, and Friedrich Fraundorfer. Deep learning in remote sensing: A comprehensive review and list of resources. IEEE Geoscience and Remote Sensing Magazine, 5(4):8–36, 2017.
  77. 77.Xiao Xiang Zhu, Jingliang Hu, Chunping Qiu, Yilei Shi, Jian Kang, Lichao Mou, Hossein Bagheri, Matthias Häberle, Yuansheng Hua, Rong Huang, et al. So2sat lcz42: A benchmark dataset for global local climate zones classification. arXiv preprint arXiv:1912.12171, 2019.

Citation

MLA
Lacoste, A., et al. “GEO-Bench: Toward Foundation Models for Earth Monitoring”. arXiv, 2023, http://arxiv.org/abs/2306.03831v2.
APA
Lacoste, A., Lehmann, N., Rodriguez, P., Sherwin, E. D., Kerner, H., Lütjens, B., Irvin, J. A., Dao, D., Alemohammad, H., Drouin, A., Gunturkun, M., Huang, G., Vazquez, D., Newman, D., Bengio, Y., Ermon, S., & Zhu, X. X. (2023). GEO-Bench: Toward Foundation Models for Earth Monitoring. arXiv. http://arxiv.org/abs/2306.03831v2
Chicago
Lacoste, A., N. Lehmann, P. Rodriguez, et al. 2023. “GEO-Bench: Toward Foundation Models for Earth Monitoring”. arXiv. http://arxiv.org/abs/2306.03831v2.
Harvard
Lacoste, A. et al. (2023) “GEO-Bench: Toward Foundation Models for Earth Monitoring”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.03831v2.
Vancouver
1. Lacoste A, Lehmann N, Rodriguez P, et al (2023) GEO-Bench: Toward Foundation Models for Earth Monitoring. arXiv

BibTeX

@article{lacoste2023geo,
  title = {GEO-Bench: Toward Foundation Models for Earth Monitoring},
  author = {Lacoste, Alexandre and Lehmann, Nils and Rodriguez, Pau and Sherwin, Evan David and Kerner, Hannah and Lütjens, Björn and Irvin, Jeremy Andrew and Dao, David and Alemohammad, Hamed and Drouin, Alexandre and Gunturkun, Mehmet and Huang, Gabriel and Vazquez, David and Newman, Dava and Bengio, Yoshua and Ermon, Stefano and Zhu, Xiao Xiang},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.03831v2},
  eprint = {2306.03831}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission