GEO-Bench: Toward Foundation Models for Earth Monitoring
Alexandre LacosteNils LehmannPau RodríguezEvan D. SherwinHannah KernerBjörn LütjensJeremy IrvinDavid DaoHamed AlemohammadAlexandre Drouin
Introduces GEO-Bench, a standardized benchmark of twelve classification and segmentation tasks alongside twenty baseline evaluations to accelerate the development and reliable assessment of foundation models for satellite Earth monitoring.
Earth monitoring via satellite imagery is critical for climate change mitigation, disaster response, and environmental protection. However, current artificial intelligence models built for standard photography perform poorly on remote sensing tasks due to overhead viewpoints, multispectral sensor bands, and temporal variations. In addition, existing research evaluates Earth observation models inconsistently across narrow, non-standardized datasets, preventing objective comparisons and slowing practical progress.
The article introduces GEO-Bench, an open-source evaluation benchmark designed to standardize and accelerate the development of machine learning foundation models for Earth monitoring across diverse real-world tasks.
To construct the benchmark, domain experts assembled twelve distinct geospatial datasets spanning global geographies—six image classification and six semantic segmentation tasks covering multispectral, radar, and aerial imagery. The datasets were standardized into accessible sizes, balanced across classes, and formatted to run on single graphics processing units to reduce compute costs and carbon emissions. The authors established an evaluation protocol that benchmarks twenty baseline models across varying training data sizes using ten random seeds, robust statistical aggregation via interquartile means, and bootstrapped confidence intervals.
The findings show that standard computer vision pre-training significantly outperforms models trained from scratch, with modern architectures such as ConvNeXt and Swin Transformer systematically achieving top performance. Remarkably, pre-trained ConvNeXt models matched the accuracy of baseline models trained from scratch while requiring only 2% of the training data, demonstrating a fifty-fold improvement in data efficiency. Conversely, existing remote sensing models pre-trained specifically on satellite data did not demonstrate significant gains over models pre-trained on standard visual datasets, showing only modest gains when incorporating multispectral bands in convolutional architectures and performance declines in vision transformers.
These results highlight substantial opportunities to reduce labeling costs and compute timelines by adopting modern architectures and pre-trained weights for environmental monitoring. The lack of distinct outperformance from current satellite-specific pre-training suggests that existing geospatial foundation models require substantial architectural improvements to properly exploit multimodal remote sensing data.
Organizations developing Earth observation systems should adopt standardized evaluation workflows and prioritize modern convolutional and transformer backbones over older architectures. Future research should develop novel multimodal pre-training frameworks capable of effectively fusing temporal revisits, weather data, and multispectral information. While the benchmark provides strong coverage, its current limitations include data gaps in regions such as South America and parts of Asia, as well as a focus on single-image inputs rather than time-series sequences.
- Paper: On the Opportunities and Risks of Foundation Models, Rishi Bommasani et al. (2021). Establishes the foundational paradigms, terminology, and evaluation challenges of large-scale pre-trained foundation models that GEO-Bench adapts to Earth monitoring.
- Paper: EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification, Patrick Helber et al. (2017). Introduces the standard multi-spectral Sentinel-2 benchmark for land use and land cover classification that forms the basis for satellite image evaluation in GEO-Bench.
- Paper: Deep learning in remote sensing: a review, Xiao Xiang Zhu et al. (2017). Provides a comprehensive foundation of deep learning architectures, multi-sensor modalities, and classical benchmarks across remote sensing tasks.
- Paper: Masked Autoencoders Are Scalable Vision Learners, Kaiming He et al. (2022). Introduces masked autoencoding as a leading self-supervised pre-training paradigm for Vision Transformers evaluated across the GEO-Bench suite.
- Paper: WILDS: A Benchmark of in-the-Wild Distribution Shifts, Pang Wei Koh et al. (2020). Establishes benchmark curation and evaluation methodologies for multi-task distribution shifts in real-world deployments, providing methodological precedent for GEO-Bench.
- Paper: Remote Sensing Image Scene Classification: Benchmark and State of the Art, Gong Cheng et al. (2017). Surveys and standardizes remote sensing scene classification datasets and baseline evaluation methodologies.
- Paper: Object Detection in Optical Remote Sensing Images: A Survey and A New Benchmark, Ke Li et al. (2019). Examines the distinct visual and spatial challenges of overhead optical remote sensing data compared to standard computer vision benchmarks.
- Paper: AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene Classification, Gui-Song Xia et al. (2016). Presents a large-scale diverse aerial scene classification dataset highlighting the necessity of diverse evaluation benchmarks in Earth observation.
- Paper: DINOv3, Oriane Siméoni et al. (2025). Develops a next-generation visual foundation model family that directly tackles the dense prediction and remote sensing performance challenges highlighted in GEO-Bench.
- Paper: Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data, Lihe Yang et al. (2024). Demonstrates how large-scale unlabeled visual data pre-training can be scaled and transferred effectively to downstream dense prediction tasks.
