GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models
Gilles Quentin HachemeGirmaw Abebe TadesseCaleb RobinsonAkram ZaytarRahul DodhiaJuan M. Lavista Ferres
Introduces GeoVision Labeler, a strictly zero-shot framework that pairs vision-language model descriptions with hierarchical LLM reasoning to classify complex satellite imagery without domain-specific training or labeled data.
Rapid and accurate classification of satellite imagery is critical for applications such as disaster response, urban planning, and environmental monitoring. However, traditional machine learning methods require large, manually labeled datasets that are expensive to build and often unavailable during emergencies. While existing zero-shot approaches attempt to classify images without labeled training examples, they still depend on specialized domain pretraining or synthetic data. The article addresses this operational bottleneck by introducing and evaluating the GeoVision Labeler (GVL), an open-source framework designed to achieve strict zero-shot geospatial image classification without requiring any task-specific fine-tuning or domain adaptation.
The GVL framework uses a two-stage, modular approach that leverages the complementary strengths of vision large language models and standard text large language models. In the first stage, a vision model examines an input image patch and generates a detailed, human-readable textual description. In the second stage, a text model reads this description and assigns it to a user-defined category, relying on a standard image-text matching fallback only if the text model fails to return a valid class. For complex classification problems with many overlapping categories, the authors implemented a recursive clustering technique that groups similar categories into broader meta-classes, allowing the system to perform hierarchical classification from coarse to fine distinctions. The framework was evaluated across three standard remote sensing benchmarks: SpaceNet v7 (531 patches), UC Merced (420 test images across 21 classes), and RESISC45 (6,300 test images across 45 classes).
The evaluation produced four key findings. First, on well-separated binary tasks, GVL achieved high accuracy without prior training; on SpaceNet v7, it reached up to 93.2% overall accuracy distinguishing buildings from non-buildings, outperforming a baseline image-text model by roughly 34 percentage points. Second, including geographic context in the prompt improved binary accuracy, whereas forcing large lists of target classes into vision model prompts diluted description quality and reduced multi-class accuracy by 10 to 19 percentage points. Third, hierarchical clustering effectively mitigated confusion between subtly different classes; top-level grouping into coarse meta-classes raised accuracy to 86.4% on UC Merced and 84.3% on RESISC45. Fourth, at the finest level of granularity across 45 distinct classes, zero-shot accuracy dropped to approximately 45.3%, underscoring the challenge of separating fine-grained categories without supervised training.
These findings indicate that GVL provides a viable, low-cost solution for rapid image screening and automated preliminary labeling, significantly reducing the time needed to extract actionable intelligence from satellite data. Because classifications are derived from intermediate textual descriptions, the pipeline offers human-readable explanations that improve decision-making transparency compared to black-box models. While GVL does not match the peak performance of fully supervised or domain-adapted models on intricate, fine-grained taxonomies, its plug-and-play architecture allows organizations to immediately deploy satellite classification workflows without collecting training labels.
Organizations should consider deploying GVL for rapid initial triage, visual search, and weak label generation to accelerate manual mapping workflows, especially in low-resource or time-critical settings. For complex label sets, teams should adopt hierarchical meta-class grouping and avoid overloading prompts with extensive class lists. Key limitations include sensitivity to prompt design, increased computational latency from multi-step hierarchical inference, and a current restriction to standard three-channel visible imagery rather than multi-spectral data. Confidence is highest for coarse and binary categorization tasks, while operational use on fine-grained classes will require further testing or future integrations with multi-spectral vision models.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. This foundational work establishes zero-shot image classification via natural language supervision, the core paradigm that GeoVision Labeler adapts and critiques for geospatial imagery.
- Paper: Remote Sensing Image Scene Classification: Benchmark and State of the Art, Gong Cheng et al. (2017). This paper introduces the RESISC45 benchmark dataset, which serves as one of the primary evaluation testbeds for multi-class hierarchical zero-shot classification in the source.
- Paper: Vision-Language Models for Vision Tasks: A Survey, Jingyi Zhang et al. (2023). This survey provides essential background on vision-language model architectures, pre-training objectives, and transfer mechanisms underpinning modular vision-LLM pipelines.
- Paper: GEO-Bench: Toward Foundation Models for Earth Monitoring, Alexandre Lacoste et al. (2023). This work formalizes standard evaluation protocols and foundation model benchmarks for Earth observation tasks, defining the challenges in satellite imagery classification addressed by GeoVision Labeler.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). This paper details open multimodal architectures (LLaVA) that integrate vision encoders with large language models to generate rich visual descriptions from images.
- Paper: Zero-Shot Learning—A Comprehensive Evaluation of the Good, the Bad and the Ugly, Yongqin Xian et al. (2017). This comprehensive evaluation outlines standard methodologies, class split definitions, and evaluation metrics for zero-shot recognition frameworks.
No sufficiently relevant recommendations were found.
