GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models

Gilles Quentin HachemeGirmaw Abebe TadesseCaleb RobinsonAkram ZaytarRahul DodhiaJuan M. Lavista Ferres

article2025arXiv1 citations

Introduces GeoVision Labeler, a strictly zero-shot framework that pairs vision-language model descriptions with hierarchical LLM reasoning to classify complex satellite imagery without domain-specific training or labeled data.

Listen

Rapid and accurate classification of satellite imagery is critical for applications such as disaster response, urban planning, and environmental monitoring. However, traditional machine learning methods require large, manually labeled datasets that are expensive to build and often unavailable during emergencies. While existing zero-shot approaches attempt to classify images without labeled training examples, they still depend on specialized domain pretraining or synthetic data. The article addresses this operational bottleneck by introducing and evaluating the GeoVision Labeler (GVL), an open-source framework designed to achieve strict zero-shot geospatial image classification without requiring any task-specific fine-tuning or domain adaptation.

The GVL framework uses a two-stage, modular approach that leverages the complementary strengths of vision large language models and standard text large language models. In the first stage, a vision model examines an input image patch and generates a detailed, human-readable textual description. In the second stage, a text model reads this description and assigns it to a user-defined category, relying on a standard image-text matching fallback only if the text model fails to return a valid class. For complex classification problems with many overlapping categories, the authors implemented a recursive clustering technique that groups similar categories into broader meta-classes, allowing the system to perform hierarchical classification from coarse to fine distinctions. The framework was evaluated across three standard remote sensing benchmarks: SpaceNet v7 (531 patches), UC Merced (420 test images across 21 classes), and RESISC45 (6,300 test images across 45 classes).

The evaluation produced four key findings. First, on well-separated binary tasks, GVL achieved high accuracy without prior training; on SpaceNet v7, it reached up to 93.2% overall accuracy distinguishing buildings from non-buildings, outperforming a baseline image-text model by roughly 34 percentage points. Second, including geographic context in the prompt improved binary accuracy, whereas forcing large lists of target classes into vision model prompts diluted description quality and reduced multi-class accuracy by 10 to 19 percentage points. Third, hierarchical clustering effectively mitigated confusion between subtly different classes; top-level grouping into coarse meta-classes raised accuracy to 86.4% on UC Merced and 84.3% on RESISC45. Fourth, at the finest level of granularity across 45 distinct classes, zero-shot accuracy dropped to approximately 45.3%, underscoring the challenge of separating fine-grained categories without supervised training.

These findings indicate that GVL provides a viable, low-cost solution for rapid image screening and automated preliminary labeling, significantly reducing the time needed to extract actionable intelligence from satellite data. Because classifications are derived from intermediate textual descriptions, the pipeline offers human-readable explanations that improve decision-making transparency compared to black-box models. While GVL does not match the peak performance of fully supervised or domain-adapted models on intricate, fine-grained taxonomies, its plug-and-play architecture allows organizations to immediately deploy satellite classification workflows without collecting training labels.

Organizations should consider deploying GVL for rapid initial triage, visual search, and weak label generation to accelerate manual mapping workflows, especially in low-resource or time-critical settings. For complex label sets, teams should adopt hierarchical meta-class grouping and avoid overloading prompts with extensive class lists. Key limitations include sensitivity to prompt design, increased computational latency from multi-step hierarchical inference, and a current restriction to standard three-channel visible imagery rather than multi-spectral data. Confidence is highest for coarse and binary categorization tasks, while operational use on fine-grained classes will require further testing or future integrations with multi-spectral vision models.

No sufficiently relevant recommendations were found.

Cover for GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models

Abstract

Classifying geospatial imagery remains a major bottleneck for applications such as disaster response and land-use monitoring-particularly in regions where annotated data is scarce or unavailable. Existing tools (e.g., RS-CLIP) that claim zero-shot classification capabilities for satellite imagery nonetheless rely on task-specific pretraining and adaptation to reach competitive performance. We introduce GeoVision Labeler (GVL), a strictly zero-shot classification framework: a vision Large Language Model (vLLM) generates rich, human-readable image descriptions, which are then mapped to user-defined classes by a conventional Large Language Model (LLM). This modular, and interpretable pipeline enables flexible image classification for a large range of use cases. We evaluated GVL across three benchmarks-SpaceNet v7, UC Merced, and RESISC45. It achieves up to 93.2% zero-shot accuracy on the binary Buildings vs. No Buildings task on SpaceNet v7. For complex multi-class classification tasks (UC Merced, RESISC45), we implemented a recursive LLM-driven clustering to form meta-classes at successive depths, followed by hierarchical classification-first resolving coarse groups, then finer distinctions-to deliver competitive zero-shot performance. GVL is open-sourced at this https URL to catalyze adoption in real-world geospatial workflows.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Datasets
  • 4 Methodology
  • 5 Results
  • 6 Discussion
  • 7 Limitations
  • 8 Conclusion
  • References
  • A Appendix

Knowls

  1. Knowl 1 — GeoVision Labeler (GVL) Two-Stage Pipeline Architecture

    model/method

    GeoVision Labeler (GVL) is a modular, strictly zero-shot geospatial image classification framework that operates without task-specific pretraining, domain adaptation, or fine-tuning. It decomposes classification into two sequential stages:

    1. Stage 1: Image Description Generation via vLLM: A vision Large Language Model (vLLM, such as Kosmos-2 or Llama 3.2-Vision-Instruct) processes an input satellite image or local image patch alongside a text prompt. The prompt includes a context directive (e.g., "This is a satellite image"), instructions to generate a detailed visual description, and optionally domain context (e.g., geographic coordinates embedded in the filename) or candidate target class names. The vLLM outputs a human-readable textual description of the visual scene.

    2. Stage 2: Semantic Text Classification via LLM: A text Large Language Model (LLM, such as GPT-4o, Llama-3.1, or Phi-3) receives the generated textual description and a predefined list of candidate target classes C={c1,c2,…,cN}\mathcal{C} = \{c_1, c_2, \dots, c_N\}. The LLM performs semantic reasoning to select the single most appropriate class label from C\mathcal{C}.

    3. Fallback Mechanism via CLIP: If the LLM classifier produces an invalid label (i.e., a string not in C\mathcal{C}), a pretrained CLIP model acts as a fallback. CLIP classifies the image directly by computing the cosine similarity between the visual embedding of the image patch and text embeddings of class prompts:

    y^=arg⁡max⁡c∈Ceimg⋅etext(c)∥eimg∥∥etext(c)∥\hat{y} = \arg\max_{c \in \mathcal{C}} \frac{\mathbf{e}_{\text{img}} \cdot \mathbf{e}_{\text{text}}(c)}{\|\mathbf{e}_{\text{img}}\| \|\mathbf{e}_{\text{text}}(c)\|}

    where eimg\mathbf{e}_{\text{img}} is the CLIP visual embedding of the image and etext(c)\mathbf{e}_{\text{text}}(c) is the text embedding for class label c∈Cc \in \mathcal{C}.

  2. Knowl 2 — LLM-Based Recursive Semantic Class Clustering Algorithm

    algorithm

    To handle large or visually ambiguous class sets without supervised retraining, GVL constructs a multi-level taxonomy of meta-classes using an LLM-driven semantic clustering procedure that recursively groups classes by semantic proximity.

    Input: Set of original class labels C={c1,c2,…,cN}\mathcal{C} = \{c_1, c_2, \dots, c_N\}, sequence of cluster target counts [K1,K2,…,KD][K_1, K_2, \dots, K_D] for hierarchy depth DD
    Output: Hierarchical taxonomy tree mapping original classes to meta-classes at depths 1,…,D1, \dots, D
    function ClusterClasses(classes S\mathcal{S}, cluster count KK)
        Prompt LLM: "Suggest KK non-overlapping category names for the following labels based on semantic similarity: S\mathcal{S}. Output in the form Cluster_1: [Name], ..., Cluster_K: [Name]."
        Parse and strip whitespace/punctuation to obtain meta-class names {M1,M2,…,MK}\{M_1, M_2, \dots, M_K\}
        for each label c∈Sc \in \mathcal{S} do
            Prompt LLM: "Assign this label cc to one of the categories {M1,…,MK}\{M_1, \dots, M_K\}. Output in the form Cluster: [Name]."
            if returned name matches MjM_j then
                Assign cc to cluster MjM_j
            else if substring matching heuristic matches MjM_j then
                Assign cc to cluster MjM_j
            else
                Assign cc to "Unknown" bucket
            end if
        end for
        Discard empty meta-classes to produce final clusters {M1,…,MK~}\{M_1, \dots, M_{\tilde{K}}\} with K~≤K\tilde{K} \le K
        return clusters and label assignments
    end function
    function BuildHierarchy(classes S\mathcal{S}, depths [K1,…,KD][K_1, \dots, K_D], current depth dd)
        clusters ←ClusterClasses(S,Kd)\leftarrow \text{ClusterClasses}(\mathcal{S}, K_d)
        if d==Dd == D or all clusters in clusters are singletons then
            return clusters
        end if
        for each meta-class Mj∈clustersM_j \in \text{clusters} do
            subclasses ←{c∈S∣c is assigned to Mj}\leftarrow \{c \in \mathcal{S} \mid c \text{ is assigned to } M_j\}
            if ∣subclasses∣>1|\text{subclasses}| > 1 then
                Mj.children←BuildHierarchy(subclasses,[K1,…,KD],d+1)M_j.\text{children} \leftarrow \text{BuildHierarchy}(\text{subclasses}, [K_1, \dots, K_D], d + 1)
            end if
        end for
        return clusters
    end function

    During inference, GVL classifies top-level meta-classes (D=0D=0) first, and subsequently refines predictions down to finer sub-classes at depths D=1,2D=1, 2.

  3. Knowl 3 — SpaceNet v7 Zero-Shot Binary Building Detection Results

    data/table

    The SpaceNet v7 benchmark evaluation consists of 59 scenes (1024×10241024 \times 1024 pixels) across 100 global locations. Each scene is partitioned into 9 equal patches (59×9=53159 \times 9 = 531 image patches) labeled as Buildings (if any pixel overlaps a building footprint) or No Buildings (otherwise), with a random timestamp selected per scene.

    The table below reports zero-shot Overall Accuracy (OA) comparing standalone CLIP with various GVL configurations (pairing vLLMs with LLM classifiers, with/without target classes in the vLLM prompt, and with/without geo-context coordinates from filenames):

    Pipeline Classifier Classes Geo-context Kosmos 2 Llama 3.2
    CLIP (Standalone) — — — 0.588
    GVL (Ours) Llama-3.1 ✓ ×\times 0.859 0.776
    GVL (Ours) Llama-3.1 ✓ ✓ 0.910 0.699
    GVL (Ours) Llama-3.1 ×\times ×\times 0.889 0.821
    GVL (Ours) Llama-3.1 ×\times ✓ 0.859 0.751
    GVL (Ours) Phi-3 ✓ ×\times 0.932 0.857
    GVL (Ours) Phi-3 ✓ ✓ 0.928 0.902
    GVL (Ours) Phi-3 ×\times ×\times 0.928 0.912
    GVL (Ours) Phi-3 ×\times ✓ 0.932 0.927
    GVL (Ours) GPT-4o ✓ ×\times 0.878 0.789
    GVL (Ours) GPT-4o ✓ ✓ 0.917 0.799
    GVL (Ours) GPT-4o ×\times ×\times 0.896 0.832
    GVL (Ours) GPT-4o ×\times ✓ 0.876 0.783

    Key takeaways:

    1. All GVL configurations substantially exceed the standalone CLIP zero-shot baseline of 0.588 OA, reaching up to 0.932 OA (Kosmos 2 + Phi-3).
    2. Kosmos 2 outperforms Llama 3.2 by up to 20 percentage points across equivalent configurations, indicating superior visual grounding for remote sensing features.
    3. Incorporating geographic context from image filenames alongside target class names boosts Kosmos 2 accuracy from 0.859 to 0.910 with Llama-3.1 and from 0.878 to 0.917 with GPT-4o.
  4. Knowl 4 — UC Merced Flat and Hierarchical Zero-Shot Classification Results

    data/table

    The UC Merced dataset contains 2,100 RGB images (256×256256 \times 256 pixels) spanning 21 land-use classes, evaluated on its 420-image test set.

    Flat 21-Class Zero-Shot Accuracy:

    Pipeline Classifier Classes in Prompt Other Baselines Kosmos 2 Llama 3.2
    CLIP (Standalone) — — 0.710 — —
    ResNet50 (Fine-tuned) — — 0.907 — —
    RS-CLIP (Adapted zero-shot) — — 0.959 — —
    GVL Llama-3.1 ✓ — 0.514 0.600
    GVL Llama-3.1 ×\times — 0.619 0.610
    GVL Phi-3 ✓ — 0.521 0.626
    GVL Phi-3 ×\times — 0.648 0.624
    GVL GPT-4o ✓ — 0.517 0.714
    GVL GPT-4o ×\times — 0.710 0.710

    Hierarchical Two-Level Classification Accuracy: The 21 classes were grouped into 5 meta-classes at depth D=0D=0, and D=1D=1 corresponds to predicting the original 21 classes:

    Clustering LLM Classifier LLM Classes in Prompt Depth DD Kosmos 2 Llama 3.2
    Llama-3.1 Llama-3.1 ×\times 0 0.738 0.755
    1 0.498 0.510
    ✓ 0 0.700 0.726
    1 0.452 0.460
    Phi-3 ×\times 0 0.612 0.617
    1 0.393 0.391
    ✓ 0 0.562 0.614
    1 0.350 0.343
    GPT-4o ×\times 0 0.679 0.679
    1 0.483 0.507
    ✓ 0 0.643 0.681
    1 0.457 0.502
    GPT-4o Llama-3.1 ×\times 0 0.795 0.743
    1 0.543 0.460
    ✓ 0 0.762 0.760
    1 0.548 0.460
    Phi-3 ×\times 0 0.857 0.807
    1 0.595 0.543
    ✓ 0 0.771 0.795
    1 0.510 0.486
    GPT-4o ×\times 0 0.864 0.807
    1 0.662 0.602
    ✓ 0 0.757 0.836
    1 0.593 0.581

    Key takeaways:

    1. On flat classification, GVL achieves a maximum OA of 0.714 (Llama 3.2 + GPT-4o).
    2. Grouping classes into coarse meta-classes (D=0D=0) elevates zero-shot OA up to 0.864 (GPT-4o clustering and classification with Kosmos 2).
    3. Meta-class taxonomies generated by GPT-4o systematically outperform those from Llama-3.1 across all classifiers and vLLMs.
  5. Knowl 5 — RESISC45 Flat and Hierarchical Zero-Shot Classification Results

    data/table

    The RESISC45 benchmark comprises 31,500 RGB images (256×256256 \times 256 pixels) covering 45 scene categories (700 images per class), evaluated on the 6,300-image test set.

    Flat 45-Class Zero-Shot Accuracy:

    Pipeline Classifier Classes in Prompt Other Baselines Kosmos 2 Llama 3.2
    CLIP (Standalone) — — 0.610 — —
    ResNet50 (Fine-tuned) — — 0.775 — —
    RS-CLIP (Adapted zero-shot) — — 0.858 — —
    GVL Llama-3.1 ✓ — 0.272 0.450
    GVL Llama-3.1 ×\times — 0.476 0.440
    GVL Phi-3 ✓ — 0.344 0.461
    GVL Phi-3 ×\times — 0.523 0.473
    GVL GPT-4o ✓ — 0.352 0.540
    GVL GPT-4o ×\times — 0.565 0.539

    Hierarchical Three-Level Classification Accuracy: The 45 classes are structured into 4 top-level meta-classes (D=0D=0), 3 mid-level meta-classes per top group (D=1D=1), and 45 original classes (D=2D=2):

    Clustering LLM Classifier LLM Classes in Prompt Depth D=0D=0 Depth D=1D=1 Depth D=2D=2 vLLM
    Llama-3.1 Llama-3.1 ×\times 0.628 / 0.563 0.486 / 0.423 0.330 / 0.289 Kosmos 2 / Llama 3.2
    ✓ 0.571 / 0.576 0.398 / 0.400 0.257 / 0.249 Kosmos 2 / Llama 3.2
    Phi-3 ×\times 0.694 / 0.650 0.561 / 0.501 0.370 / 0.331 Kosmos 2 / Llama 3.2
    ✓ 0.608 / 0.621 0.461 / 0.442 0.308 / 0.280 Kosmos 2 / Llama 3.2
    GPT-4o ×\times 0.606 / 0.581 0.494 / 0.470 0.325 / 0.323 Kosmos 2 / Llama 3.2
    ✓ 0.557 / 0.603 0.421 / 0.465 0.260 / 0.307 Kosmos 2 / Llama 3.2
    GPT-4o Llama-3.1 ×\times 0.818 / 0.803 0.586 / 0.563 0.390 / 0.371 Kosmos 2 / Llama 3.2
    ✓ 0.790 / 0.810 0.550 / 0.577 0.350 / 0.365 Kosmos 2 / Llama 3.2
    Phi-3 ×\times 0.828 / 0.801 0.592 / 0.530 0.401 / 0.346 Kosmos 2 / Llama 3.2
    ✓ 0.811 / 0.791 0.571 / 0.541 0.358 / 0.341 Kosmos 2 / Llama 3.2
    GPT-4o ×\times 0.843 / 0.823 0.648 / 0.622 0.453 / 0.446 Kosmos 2 / Llama 3.2
    ✓ 0.809 / 0.825 0.618 / 0.635 0.413 / 0.441 Kosmos 2 / Llama 3.2

    Key takeaways:

    1. On flat classification, excluding class names from the vLLM prompt achieves the highest zero-shot OA of 0.565 (Kosmos 2 + GPT-4o).
    2. Coarse meta-class classification (D=0D=0) reaches an OA of 0.843, mitigating confusion among fine-grained classes.
    3. Across depths, performance degrades as task granularity increases (D=0→D=1D=0 \to D=1 drops by ≈0.20\approx 0.20, and D=1→D=2D=1 \to D=2 drops by another ≈0.18\approx 0.18--0.200.20).
  6. Knowl 6 — LLM-Derived Meta-Class Taxonomies for UC Merced and RESISC45

    data/table

    Zero-shot semantic clustering using GPT-4o and Llama-3.1 yielded the following meta-class cluster assignments for the benchmark datasets:

    UC Merced Taxonomies (D=0→D=1D=0 \to D=1):

    • GPT-4o Clusters:
      • Natural Landscapes: agricultural, beach, chaparral, forest, river
      • Recreational Areas: baseballdiamond, golfcourse, tenniscourt
      • Residential Areas: denseresidential, mediumresidential, mobilehomepark, sparseresidential
      • Transportation: airplane, freeway, harbor, intersection, overpass, parkinglot, runway
      • Urban Infrastructure: buildings, storagetanks
    • Llama-3.1 Clusters:
      • Manmade Structures: baseballdiamond, buildings, harbor, intersection, overpass, parkinglot, storagetanks
      • Natural Environments: agricultural, beach, chaparral, forest, river
      • Recreational Facilities: golfcourse, tenniscourt
      • Residential Areas: denseresidential, mediumresidential, mobilehomepark, sparseresidential
      • Transportation Infrastructure: airplane, freeway, runway

    RESISC45 Taxonomy via GPT-4o (D=0→D=1→D=2D=0 \to D=1 \to D=2):

    • Natural Landscapes:
      • Agricultural and Terrain: circular_farmland, rectangular_farmland
      • Aquatic and Water Bodies: lake, river, sea_ice, wetland
      • Natural Landscape and Vegetation: beach, chaparral, cloud, desert, forest, island, meadow, mountain, snowberg
    • Recreational Facilities:
      • Ball Sports: baseball_diamond, basketball_court, stadium, tennis_court
      • Golf: golf_course
      • Track and Field: ground_track_field
    • Transportation & Infrastructure:
      • Industrial Facilities: ship, storage_tank, thermal_power_station
      • Traffic Structures: bridge, intersection, overpass, roundabout
      • Transportation Infrastructure: airplane, airport, freeway, harbor, parking_lot, railway, railway_station, runway
    • Urban Structures:
      • Commercial & Industrial Zones: commercial_area, industrial_area
      • Public & Historic Buildings: church, palace
      • Residential Areas: dense_residential, medium_residential, mobile_home_park, sparse_residential, terrace
  7. Knowl 7 — Prompt Sensitivity and Class Enumeration Dilution in Vision-Language Models

    empirical result

    Empirical evaluation across binary and multi-class benchmarks reveals that the effect of enumerating candidate class names within the vision Large Language Model (vLLM) prompt depends strongly on label set cardinality:

    1. Small / Binary Taxonomies Benefit from Class Enumeration: For the binary Buildings vs. No Buildings task on SpaceNet v7, providing the candidate classes in the vLLM prompt sharpens visual focus and improves zero-shot classification accuracy. When combined with geographic context extracted from image filenames, Kosmos 2 + GPT-4o OA increases from 0.878 to 0.917, and Kosmos 2 + Llama-3.1 OA increases from 0.859 to 0.910.

    2. Large Taxonomies Suffer from Semantic Dilution: When classifying across 21 classes (UC Merced) or 45 classes (RESISC45), inserting the exhaustive list of class names into the vLLM prompt causes severe performance degradation, particularly for Kosmos 2:

      • On UC Merced (21 classes), forcing all class names into the prompt reduces Kosmos 2 + GPT-4o OA by 19.3 percentage points (0.710→0.5170.710 \to 0.517), Kosmos 2 + Phi-3 OA by 12.7 percentage points (0.648→0.5210.648 \to 0.521), and Kosmos 2 + Llama-3.1 OA by 10.5 percentage points (0.619→0.5140.619 \to 0.514).
      • On RESISC45 (45 classes), forcing all class names lowers Kosmos 2 + GPT-4o OA by 21.3 percentage points (0.565→0.3520.565 \to 0.352), Kosmos 2 + Phi-3 OA by 17.9 percentage points (0.523→0.3440.523 \to 0.344), and Kosmos 2 + Llama-3.1 OA by 20.4 percentage points (0.476→0.2720.476 \to 0.272).

    Consequently, open-ended visual description generation without explicit class enumeration produces significantly superior textual features for downstream LLM classification in large label spaces.

  8. Knowl 8 — Limitations of GeoVision Labeler

    limitation

    GeoVision Labeler (GVL) operates under several key limitations:

    1. Prompt Tuning Dependency: Long lists of target categories dilute vLLM descriptive accuracy, requiring dataset-specific prompt engineering to determine whether to include class names.
    2. Inference Latency in Hierarchical Pipelines: Multi-level recursive clustering and classification require multiple sequential calls to vLLMs and LLMs, which increases computational expense, inference latency, and system orchestration complexity.
    3. Performance Gap to Supervised Baselines: In fine-grained multi-class regimes, GVL's pure zero-shot performance remains below supervised and adapted models (e.g., 0.714 OA on UC Merced vs. 0.907 for fine-tuned ResNet50 and 0.959 for RS-CLIP; 0.565 OA on RESISC45 vs. 0.775 for ResNet50 and 0.858 for RS-CLIP).
    4. Domain Shift and Environmental Vulnerability: Off-the-shelf vision and language models are predominantly pretrained on general web imagery rather than satellite data; environmental factors such as cloud cover, seasonal foliage variations, and sensor artifacts degrade description fidelity.
    5. Spatial Patch Scale Constraints: Patch sizes below 224×224224 \times 224 pixels lack sufficient spatial context for accurate vLLM description, while large unpartitioned scenes dilute visual attention over fine features.
    6. RGB Format Restriction: GVL is currently restricted to 3-channel RGB imagery and cannot process multispectral, hyperspectral, or near-infrared bands standard in satellite remote sensing.

Coverage note — All major contributed methods, clustering algorithms, benchmark results across SpaceNet v7, UC Merced, and RESISC45, prompt engineering analyses, derived meta-class taxonomies, and stated limitations were included; qualitative confusion matrix figures were synthesized directly into the corresponding empirical result knowls.

References

  1. 1.Abdin, M., Aneja, J., Awadalla, H., Awadallah, A., Awan, A. A., Bach, N., Bahree, A., Bakhtiari, A., Bao, J., Behl, H., et al. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219, 2024.
  2. 2.Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  3. 3.Al Shafian, S. and Hu, D. Integrating machine learning and remote sensing in disaster management: A decadal review of post-disaster building damage assessment. Buildings, 14(8):2344, 2024.
  4. 4.Bilal, A., Ebert, D., and Lin, B. Llms for explainable ai: A comprehensive survey. arXiv preprint arXiv:2504.00125, 2025.
  5. 5.Corley, I., Robinson, C., Dodhia, R., Ferres, J. M. L., and Najafirad, P. Revisiting pre-trained remote sensing model benchmarks: resizing and normalization matters. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3162–3172, 2024.
  6. 6.Ge, Z., McCool, C., Sanderson, C., and Corke, P. Subset feature learning for fine-grained category classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pp. 46–52, 2015.
  7. 7.Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., et al. The Llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.
  8. 8.He, X. and Peng, Y. Fine-grained image classification via combining vision and language. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5994–6002, 2017.
  9. 9.Li, D., Wang, S., He, Q., and Yang, Y. Cost-effective land cover classification for remote sensing images. Journal of Cloud Computing, 11(1):62, 2022.
  10. 10.Li, H., Cui, Z., Zhu, Z., Chen, L., Zhu, J., Huang, H., and Tao, C. Rs-metanet: Deep meta metric learning for few-shot remote sensing scene classification. arXiv preprint arXiv:2009.13364, 2020.
  11. 11.Li, J., Cai, Y., Li, Q., Kou, M., and and, T. Z. A review of remote sensing image segmentation by deep learning methods. International Journal of Digital Earth, 17(1): 2328827, 2024a.
  12. 12.Li, X., Wen, C., Hu, Y., and Zhou, N. Rs-clip: Zero shot remote sensing scene classification via contrastive vision-language supervision. International Journal of Applied Earth Observation and Geoinformation, 124: 103497, 2023.
  13. 13.Li, X., Deng, R., Tang, Y., Bao, S., Yang, H., and Huo, Y. Leverage weekly annotation to pixel-wise annotation via zero-shot segment anything model for molecularempowered learning. In Medical Imaging 2024: Digital and Computational Pathology, volume 12933, pp. 133–139, 2024b.
  14. 14.Mehmood, M., Shahzad, A., Zafar, B., Shabbir, A., and Ali, N. Remote sensing image classification: A comprehensive review and applications. Mathematical Problems in Engineering, 2022(1):5880959, 2022.
  15. 15.Menon, S. and Vondrick, C. Visual classification via description from large language models. arXiv preprint arXiv:2210.07183, 2022.
  16. 16.Mirza, M. J., Karlinsky, L., Lin, W., Possegger, H., Kozinski, M., Feris, R., and Bischof, H. Lafter: Label-free tuning of zero-shot classifier using language and unlabeled image collections. Advances in Neural Information Processing Systems, 36:5765–5777, 2023.
  17. 17.Navin, M. S. and Agilandeeswari, L. Comprehensive review on land use/land cover change classification in remote sensing. Journal of Spectral Imaging, 9, 2020.
  18. 18.Neumann, M., Pinto, A. S., Zhai, X., and Houlsby, N. In-domain representation learning for remote sensing. arXiv preprint arXiv:1911.06721, 2019.
  19. 19.Nguyen, H., Clement, T., Nguyen, L., Kemmerzell, N., Truong, B., Nguyen, K., Abdelaal, M., and Cao, H. Langxai: Integrating large vision models for generating textual explanations to enhance explainability in visual perception tasks. In Larson, K. (ed.), Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24, pp. 8754–8758. International Joint Conferences on Artificial Intelligence Organization, 8 2024. Demo Track.
  20. 20.Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., and Wei, F. Kosmos-2: Grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824, 2023a.
  21. 21.Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., and Wei, F. Kosmos-2: Grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824, 2023b.
  22. 22.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pp. 8748–8763, 2021.
  23. 23.Rambabu, D., Gayathri, C., Datla, R., and Babu, S. Superclip: Semantic attribute-guided transformer with superresolution and clip for zero-shot remote sensing scene classification. IEEE Geoscience and Remote Sensing Letters, 22:1–5, 2025.
  24. 24.Snæbjarnarson, V., Du, K., Stoehr, N., Belongie, S., Cotterell, R., Lang, N., and Frank, S. Taxonomy-aware evaluation of vision-language models. arXiv preprint arXiv:2504.05457, 2025.
  25. 25.Song, J., Gao, S., Zhu, Y., and Ma, C. A survey of remote sensing image classification based on cnns. Big Earth Data, 3(3):232–254, 2019.
  26. 26.Sun, G., Cholakkal, H., Khan, S., Khan, F., and Shao, L. Fine-grained recognition: Accounting for subtle differences between similar classes. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pp. 12047–12054, 2020.
  27. 27.Sun, H., Zhen, Z., Zhao, P., and Zhang, Y. A review of zero-shot image classification. In Hassanien, A. E., Zheng, D., Zhao, Z., and Fan, Z. (eds.), Business Intelligence and Information Technology, 2024a.
  28. 28.Sun, X., Tian, Y., and Li, H. Zero-shot image classification via visual–semantic feature decoupling. Multimedia Systems, 30(2):82, 2024b.
  29. 29.Tao, C., Qi, J., Lu, W., Wang, H., and Li, H. Remote sensing image scene classification with self-supervised paradigm under limited labeled samples. IEEE Geoscience and Remote Sensing Letters, 19:1–5, 2020.
  30. 30.Van Etten, A. and Hogan, D. The spacenet multi-temporal urban development challenge. arXiv preprint arXiv:2102.11958, 2021.
  31. 31.Wang, Q.-W., Xie, Y., Zhang, L., Liu, Z., and Xia, S.-T. Pre-trained vision-language models as noisy partial annotators. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pp. 21189–21197, 2025.
  32. 32.Yang, Y. and Newsam, S. Bag-of-visual-words and spatial extensions for land-use classification. In Proceedings of the 18th SIGSPATIAL International Conference on Advances in Geographic Information Systems, pp. 270–279, 2010.
  33. 33.Ye, Z., Hu, F., Lyu, F., Li, L., and Huang, K. Disentangling semantic-to-visual confusion for zero-shot learning. IEEE Transactions on Multimedia, 24:2828–2840, 2021.
  34. 34.Yin, J., Dong, J., Hamm, N. A., Li, Z., Wang, J., Xing, H., and Fu, P. Integrating remote sensing and geospatial big data for urban land use mapping: A review. International Journal of Applied Earth Observation and Geoinformation, 103:102514, 2021.
  35. 35.Zhang, Y., Unell, A., Wang, X., Ghosh, D., Su, Y., Schmidt, L., and Yeung-Levy, S. Why are visually-grounded language models bad at image classification? arXiv preprint arXiv:2405.18415, 2024.

Citation

MLA
Hacheme, G. Q., et al. “GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models”. arXiv, 2025, http://arxiv.org/abs/2505.24340v1.
APA
Hacheme, G. Q., Tadesse, G. A., Robinson, C., Zaytar, A., Dodhia, R., & Ferres, J. M. L. (2025). GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models. arXiv. http://arxiv.org/abs/2505.24340v1
Chicago
Hacheme, G. Q., G. A. Tadesse, C. Robinson, A. Zaytar, R. Dodhia, and J. M. L. Ferres. 2025. “GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models”. arXiv. http://arxiv.org/abs/2505.24340v1.
Harvard
Hacheme, G.Q. et al. (2025) “GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2505.24340v1.
Vancouver
1. Hacheme GQ, Tadesse GA, Robinson C, Zaytar A, Dodhia R, Ferres JML (2025) GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models. arXiv

BibTeX

@article{hacheme2025geovision,
  title = {GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models},
  author = {Hacheme, Gilles Quentin and Tadesse, Girmaw Abebe and Robinson, Caleb and Zaytar, Akram and Dodhia, Rahul and Ferres, Juan M. Lavista},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2505.24340v1},
  eprint = {2505.24340}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/