OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation
Quick Answer
OVEarth-Bench introduces a comprehensive evaluation framework for open-vocabulary Earth observation, emphasizing broader category coverage and diverse query types.
Quick Take
Current methods show limited performance, with MLLM-based approaches leading in effectiveness, while EO-specific models often underperform compared to general models. The findings stress the need for more realistic benchmarks to enhance future method development.
Key Points
- OVEarth-Bench enhances evaluation with broad hierarchical category coverage and diverse query types.
- Current methods show limited performance; broader categories lead to more stable rankings.
- MLLM-based methods achieve the best overall performance in the evaluation.
- EO-specific methods generally underperform compared to general models.
- The benchmark aims to guide future open-vocabulary EO method design.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchmarks, however, usually cover narrow category vocabularies or limited query forms. To fill this gap, we introduce OVEarth-Bench, which extends existing evaluation in two directions: category breadth, through broad hierarchical category coverage with positive and negative expressions, and query diversity, through vocabulary, referring, and reasoning queries. The benchmark supports mask and box localization under a unified zero-shot protocol. We evaluate a broad set of general and EO-specific methods. The evaluation reveals that: (1) the performance of current methods remains limited, while broader category coverage yields more stable model rankings; (2) MLLM-based methods achieve the strongest overall performance; and (3) EO-specific methods generally underperform general models and rarely match the strongest methods. These findings provide guidance for future open-vocabulary EO method design and highlight the importance of developing more realistic, diverse, high-quality, and large-scale benchmarks for reliable evaluation. Our data and evaluation package are released at this https URL.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.27278 [cs.CV] |
| (or arXiv:2607.27278v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.27278 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Kaiyu Li [view email]
[v1]
Wed, 29 Jul 2026 13:19:12 UTC (7,043 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.