EO-Agents: A Three-Agent LLM Pipeline for Earth Observation Hypothesis Generation
Quick Answer
This paper shows that The EO-Agents pipeline utilizes a three-agent LLM system to generate scientifically grounded hypotheses from NASA's Earth Observation Knowledge Graph, producing 160 hypotheses across various Earth science domains.
Quick Take
A factorial experiment reveals stable hypothesis rankings across models GPT-5.2 and Claude Sonnet 4.6, while highlighting the variability in absolute scores based on judge identity.
Key Points
- The pipeline ranks dataset pairings using a heterogeneous graph neural network.
- 160 hypotheses generated span ecohydrology, glaciology, and more.
- Model-predicted dataset pairings are nearly as plausible as real co-usages.
- Hypothesis rankings remain stable across different .
- Single-judge evaluations reveal limitations in assessing hypothesis quality.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language models have recently been explored for scientific hypothesis generation, but most prior work relies on unstructured literature and free-form textual claims. We present a pipeline for Earth observation that grounds hypothesis generation directly in the NASA Earth Observation Knowledge Graph. A heterogeneous graph neural network trained on historical co-usage relations ranks candidate dataset pairings, and a three-agent LLM pipeline filters, generates, and evaluates structured research hypotheses. Applied to 1,475 NASA datasets, the system produces 160 hypotheses spanning multiple Earth-science domains, including ecohydrology, glaciology, aerosol--cloud interactions, vegetation phenology, and stratospheric chemistry. Model-predicted novel dataset pairings are rated nearly as plausible as held-out real co-usages from the literature, indicating that the pipeline surfaces scientifically coherent yet unexplored combinations. A 2*2*2 factorial experiment across GPT-5.2 and Claude Sonnet 4.6 shows that hypothesis rankings remain stable, while absolute scores depend strongly on judge identity, highlighting limitations of single-judge LLM evaluation.
| Comments: | Accepted at the ICML 2026 AI for Science Workshop |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.01584 [cs.AI] |
| (or arXiv:2607.01584v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.01584 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mahyar Ghazanfari [view email]
[v1]
Thu, 2 Jul 2026 01:31:18 UTC (757 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.