Search-on-Graph-R1: Training Large Language Models to Search Knowledge Graphs with Reinforcement Learning
Quick Answer
This paper shows that Search-on-Graph-R1 (SOGR1) is an 8B model that outperforms existing frontier LLMs in knowledge graph question answering (KGQA) by integrating supervised fine-tuning and reinforcement learning.
Quick Take
It achieves superior results on benchmarks like CWQ without auxiliary modules during inference, demonstrating effective navigation through knowledge graphs with fewer search calls.
Key Points
- SOGR1 integrates supervised fine-tuning and reinforcement learning for enhanced KGQA.
- Achieves best results on CWQ compared to other frozen frontier-.
- Utilizes a live Freebase server for grounded knowledge graph navigation.
- Demonstrates fewer search calls needed to reach answers than its SFT initialization.
- No auxiliary modules or LLM judges are used during inference or training.
DeepSignal Analysis
What happened
Search-on-Graph-R1 (SOGR1) is an 8 billion parameter model that integrates supervised fine-tuning and reinforcement learning to improve knowledge graph question answering (KGQA). It outperforms existing large language models (LLMs) on benchmarks like CWQ without requiring auxiliary modules during inference.
Key evidence
- SOGR1 achieves superior results on benchmarks such as WebQSP, CWQ, and GrailQA, surpassing every frozen frontier-LLM system in comparison.
- The model internalizes navigation through knowledge graphs, using a teacher model that follows a known answer path with a live Search tool.
- SOGR1 learns to reach answers in fewer search calls compared to its supervised fine-tuning initialization, indicating effective reinforcement learning.
Why it matters
The development of SOGR1 highlights a shift towards more efficient models in KGQA, reducing the reliance on costly frontier-scale inference. This could lead to broader adoption of knowledge graph technologies in applications where quick and accurate information retrieval is essential. The model's ability to outperform larger models with fewer resources may influence future research and development in the field.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Knowledge graph question answering (KGQA) requires navigating from topic entities to an answer several relations away. Recent methods prompt a frontier LLM to explore the graph through a retrieval tool, but their reliance on frontier-scale inference makes them costly to deploy. We present Search-on-Graph-R1 (\sogrone{}), which internalizes this navigation into a compact 8B model through supervised fine-tuning (SFT) followed by reinforcement learning (RL). Our central idea is to scaffold a frontier teacher with each question's gold SPARQL query, so the teacher traverses a known answer-bearing path with a live \texttt{Search} tool rather than having to discover the path itself. Since every call executes against a live Freebase server, the resulting trajectories are grounded in the knowledge graph by construction. On WebQSP, CWQ, and GrailQA, \sogrone{} at 8B surpasses every frozen frontier-LLM system in our comparison and posts the strongest results on CWQ of any system we compare against. It does so using no auxiliary module at inference and no LLM judge during training. Isolating each training stage shows that SFT and RL contribute complementary gains, our approach transfers across model families, and RL learns to reach answers in fewer \texttt{Search} calls than its SFT initialization.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.18481 [cs.CL] |
| (or arXiv:2607.18481v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.18481 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jia Ao Sun [view email]
[v1]
Mon, 20 Jul 2026 19:58:32 UTC (190 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.