Harness-G: A Graph-Structured Harness for Search Agents
Quick Answer
Harness-G introduces a graph-structured retrieval framework that enhances RL search agents' query generation, achieving a 10.74 point F1 improvement over Graph-R1 at 1.5B parameters.
Quick Take
This method reduces retrieval aliasing and implements Structured Non-myopic Credit (SNC) for better action evaluation, leading to superior performance across six QA benchmarks.
Key Points
- Harness-G reformulates query generation as finite action selection, improving retrieval accuracy.
- The framework reduces linguistic aliasing, allowing for better comparison of retrieval alternatives.
- Structured Non-myopic Credit (SNC) assigns gains to earlier actions based on selected alternatives.
- Harness-G outperforms Graph-R1 by 10.74 points at 1.5B parameters and 3.98 points at 3B.
- Achieves highest average F1 across six QA benchmarks, demonstrating significant performance gains.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Reinforcement learning (RL) search agents commonly model retrieval as free-form natural-language query generation and optimize multi-turn interactions using final-answer rewards. Current studies mainly improve training with denser or more structured credit signals, but rarely examine whether retrieval is properly formulated at the policy-environment interface. We observe pronounced retrieval aliasing during Search-R1 training: rollouts for the same question continue to generate distinct query strings, yet their accumulated evidence sets increasingly overlap. We call this phenomenon retrieval-equivalence collapse; in this regime, trajectories approach utility equivalence with respect to retrieval decisions, leaving within-group returns with little effective retrieval contrast. To address this problem, we propose Harness-G, a graph-structured retrieval framework that redesigns this interface. It reformulates free-form query generation as finite action selection: the policy selects an evidence sentence or entity, or chooses to answer, while the environment constructs the menu, tracks retrieval state, and validates and executes each choice. This interface reduces linguistic aliasing and makes same-state alternatives directly comparable. Building on this interface, we introduce Structured Non-myopic Credit (SNC), which uses a frozen answer scorer to compare the selected action with its alternatives and assigns downstream gains to the earlier actions that enabled them. Across six QA benchmarks, Harness-G achieves the highest average F1 at both evaluated model scales, outperforming the strongest baseline, Graph-R1, by 10.74 points at 1.5B and 3.98 points at 3B.
| Comments: | Code:this https URL |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.27652 [cs.CL] |
| (or arXiv:2607.27652v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.27652 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yanning Hou [view email]
[v1]
Thu, 30 Jul 2026 04:05:05 UTC (2,930 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.