EPOCH: Reliable Discovery through Evidence-Governed Search
Quick Answer
EPOCH is an evidence-governed architecture that enhances AI research agents' discovery capabilities, achieving a mean normalized score of 0.65 on AlgoTune, surpassing the previous baseline of 0.53.
Quick Take
It demonstrates significant advancements across ten discovery problems, yielding improved algorithms and proof-supported results, thereby promoting more reliable scientific discoveries.
Key Points
- EPOCH implements an evidence-governed discovery loop with explicit task contracts and active falsification.
- Achieves state-of-the-art performance on AlgoTune with a mean normalized score of 0.65.
- Delivers substantial task-specific advancements in algorithms and proof-supported results.
- Demonstrates favorable behavior under official-test replay, leading the descriptive aggregate on AgentHPO.
- Highlights the necessity of evidence governance for trustworthy AI research outcomes.
Paper Resources
📖 Reader Mode
~2 min readAbstract:AI research agents are increasingly used to search over programs, mathematical constructions, and proofs. However, existing systems typically optimize evaluator feedback without adequately governing how that feedback is interpreted, challenged, and reused. As a result, promising but fragile candidates can be promoted as discoveries, while benchmark improvements, finite certificates, and theorem-level claims are too easily conflated. We introduce EPOCH, an evidence-governed architecture designed to close this gap. EPOCH implements an evidence-governed discovery loop by combining explicit task contracts, typed memory, active falsification, admission checks, and independent replay, so that each candidate is evaluated against the strength and scope of the claim it supports. EPOCH achieves state-of-the-art aggregate performance on AlgoTune, substantially exceeding the strongest baseline in mean normalized score (0.65 vs. 0.53), and attains the highest mean score on the internal Math14 suite (0.57). It further shows favorable held-out behavior under official-test replay and leads the descriptive aggregate on AgentHPO. Across ten discovery problems, EPOCH delivers substantial task-specific advances, including improved executable constructions, optimized algorithms, counterexamples, and proof-supported results. These advances demonstrate its ability to convert search into concrete progress across mathematical and computational domains. Together, the results suggest that evidence governance is a necessary step toward AI research agents that produce not only stronger solutions, but also more trustworthy scientific discoveries.
| Comments: | 49 pages, 16 figures, including supplementary material |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.06986 [cs.AI] |
| (or arXiv:2610.06986v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.06986 arXiv-issued DOI via DataCite |
Submission history
From: Binjie Guo [view email]
[v1]
Sun, 4 Oct 2026 06:51:12 UTC (541 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.