
Science One Framework: A verifiable autonomous research framework via Chain-of-Evidence
Quick Answer
The Science One Framework introduces a verifiable autonomous research system using Chain-of-Evidence, achieving zero phantom references and fully verifiable scores, outperforming existing models like Sakana's AI-Scientist on benchmarks such as MLE-Bench and Parameter-Golf.
Key Points
- Chain-of-Evidence ensures each research claim is backed by verifiable evidence.
- Science One Framework eliminates hallucinated references by grounding citations through the Semantic Scholar API.
- The CoE Audit rigorously checks generated papers for integrity and alignment with code.
- Science One achieved state-of-the-art performance on five optimization tasks in the ADRS benchmark.
- Baseline systems hallucinate up to 21% of their references, highlighting the need for improved verifiability.
DeepSignal Analysis
What happened
Google Research introduced the Science One Framework, which utilizes a Chain-of-Evidence to enhance verifiability in AI-driven research. This framework reportedly achieves zero phantom references and fully verifiable scores, outperforming existing models like Sakana's AI-Scientist on benchmarks such as MLE-Bench and Parameter-Golf.
Key evidence
- The Science One Framework eliminates reliance on model memory by building a citation graph via the Semantic Scholar API, ensuring all references are real.
- In tests, the Science One Framework achieved perfect score verification and the highest method-code alignment, while baseline systems had hallucination rates up to 21%.
- The framework matched or exceeded human expert performance on five tasks from the Automated Design of Research Systems benchmark, achieving the best overall score on two tasks.
Why it matters
As AI systems increasingly engage in scientific research, ensuring the trustworthiness of their outputs is critical. The Science One Framework's focus on verifiability addresses a significant challenge in the field, potentially setting a new standard for future AI research systems. By integrating evidence chains at the claim production stage, it aims to enhance the reliability of AI-generated research.
Paper Resources
Source Excerpt
(LLMs) are increasingly being deployed not just as coding assistants but as autonomous agents capable of conducting end-to-end scientific research workflows. Recent systems (e. g. , Sakana’s AI-Scientist, AutoResearchClaw, DeepScientist, AI-Researcher) can review literature, formulate hypotheses, execute experiments and write complete manuscripts that are comparable to human-authored papers. However, as the surface-level quality of these AI-generated manuscripts improves, a c
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Google Research
See more →
Towards a quantum computer that learns from its errors
Google Research introduces a reinforcement learning framework for quantum error correction, enhancing logical stability by 3.5 times on the Willow superconducting processor. This approach allows continuous calibration during computation, addressing the limitations of traditional quantum control methods.