Skill-Contracted Agents for Evidence-Aware Materials Literature Analysis
Quick Answer
AlphaAgent is a skill-driven framework for materials literature analysis that separates retrieval from report generation, significantly outperforming baseline systems in a blind evaluation on 40 questions, particularly in mechanistic explanation and credibility awareness.
Key Points
- AlphaAgent utilizes explicit skill contracts for improved task separation.
- It queries over 300,000 papers from the Journal Citation Reports.
- The framework achieved substantial gains in analytical reasoning tasks.
- Mechanistic explanation and credibility awareness were notably enhanced.
- Results indicate improved literature analysis for materials research.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Materials science literature analysis requires simultaneous attention to composition, processing, characterization, and property relationships, yet conventional retrieval-augmented generation pipelines struggle to reconcile heterogeneous tasks within a single retrieve-then-generate architecture. Here we present AlphaAgent, a skill-driven agent framework that decouples retrieval-based question answering from paper-level report generation through explicit skill contracts. A dedicated retrieval skill rewrites user requests into material-specific search intents, queries a curated index of more than 300,000 papers from the Journal Citation Reports Metallurgy and Metallurgical Engineering category, and reformulates queries when initial evidence is insufficient. A separate report-generation skill parses full-text PDFs to produce structured per-paper analytical reports and cross-paper summaries. In a blind evaluation on 40 materials-science questions, half of which required deep analytical reasoning, AlphaAgent substantially outperformed a baseline system matched for underlying model, document index, and retrieval scale, with the largest gains in mechanistic explanation and awareness of credibility boundaries. These results indicate that explicit task separation, refined retrieval intent, and evidence-aware generation improve large-language-model-based literature analysis for materials research.
| Comments: | 9 pages, 5 figures |
| Subjects: | Computation and Language (cs.CL); Information Retrieval (cs.IR) |
| Cite as: | arXiv:2607.20431 [cs.CL] |
| (or arXiv:2607.20431v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.20431 arXiv-issued DOI via DataCite |
Submission history
From: Peng Kang PhD [view email]
[v1]
Sun, 10 May 2026 13:00:32 UTC (4,444 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.