Agentic evolution of physically constrained foundation models
Quick Answer
This paper shows that A new multi-agent discovery engine autonomously designs hardware-compliant systems, evolving methods like Q-Enhance and MoE-Salient-AQ that outperform human heuristics.
Quick Take
It successfully deployed a 235-billion-parameter model on a dual-A100 server, reducing memory needs by 75% with only a 0.64% accuracy drop.
Key Points
- Engine evolved two hardware-aware compression methods: Q-Enhance and MoE-Salient-AQ.
- Q-Enhance reduces long-context accuracy loss in dense models effectively.
- MoE-Salient-AQ outperforms manual sparse designs by 3.7% in sub-3-bit regimes.
- Successfully deployed a 235-billion-parameter model on a dual-A100 server.
- Achieved 75% memory reduction with a minimal accuracy degradation of 0.64%.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Artificial intelligence increasingly drives automated scientific discovery, yet contemporary generalist agents lack physical grounding, frequently hallucinating hardware-incompatible designs. Here, we present a physically grounded, multi-agent discovery engine that autonomously architects hardware-compliant computing systems. Anchored by an Evolutionary Knowledge Graph structuring past scientific innovations, the framework extracts an "algorithmic Chain-of-Thought" to transform blind stochastic search into directed structural evolution. Applied to the extreme testbed of foundation model deployment, the engine evolved two hardware-aware compression methodologies surpassing human-engineered heuristics: Q-Enhance mitigates long-context accuracy loss in dense models, and MoE-Salient-AQ outperforms state-of-the-art manual sparse Mixture-of-Experts designs by 3.7% at sub-3-bit regimes. Utilizing a bandwidth-efficient Sensitivity Profile, we successfully deployed a massive 235-billion-parameter model onto a constrained dual-A100 server, reducing memory requirements by 75% with a marginal 0.64% accuracy degradation. By transforming unconstrained combinatorial search into knowledge-driven autonomy, this establishes a scalable hardware-software co-design paradigm for machine-driven discovery within strict physical boundaries.
| Comments: | 29 pages, 5 main figures and 4 extended data figures |
| Subjects: | Artificial Intelligence (cs.AI); Hardware Architecture (cs.AR); Machine Learning (cs.LG); Multiagent Systems (cs.MA) |
| Cite as: | arXiv:2606.25532 [cs.AI] |
| (or arXiv:2606.25532v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2606.25532 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jiangwei Zhang [view email]
[v1]
Wed, 24 Jun 2026 08:07:59 UTC (3,228 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.