OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining
Quick Answer
OPTScientist introduces a multi-agent framework for discovering optimizers in a typed DSL, overcoming limitations of existing methods.
Quick Take
It successfully identifies RS-MR, a reduced-state matrix optimizer that enhances transformer pretraining performance beyond strong baselines, paving the way for automated optimizer science.
Key Points
- OPTScientist uses a approach with roles like Theorist and Designer.
- It formulates optimizer design as a constrained scientific search process.
- The framework discovered RS-MR, improving transformer pretraining performance.
- Combines evolutionary search with DSL extensions to address representational bottlenecks.
- Results indicate a new direction for automated optimizer design grounded in theory.
DeepSignal Analysis
What happened
OPTScientist is a multi-agent framework designed to discover optimizers for deep learning, specifically for transformer pretraining. It addresses the limitations of existing methods by using a typed domain-specific language (DSL) and a collaborative approach involving four distinct roles: Theorist, Designer, Engineer, and Reviewer. The framework successfully identified RS-MR, an optimizer that outperforms established baselines.
Key evidence
- OPTScientist employs a multi-agent framework that includes Theorist, Designer, Engineer, and Reviewer roles to facilitate optimizer discovery.
- The framework utilizes a typed domain-specific language (DSL) to express candidate updates, addressing issues found in both unconstrained code spaces and narrowly parameterized families.
- RS-MR, the discovered optimizer, improves transformer pretraining performance beyond strong baselines according to the native evaluation protocol used in the study.
Why it matters
The introduction of OPTScientist represents a significant advancement in the field of automated optimizer discovery, which has traditionally faced challenges in balancing flexibility and stability. By combining evolutionary search with a theory-guided approach, this framework could lead to more effective and interpretable optimizers. The successful identification of RS-MR suggests potential for further innovations in optimizer design, which could enhance the efficiency of deep learning models.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Designing optimizers for modern deep learning remains a challenging scientific problem, requiring the joint consideration of optimization geometry, state dynamics, numerical stability, implementation constraints, and empirical generalization. Existing automated optimizer discovery methods typically search either over unconstrained code spaces or within narrowly parameterized optimizer families. The former is flexible but often produces invalid or uninterpretable programs, while the latter is stable but limits novelty. We introduce OPTScientist, a theory-guided multi-agent framework for optimizer discovery in a typed domain-specific language (DSL). OPTScientist formulates optimizer design as a constrained scientific search process, where candidate updates are expressed through direction, scaling, preconditioning, regularization, state, and grouping modules. Four role agents, Theorist, Designer, Engineer, and Reviewer, collaborate within a single orchestration loop to propose hypotheses, synthesize DSL candidates, compile and evaluate optimizers, and critique results. To overcome the limitations of a fixed search space, OPTScientist combines evolutionary search over optimizer programs with a second-stage mechanism that proposes small DSL extensions when repeated failures reveal representational bottlenecks. Using this framework, we discover RS-MR, a reduced-state matrix optimizer that improves transformer pretraining over strong baselines under our native evaluation protocol. Our results suggest a path toward automated optimizer science grounded in theory, typed programs, compiler validation, and closed-loop experimentation.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.20486 [cs.AI] |
| (or arXiv:2607.20486v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.20486 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zhongzheng Li [view email]
[v1]
Tue, 2 Jun 2026 15:21:53 UTC (2,854 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.