Synthesis Through Simulation: Generating Coherent Enterprise Data via Scalable Agent-System Interaction
Quick Answer
This paper shows that The Synthesis Through Simulation (STS) paradigm enables schema-free data synthesis using LLM agents to generate enterprise data through policy-enforcing APIs, achieving 0.88 average marginal fidelity and 100% constraint satisfaction across ten environments without requiring database schemas.
Key Points
- STS guarantees structural validity by generating data within defined enterprise environments.
- The Generalist Populator (GP) achieves high fidelity and scalability without database schema access.
- Statistical synthesizers fail in seven environments due to seed data limitations.
- Schema-privileged agents struggle with 82% of tasks in tightly coupled workflows.
DeepSignal Analysis
What happened
The Synthesis Through Simulation (STS) paradigm allows for schema-free data synthesis using LLM agents. This approach generates enterprise data through policy-enforcing APIs, achieving an average marginal fidelity of 0.88 and 100% constraint satisfaction across ten environments without needing database schemas.
Key evidence
- STS enables data generation by executing operations against policy-enforcing APIs in simulated environments, ensuring structural validity.
- The Generalist Populator (GP) achieves an average marginal fidelity of 0.88 and 100% constraint satisfaction across ten environments without access to database schemas.
- Statistical synthesizers are inapplicable to seven environments due to seed data requirements, while schema-privileged agents fail 82% of trajectories in tightly coupled workflows.
Why it matters
This development addresses significant challenges in enterprise AI, particularly the limitations imposed by business and legal constraints on data access. By decoupling validity enforcement from distribution modeling, STS offers a scalable solution for generating valid enterprise data, which could enhance operational efficiency and data-driven decision-making in various industries.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Tool-calling agents have become central to enterprise AI, yet training and evaluating them at scale remains severely constrained due to business and legal restrictions on enterprise systems, data, and database schemas. Tabular data synthesis offers a natural alternative, but its effectiveness is fundamentally limited by structural validity and schema availability, while procedure-based approaches yield the opposite weakness, typically lacking distributional fidelity without per-domain authoring. We introduce **Synthesis Through Simulation** (STS), a **schema--free** data synthesis paradigm in which an LLM agent generates data by executing operations against policy-enforcing APIs within simulated enterprise environments. Because data is generated through the same environment that defines what is valid, STS guarantees structural validity by construction while decoupling validity enforcement from distribution modeling, allowing each to be addressed independently. The **Generalist Populator** (GP), STS's domain-agnostic agent, addresses the remaining challenges of distributional fidelity and synthesis scalability: GP achieves **0.88** average marginal fidelity and **100\% constraint satisfaction** across all ten environments *without access to DB schemas*, while statistical synthesizers are inapplicable to seven due to necessary seed data requirements, and schema-privileged agents fail 82\% of trajectories on airline environment's tightly coupled workflows due to brittle task composition. We open-source the full framework, all ten environments, and generated datasets at this https URL.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.10549 [cs.AI] |
| (or arXiv:2610.10549v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10549 arXiv-issued DOI via DataCite |
Submission history
From: Ashutosh Hathidara [view email]
[v1]
Thu, 24 Sep 2026 13:55:22 UTC (419 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.