How Can AI Find My Model? A Model-Finding Experimental Study Considering Data Formats, Embeddings, and Retrieval Strategies
Quick Answer
This study explores how data representation, transformer-based embeddings, and retrieval strategies impact the discovery of simulation models through natural language queries.
Quick Take
Results indicate that open-source embedding models perform well, and reranking methods are crucial as query complexity increases, providing a baseline for AI-driven model discovery.
Key Points
- Data representation significantly affects model discovery performance.
- Open-source embedding models achieve high performance in retrieval tasks.
- Reranking methods are essential for complex queries.
- The study uses recall@5 and nDCG@5 as evaluation metrics.
- Findings contribute to AI-driven composability and interoperability.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Discovering simulation models for reuse remains a fundamental challenge in Modeling and Simulation (M&S). When many models coexist, identifying those that align with a given modeling intent remains difficult. Recent advances in Artificial Intelligence (AI), particularly retrieval-based approaches, offer a promising pathway to operate at this semantic layer. In this paper, we present an experimental study investigating the impact of data representation, transformer-based embedding models, and retrieval strategies on the discovery of simulation models using natural language queries. We evaluated performance across multiple query types using standard information retrieval metrics, including recall@5 and nDCG@5. Results show that data representation matters, open-source embedding models can achieve high performance, and reranking methods are important, especially as query complexity increases. This work provides a baseline for AI-driven model discovery and discusses its role in advancing toward AI-driven composability and interoperability.
| Comments: | Accepted for publication in Proceedings of the 2026 Winter Simulation Conference (WSC 2026). The final published version will appear in IEEE Xplore |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.30846 [cs.AI] |
| (or arXiv:2606.30846v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2606.30846 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jhon G. Botello [view email]
[v1]
Mon, 29 Jun 2026 19:23:32 UTC (1,078 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.