Scenario Generation for Testing of Autonomous Driving Systems Using Real-World Failure Records
Quick Answer
This study introduces a novel scenario generation pipeline for Autonomous Driving Systems (ADS) testing, leveraging historical failure records in natural language.
Quick Take
By utilizing modular -based synthetic scenario generation, the method produces diverse scenarios compatible with testing constraints, successfully applying it to generate 20 scenarios for the Metadrive simulator using NHTSA ADS crash data.
Key Points
- Proposes a scenario generation pipeline using historical ADS failure records.
- Utilizes modular LLM-based generation for diverse scenario creation.
- Successfully generates 20 scenarios for Metadrive simulator testing.
- Combines 4 road types and 3 non-ego vehicle movements in scenarios.
- Code available on GitHub for further use and exploration.
Paper Resources
📖 Reader Mode
~2 min readAbstract:To ensure safe on-road behavior, pre-deployment testing and failure discovery of Autonomous Driving Systems (ADS) is crucial. Present day simulation based testing methods focus largely on mathematical models for efficient search of optimal scenarios, assuming a fixed scenario representation. On the other hand, real-world testing involves substantial manual effort to design scenario templates for testing. These templates represent distinct failure scenarios consisting of pre-deployment vehicle movements, map types, etc. Historical failure records for ADS are a reliable source of real-world failure conditions, which can be used for scenario generation. In this work, we propose a scenario generation pipeline using categorical and contextual information available from historical records in natural language format. Our approach consists of modular LLM based synthetic scenario generation, compatible with the testing constraints of a given system. We successfully apply our method to generate a diverse set of scenarios for testing autonomous navigation on Metadrive simulator using the NHTSA ADS crash records. Our approach results in accurate and diverse scenario generation with a combination of 4 road types, 3 non ego vehicle movement types, including on road anomalies in the form of working zones. Generated scenarios align with the provided testing conditions, and reveals interesting failures of the system within a limited testing budget of 20 scenarios. Code is available at this https URL.
| Comments: | 9 pages, Appendix included. Paper accepted and presented at NeuS 2026 |
| Subjects: | Artificial Intelligence (cs.AI); Robotics (cs.RO) |
| Cite as: | arXiv:2606.31131 [cs.AI] |
| (or arXiv:2606.31131v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2606.31131 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Anjali Parashar [view email]
[v1]
Tue, 30 Jun 2026 04:56:25 UTC (3,201 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.