Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation
Quick Answer
This paper introduces a black-box evaluation framework for assessing LLMs in generating Design Structure Matrices (DSMs) from technical documentation.
Quick Take
The framework benchmarks generated DSMs against validated matrices, revealing ' potential and limitations in producing structurally plausible DSMs while highlighting issues with ambiguity and prompt formulation. Results indicate high reproducibility under structured inputs, establishing a transparent benchmark for Auto-DSM pipelines.
Key Points
- Introduces a reproducible methodology for benchmarking generated DSMs against ground-truth matrices.
- Evaluates DSMs using structural, classification, and stability metrics, culminating in a Composite Quality Score.
- Controlled experiments reveal LLMs' sensitivity to ambiguity and inconsistent dependency definitions.
- Demonstrates both the potential and limitations of LLM-driven DSM automation in engineering workflows.
- Provides a transparent benchmark for auditing Auto-DSM pipelines.
Paper Resources
📖 Reader Mode
~2 min readAbstract:This paper presents a black-box evaluation framework to systematically assess the ability of Large Language Models (LLMs) to generate Design Structure Matrices (DSMs) from structured technical documentation. Motivated by the closed-source nature of current Auto-DSM pipelines, the framework introduces a reproducible methodology that benchmarks generated DSMs (GEN-DSMs) against manually validated ground-truth matrices (GT-DSMs). The evaluation integrates both single-run and multi-run perspectives, combining structural metrics (Completeness, Correctness, Coupling Density), classification metrics (Selective Accuracy, Abstention Coverage), and stability measures (Entropy, Fleiss' $\kappa$). To synthesize these aspects, a Composite Quality Score (Q) is proposed. Controlled experiments are conducted on two datasets: a fictive abstract system and a real-world refrigerator decomposition, covering variations in phrasing, parameter-dataset alignment, and system complexity. Results show that LLMs can produce structurally plausible DSMs and achieve high reproducibility under well-structured inputs, but remain sensitive to ambiguity, inconsistent dependency definitions, and prompt formulation. The findings highlight systematic sources of hallucination and abstention failure, demonstrating both the potential and current limitations of LLM-driven DSM automation. The proposed framework provides a transparent benchmark for auditing Auto-DSM pipelines and establishes foundations for integrating LLM-based decomposition methods into model-based systems engineering (MBSE) workflows.
| Subjects: | Artificial Intelligence (cs.AI); Hardware Architecture (cs.AR); Computational Engineering, Finance, and Science (cs.CE); Systems and Control (eess.SY) |
| Cite as: | arXiv:2607.05985 [cs.AI] |
| (or arXiv:2607.05985v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.05985 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Theo Hofman [view email]
[v1]
Tue, 7 Jul 2026 08:17:49 UTC (1,631 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.