DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings
Quick Answer
DrawingVQA introduces a benchmark for evaluating multimodal large language models (MLLMs) on construction drawings, revealing significant performance gaps, especially in complex reasoning tasks.
Quick Take
It includes 33 construction drawings and 92 curated Q&A pairs across three reasoning depths, highlighting the need for AI integration in engineering workflows.
Key Points
- First benchmark for MLLMs on real-world construction drawings.
- Includes 33 'Issued for Construction' drawings and 92 Q&A pairs.
- Evaluates three reasoning depths: perceptual, contextual, and expert.
- Significant performance gap between MLLMs and human experts.
- Lays foundation for AI-driven understanding in engineering workflows.
DeepSignal Analysis
What happened
DrawingVQA is a newly introduced benchmark aimed at assessing multimodal large language models (MLLMs) specifically on construction drawings. It includes 33 construction drawings and 92 curated question-answer pairs that span three levels of reasoning. Evaluations indicate a significant performance gap between MLLMs and expert performance, particularly in complex reasoning tasks.
Key evidence
- DrawingVQA consists of 33 'Issued for Construction' drawings and 92 curated Q&A pairs across three reasoning depths: perceptual understanding, contextual interpretation, and domain-expert reasoning.
- The benchmark reveals a substantial gap between the performance of state-of-the-art MLLMs and expert performance, especially at higher reasoning depths.
- DrawingVQA is the first benchmark to explicitly map engineering workflows to AI reasoning competencies, highlighting the integration of AI in engineering.
Why it matters
The introduction of DrawingVQA addresses a critical need for evaluating AI capabilities in understanding complex construction drawings, which are essential in various engineering fields. The significant performance gaps identified suggest that current MLLMs may not yet be suitable for practical applications in engineering workflows. This benchmark could drive future research and development towards improving AI's reasoning abilities in specialized domains.
Paper Resources
📖 Reader Mode
~2 min readAbstract:We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings -- a core media in architecture, civil, and many other engineering practices. Unlike natural images or schematic floor plans, construction drawings fuse abstract geometry, symbolic notation, tabular data, annotations, and domain-specific text, forming a uniquely complex visual-textual domain core to engineering workflows. DrawingVQA bridges this gap with 33 "Issued for Construction" drawings and 92 expertly curated question-answer pairs, spanning three reasoning depths: perceptual understanding, contextual interpretation, and domain-expert reasoning. To evaluate model capabilities, we present a dual categorization framework to jointly analyze performance across seven construction-engineering and four MLLM capability dimensions -- the first to explicitly map engineering workflows to AI reasoning competencies. Evaluations of state-of-the-art MLLMs reveal a substantial gap between model and expert performance, particularly at higher reasoning depths. This benchmark lays a foundation for domain-specialized multimodal reasoning to allow for advancement on integration of AI-driven understanding and real-world engineering workflows.
| Comments: | CVPR 2026 Findings accepted paper |
| Subjects: | Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.15418 [cs.AI] |
| (or arXiv:2607.15418v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.15418 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yoonhwa Jung [view email]
[v1]
Thu, 16 Jul 2026 19:44:03 UTC (9,031 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.