DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings
Quick Answer
DrawingVQA introduces a benchmark for evaluating multimodal large language models (MLLMs) on construction drawings, revealing significant performance gaps, especially in complex reasoning tasks.
Quick Take
It includes 33 construction drawings and 92 curated Q&A pairs across three reasoning depths, highlighting the need for AI integration in engineering workflows.
Key Points
- First benchmark for MLLMs on real-world construction drawings.
- Includes 33 'Issued for Construction' drawings and 92 Q&A pairs.
- Evaluates three reasoning depths: perceptual, contextual, and expert.
- Significant performance gap between MLLMs and human experts.
- Lays foundation for AI-driven understanding in engineering workflows.
DeepSignal Analysis
What happened
DrawingVQA is a newly introduced benchmark aimed at assessing multimodal large language models (MLLMs) specifically on construction drawings. It includes 33 construction drawings and 92 curated question-answer pairs that span three levels of reasoning. Evaluations indicate a significant performance gap between MLLMs and expert performance, particularly in complex reasoning tasks.
Key evidence
- DrawingVQA consists of 33 'Issued for Construction' drawings and 92 curated Q&A pairs across three reasoning depths: perceptual understanding, contextual interpretation, and domain-expert reasoning.
- The benchmark reveals a substantial gap between the performance of state-of-the-art MLLMs and expert performance, especially at higher reasoning depths.
- DrawingVQA is the first benchmark to explicitly map engineering workflows to AI reasoning competencies, highlighting the integration of AI in engineering.
Why it matters
The introduction of DrawingVQA addresses a critical need for evaluating AI capabilities in understanding complex construction drawings, which are essential in various engineering fields. The significant performance gaps identified suggest that current MLLMs may not yet be suitable for practical applications in engineering workflows. This benchmark could drive future research and development towards improving AI's reasoning abilities in specialized domains.
Paper Resources
Source Excerpt
We introduce DrawingVQA, the first benchmark designed to evaluate multimodal (MLLMs) on real-world construction drawings -- a core media in architecture, civil, and many other engineering practices. Unlike natural images or schematic floor plans, construction drawings fuse abstract geometry, symbolic notation, tabular data, annotations, and domain-specific text, forming a uniquely complex visual-textual domain core to engineering workflows. DrawingVQA bridges this gap with
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →Automatic Ordinary Differential Equations Discovery For Biological Systems Using Powered Agentic System
The MEDA system utilizes large language models and symbolic regression to autonomously discover ordinary differential equations for biological systems, achieving strong structural recovery and biologically plausible models. It outperforms existing methods by integrating domain knowledge and mechanistic constraints, demonstrating effective retrieval and extrapolation capabilities.