A new benchmark called DrawingVQA evaluates how well multimodal AI models can understand real-world construction drawings, which combine geometry, symbols, tables, and domain-specific text. The dataset includes 33 professional construction drawings with 92 expert-curated questions across three reasoning levels, revealing significant gaps between current state-of-the-art models and human expert performance, especially on complex domain-specific tasks.
Why it matters: As AI integration accelerates across engineering and construction industries, this benchmark establishes critical performance standards and identifies where MLLMs fall short on real-world technical documents—essential data for enterprises evaluating AI adoption in these sectors.