1 Institute of Concrete Structures, TUD Dresden University of Technology, Dresden, Germany
2 Technische Universität Dresden, Dresden, Germany
3 TUD Dresden University of Technology, Dresden, Germany
4 Institute of Concrete Structures, TU Dresden, Dresden, Germany
DOI: 10.35490/EC3.2026.252
Abstract: Title-blocks in engineering drawings contain essential metadata, but automated interpretation remains difficult because of historical scans and heterogeneous layouts. This paper presents a vision-language model (VLM)-based workflow for title-block detection and structured information extraction using the open-source Qwen2.5-VL backbone. The workflow combines promptable visual grounding on full drawings with schema-constrained extraction from cropped title-block images. The study focuses on three representative fields: scale, drawing name, and route. Results show promising performance for both tasks and indicate that open-source VLMs can provide a simplified alternative to conventional multi-stage pipelines for architecture, engineering, and construction (AEC) document analysis.
Keywords: Engineering drawings, open-source vision-language model (VLM), structured information extraction, title-block detection