Vision–Language Models in Construction: Assessing Prompt Robustness in Open-Vocabulary Object Detection
DOI: 10.35490/EC3.2026.304
Abstract: Vision–language models offer a promising alternative to supervised computer vision by enabling open-set object recognition without retraining or large annotated datasets. Recent advances in open-vocabulary object detection models, such as GroundingDINO, support object grounding through natural language prompts; however, their sensitivity to prompt formulation remains underexplored in the AEC domain. This study evaluates the reliability of GroundingDINO for construction objects detection by systematically analysing prompt variation. Lexical variation and prompt multiplicity are examined. A custom dataset of 48 construction objects was evaluated under eight prompt configurations. Results show that lexical variation significantly affects detection performance, with “regular base term” achieving the highest F1 score of 56.1%, and multiple prompting decreases performance by about 2%, highlighting the effectiveness and limitations of AEC applications.
Keywords: Construction Engineering, Open-vocabulary Object Detection, prompt engineering, vision language models