Vision–Language Models in Construction: Assessing Prompt Robustness in Open-Vocabulary Object Detection