Enabling Cross-Study Comparison: A Method for Automated BIM-Qa Evaluation
DOI: 10.35490/EC3.2026.347
Abstract: Driven by LLM advances, Question Answering (QA) systems for Building Information Models (BIM) have proliferated, yet no standardized methods exist for cross-study comparison. We present an automated evaluation pipeline employing LLM judges to assess BIM-QA answers across five quality criteria and validate it through inter-rater agreement analysis with two domain experts and two LLM judges on 74 IFC-Bench question-answer pairs. LLM judges achieve high inter-rater reliability (Krippendorff’s α = 0.70–1.00), exceeding human experts (α = 0.32–0.57), and show good agreement with expert consensus (α = 0.48–0.85), positioning the pipeline as a suitable solution for automated, reproducible BIM-QA benchmarking.
Keywords: Benchmarking, Building Information Modeling, Evaluation, Large Language Models, Question Answering