BIBIMBAP: A Benchmark for Instructional BIM-Based Automated Programming BIBIMBAP: A Benchmark for Instructional BIM-Based Automated Programming
DOI: 10.35490/EC3.2026.271
Abstract: Large Language Models (LLMs) are emerging as natural-language interfaces for Building Information Modeling (BIM), successfully achieving information retrieval, but no benchmark exists to evaluate their ability to edit Industry Foundation Classes (IFC). We introduce BIBIMBAP, a Text-to-BIM benchmark with 100 curated tasks covering Create-Read-Update-Delete (CRUD) operations across spatial, geometric, topological, numeric, and conceptual categories. Each task includes a natural-language prompt, an IFC model, expected structured outputs, and test scripts for automated evaluation. Baseline results show limited performance (best: 50.2%) and frequent failures in constraint preservation and spatial reasoning, with some correct outputs achieved despite flawed reasoning.
Keywords: Benchmark, BIM, IFC, LLMs, Text-to-BIM