BIBIMBAP: A Benchmark for Instructional BIM-Based Automated Programming BIBIMBAP: A Benchmark for Instructional BIM-Based Automated Programming

Clemens Kujat1, Bharathi Kannan Nithyanantham1, Stefan Telgmann1, Tobias Sesterhenn2,3, Jörn Plönnigs1, Stefan Lüdtke1, Christian Bartelt2,3
1 University of Rostock, Rostock, Germany
2 Technical University of Clausthal, Clausthal-Zellerfeld, Germany
3 Technische Universität Clausthal, Clausthal-Zellerfeld, Germany
DOI: 10.35490/EC3.2026.271
Abstract: Large Language Models (LLMs) are emerging as natural-language interfaces for Building Information Modeling (BIM), successfully achieving information retrieval, but no benchmark exists to evaluate their ability to edit Industry Foundation Classes (IFC). We introduce BIBIMBAP, a Text-to-BIM benchmark with 100 curated tasks covering Create-Read-Update-Delete (CRUD) operations across spatial, geometric, topological, numeric, and conceptual categories. Each task includes a natural-language prompt, an IFC model, expected structured outputs, and test scripts for automated evaluation. Baseline results show limited performance (best: 50.2%) and frequent failures in constraint preservation and spatial reasoning, with some correct outputs achieved despite flawed reasoning.
Keywords: Benchmark, BIM, IFC, LLMs, Text-to-BIM

Presentation video

Successfully submitted

Your submission has been received. We will review your details and contact you soon.