3D Segmentation Agent Integrating 3D Gaussian Splatting and Vision-Language Models for as-Built Analysis

Boyu Wang1, Jun Ma1, Borja García de Soto2
1 University of Hong Kong, Hong Kong, China
2 New York University Abu Dhabi, Abu Dhabi, United Arab Emirates
DOI: 10.35490/EC3.2026.419
Abstract: The transition from raw as-built data to semantic object representations is a bottleneck in construction digitalization, currently hampered by manual dependencies and the “black-box” nature of existing point cloud segmentation networks. This paper introduces an autonomous self-correcting 3D segmentation agent tailored for construction environments, utilizing 3D Gaussian Splatting (3DGS) as the underlying representation and Vision Language Models (VLMs) as the reasoning core. Unlike static segmentation methods, our agent adopts a dynamic approach by performing active scene roaming. Through a closed-loop feedback mechanism, the VLM analyzes visual data, detects objects, and rigorously assesses the plausibility of segmentation results. The agent iteratively corrects its own errors until a high-fidelity, logically consistent segmentation is achieved. This method demonstrates superior interpretability and accuracy, enabling downstream applications such as as-built verification, progress tracking, and quantity takeoff. By bridging generative rendering with semantic reasoning, our approach provides a robust framework for automated 3D scene understanding.
Keywords: 3D Gaussian Splatting, 3D Segmentation, AI Agent, As-built Modeling, Vision-Language Model

Presentation video

Successfully submitted

Your submission has been received. We will review your details and contact you soon.