Wenrui Zhu, Junqi Yu, Chunyong Feng, Zhengping Dong, Zhengwei Song, Jingyu Gao, Ruifeng Huang · Developments in the Built Environment 2026 · 2026
DOI: 10.1016/j.dibe.2026.101047
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Vision-language models (VLMs) offer a promising paradigm for zero-shot scenario understanding without task-specific training, yet existing benchmarks rely on costly manual annotation and generic scenes that fail to reflect the unstructured, dynamic nature of indoor construction sites. This case study introduces a label-free evaluation benchmark dataset, which comprises 1154 raw video clips captured from two real unstructured indoor construction sites in Xi’an, China. Nine state-of-the-art open-source VLMs are evaluated via five standardized prompt-driven tasks. The evaluation framework bypasses ground-truth annotations by introducing six proxy metrics (self-consistency, cross-model consensus, domain vocabulary coverage, hallucination suppression rate, structured parsing success, and inference efficiency) to objectively quantify model performance. Data collection and VLM inference are carried out collaboratively between a cloud server and the edge devices onboard the inspection robots, providing practical guidance on model selection for demand-driven construction scenario understanding tasks. The data and code are publicly available at https://github.com/redmogirill111/VLM-bench-for-indoor-construction-scenario-understanding .
No comments yet — start the discussion below.