Chao Fan, Qipei Mei, Xiaonan Wang, Xinming Li · Applied Sciences 2026 · 2026
DOI: 10.3390/app16199887
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Construction workers are frequently exposed to postural ergonomic risks, yet conventional AI-based ergonomic risk assessment (ERA) systems generally provide risk classifications or scores rather than interpretable and interactive explanations. This study investigates whether domain adaptation can improve a generic vision–language model’s ability to identify and explain construction ergonomic risks from images. A dataset with 1900 construction image–text pairs was curated. The domain-adapted model was compared with the baseline model. Performance was evaluated using visual question answering (VQA) accuracy, nine image-captioning metrics, and a blinded human evaluation involving 50 participants with varying levels of ergonomics knowledge. ErgoChat improved performance across the nine caption-evaluation metrics and VQA. In the human evaluation, ErgoChat-generated descriptions were selected as more accurate in 81.69% of comparisons. Limitations include dataset size and class imbalance, potential VLM hallucination, and visual-perception errors. Future research will expand and balance real-site data, evaluate prompt robustness, and strengthen integration with structured ergonomic assessment procedures. The primary contribution is therefore a construction-ergonomics-specific domain-adaptation and evaluation framework that enables interactive VQA and natural-language ergonomic risk reasoning.
No comments yet — start the discussion below.