Vision–Language Model (VLM)-Based Interactive Visual Querying for Construction Ergonomic Risk Reasoning
Construction workers are frequently exposed to postural ergonomic risks, yet conventional AI-based ergonomic risk assessment (ERA) systems generally provide risk classifications or scores rather than interpretable and interactive explanations. This study investigates whether domain adaptation can improve a generic vision–language model’s ability to identify and explain construction ergonomic risks from images. A dataset with 1900 construction image–text pairs was curated. The domain-adapted model was compared with the baseline model. Performance was evaluated using visual question answering (VQA) accuracy, nine image-captioning metrics, and a blinded human evaluation involving 50 participants with varying levels of ergonomics knowledge. ErgoChat improved performance across the nine caption-evaluation metrics and VQA. In the human evaluation, ErgoChat-generated descriptions were selected as more accurate in 81.69% of comparisons. Limitations include dataset size and class imbalance, potential VLM hallucination, and visual-perception errors. Future research will expand and balance real-site data, evaluate prompt robustness, and strengthen integration with structured ergonomic assessment procedures. The primary contribution is therefore a construction-ergonomics-specific domain-adaptation and evaluation framework that enables interactive VQA and natural-language ergonomic risk reasoning.
Authors
- Qipei Mei (ORCID: https://orcid.org/0000-0003-1409-3562)
- Chao Fan (ORCID: https://orcid.org/0000-0003-4367-4219)
- Xinming Li (ORCID: https://orcid.org/0000-0001-6802-033X)
- Xiaonan Wang
Institutions
- University of Alberta (CA)
Publication Details
- Journal
- Applied Sciences
- Published
- 2026-10-06
- DOI
- https://doi.org/10.3390/app16199887
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00