Vision–Language Model (VLM)-Based Interactive Visual Querying for Construction Ergonomic Risk Reasoning

Construction workers are frequently exposed to postural ergonomic risks, yet conventional AI-based ergonomic risk assessment (ERA) systems generally provide risk classifications or scores rather than interpretable and interactive explanations. This study investigates whether domain adaptation can improve a generic vision–language model’s ability to identify and explain construction ergonomic risks from images. A dataset with 1900 construction image–text pairs was curated. The domain-adapted model was compared with the baseline model. Performance was evaluated using visual question answering (VQA) accuracy, nine image-captioning metrics, and a blinded human evaluation involving 50 participants with varying levels of ergonomics knowledge. ErgoChat improved performance across the nine caption-evaluation metrics and VQA. In the human evaluation, ErgoChat-generated descriptions were selected as more accurate in 81.69% of comparisons. Limitations include dataset size and class imbalance, potential VLM hallucination, and visual-perception errors. Future research will expand and balance real-site data, evaluate prompt robustness, and strengthen integration with structured ergonomic assessment procedures. The primary contribution is therefore a construction-ergonomics-specific domain-adaptation and evaluation framework that enables interactive VQA and natural-language ergonomic risk reasoning.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-10-06
DOI
https://doi.org/10.3390/app16199887
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Vision–Language Model (VLM)-Based Interactive Visual Querying for Construction Ergonomic Risk Reasoning

Qipei Mei, Chao Fan, Xinming Li, Xiaonan Wang
Applied Sciences
Multimodal Machine Learning Applications
article

Vision–Language Model (VLM)-Based Interactive Visual Querying for Construction Ergonomic Risk Reasoning

Qipei Mei, Chao Fan, Xinming Li, Xiaonan Wang
article en

Abstract

Construction workers are frequently exposed to postural ergonomic risks, yet conventional AI-based ergonomic risk assessment (ERA) systems generally provide risk classifications or scores rather than interpretable and interactive explanations. This study investigates whether domain adaptation can improve a generic vision–language model’s ability to identify and explain construction ergonomic risks from images. A dataset with 1900 construction image–text pairs was curated. The domain-adapted model was compared with the baseline model. Performance was evaluated using visual question answering (VQA) accuracy, nine image-captioning metrics, and a blinded human evaluation involving 50 participants with varying levels of ergonomics knowledge. ErgoChat improved performance across the nine caption-evaluation metrics and VQA. In the human evaluation, ErgoChat-generated descriptions were selected as more accurate in 81.69% of comparisons. Limitations include dataset size and class imbalance, potential VLM hallucination, and visual-perception errors. Future research will expand and balance real-site data, evaluate prompt robustness, and strengthen integration with structured ergonomic assessment procedures. The primary contribution is therefore a construction-ergonomics-specific domain-adaptation and evaluation framework that enables interactive VQA and natural-language ergonomic risk reasoning.

Applied SciencesVol. 16(19)
University of Alberta (CA)
Openalex Percentile: Top 14%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Vision–Language Model (VLM)-Based Interactive Visual Querying for Construction Ergonomic Risk Reasoning — Qipei Mei, Chao Fan, et al. · Applied Sciences (2026) | TGRS Research Map | TGRS