Optimizing Vision-Based CVS Risk Assessment: Benchmarking Geometric and Deep Learning Architectures for Real-Time Applications
This study benchmarks geometric and deep learning architectures for vision-based assessment of behavioral risk factors associated with Computer Vision Syndrome (CVS). Eye blinking, head pose deviation, and body twist were the three behaviors assessed. With accuracies of 0.9034 for blink detection and 0.9037 for head pose estimation, the results demonstrate that Dlib offered the most dependable performance for facial analysis. For head pose analysis, the highest F1-score was 0.7896 at a yaw threshold of 11°, with stable performance across the 9°–11° range. MediaPipe showed greater practical suitability for real-time body twist detection, with processing speeds of 29.98–32.79 FPS, although its maximum accuracy was 0.4668. In contrast, the CNN-based model was constrained by low processing speed, ranging from 4.22 to 7.32 FPS, making it less suitable for immediate deployment. These findings indicate that effective CVS risk assessment requires a balance between detection reliability and computational efficiency. Dlib is more appropriate for accuracy-sensitive facial measurements, whereas MediaPipe offers better practical value for real-time posture and movement monitoring.
Authors
- Chonnikarn Rodmorn
- Mathuros Panmuang
Institutions
- Rajamangala University of Technology (TH)
- King Mongkut's University of Technology North Bangkok (TH)
Publication Details
- Journal
- WSEAS Transactions on Information Science and Applications archive
- Published
- 2026-09-17
- DOI
- https://doi.org/10.37394/23209.2026.23.50
- Primary Topic
- Ergonomics and Musculoskeletal Disorders
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- King Mongkut's University of Technology North Bangkok
- Rajamangala University of Technology Thanyaburi
- Faculty of Applied Science, King Mongkut's University of Technology North Bangkok