Facial video-based non-contact stress recognition: a multi-stage framework utilizing contrastive learning and attention mechanisms
Psychological stress plays a critical role in cardiovascular health, motivating the need for reliable, non-contact monitoring systems suitable for real-world environments. While remote photoplethysmography (rPPG) has enabled contactless physiological sensing, existing approaches often struggle to model the complex dynamics of multi-level stress responses. In this paper, we propose MS-CAM-Net (Multi-Stage Contrastive Attention Mechanism Network), a multi-stage learning framework for stress recognition from facial videos. The framework consists of three stages: self-supervised feature learning, attention-based rPPG reconstruction, and multi-task stress classification. Under the original fixed subject-independent test split, MS-CAM-Net achieved 99.2% accuracy for binary stress-state recognition and 86.7% accuracy for three-class stress-level classification. These values were numerically higher than the highest corresponding accuracies listed in the cross-study comparison by 3.1 and 1.6 percentage points, respectively. These results highlight the effectiveness of multi-stage physiological modeling for contactless stress monitoring under the reported evaluation setting.
Authors
- Khalil Alipour (ORCID: https://orcid.org/0000-0003-4456-1179)
- Mohammad Ghamari (ORCID: https://orcid.org/0000-0002-7752-4270)
- Bahram Tarvirdizadeh (ORCID: https://orcid.org/0000-0001-9851-3805)
- Alaa Hajr
- Hadi Zare
Institutions
- California Polytechnic State University (US)
- University of Tehran (IR)
- Iran University of Science and Technology (IR)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-11
- DOI
- https://doi.org/10.1038/s41598-026-70838-2
- Primary Topic
- Non-Invasive Vital Sign Monitoring
- Type
- article
- Field-Weighted Citation Impact
- 0.00