AlamViT: A Lightweight Spatiotemporal Facial Expression Model for Automated Pain Recognition
Automated facial pain recognition could support future monitoring tools for patients who cannot reliably self-report pain. Existing approaches either neglect temporal dynamics by processing single frames or employ computationally demanding models. This paper presents AlamViT, a computationally efficient video model that integrates Swin-style windowed spatial attention into a MobileViT backbone and appends a one-dimensional Swin temporal encoder. To the best of the authors’ knowledge, this particular compact spatial–temporal integration has not previously been evaluated for facial pain recognition. The model is evaluated independently on the BioVid Heat Pain Database and AI4PAIN using subject-disjoint partitions. A three-seed BioVid experiment using consistently reconstructed preprocessing achieved 54.07%±4.62% hold-out test accuracy. For context, the original finalized single-run BioVid experiment reported 60.13% test accuracy. On AI4PAIN, the binary model attained 90.97% on the validation set after epoch and threshold selection on that set; this value is not a hold-out test or generalization estimate because binary test labels were unavailable. A three-class variant achieved 51.33% test accuracy on AI4PAIN. The baseline contains 3.11 M parameters and requires approximately 2.2 GFLOPs per frame. These measurements characterize algorithmic efficiency; deployment performance on clinical or edge hardware was not evaluated.
Authors
- Talal Bonny (ORCID: https://orcid.org/0000-0003-1111-0304)
- Tamer Rabie (ORCID: https://orcid.org/0000-0003-4003-1592)
- Alya Alabdouli
- Mohammed Baziyad
Institutions
- University of Sharjah (AE)
Publication Details
- Journal
- Mathematics
- Published
- 2026-09-17
- DOI
- https://doi.org/10.3390/math14183376
- Primary Topic
- Emotion and Mood Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00