AlamViT: A Lightweight Spatiotemporal Facial Expression Model for Automated Pain Recognition

Automated facial pain recognition could support future monitoring tools for patients who cannot reliably self-report pain. Existing approaches either neglect temporal dynamics by processing single frames or employ computationally demanding models. This paper presents AlamViT, a computationally efficient video model that integrates Swin-style windowed spatial attention into a MobileViT backbone and appends a one-dimensional Swin temporal encoder. To the best of the authors’ knowledge, this particular compact spatial–temporal integration has not previously been evaluated for facial pain recognition. The model is evaluated independently on the BioVid Heat Pain Database and AI4PAIN using subject-disjoint partitions. A three-seed BioVid experiment using consistently reconstructed preprocessing achieved 54.07%±4.62% hold-out test accuracy. For context, the original finalized single-run BioVid experiment reported 60.13% test accuracy. On AI4PAIN, the binary model attained 90.97% on the validation set after epoch and threshold selection on that set; this value is not a hold-out test or generalization estimate because binary test labels were unavailable. A three-class variant achieved 51.33% test accuracy on AI4PAIN. The baseline contains 3.11 M parameters and requires approximately 2.2 GFLOPs per frame. These measurements characterize algorithmic efficiency; deployment performance on clinical or edge hardware was not evaluated.

Authors

Institutions

Publication Details

Journal
Mathematics
Published
2026-09-17
DOI
https://doi.org/10.3390/math14183376
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

AlamViT: A Lightweight Spatiotemporal Facial Expression Model for Automated Pain Recognition

Talal Bonny, Tamer Rabie, Alya Alabdouli, Mohammed Baziyad
Mathematics
Emotion and Mood Recognition
article

AlamViT: A Lightweight Spatiotemporal Facial Expression Model for Automated Pain Recognition

Talal Bonny, Tamer Rabie, Alya Alabdouli, Mohammed Baziyad
article en

Abstract

Automated facial pain recognition could support future monitoring tools for patients who cannot reliably self-report pain. Existing approaches either neglect temporal dynamics by processing single frames or employ computationally demanding models. This paper presents AlamViT, a computationally efficient video model that integrates Swin-style windowed spatial attention into a MobileViT backbone and appends a one-dimensional Swin temporal encoder. To the best of the authors’ knowledge, this particular compact spatial–temporal integration has not previously been evaluated for facial pain recognition. The model is evaluated independently on the BioVid Heat Pain Database and AI4PAIN using subject-disjoint partitions. A three-seed BioVid experiment using consistently reconstructed preprocessing achieved 54.07%±4.62% hold-out test accuracy. For context, the original finalized single-run BioVid experiment reported 60.13% test accuracy. On AI4PAIN, the binary model attained 90.97% on the validation set after epoch and threshold selection on that set; this value is not a hold-out test or generalization estimate because binary test labels were unavailable. A three-class variant achieved 51.33% test accuracy on AI4PAIN. The baseline contains 3.11 M parameters and requires approximately 2.2 GFLOPs per frame. These measurements characterize algorithmic efficiency; deployment performance on clinical or edge hardware was not evaluated.

MathematicsVol. 14(18)
University of Sharjah (AE)
Openalex Percentile: Top 8%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

AlamViT: A Lightweight Spatiotemporal Facial Expression Model for Automated Pain Recognition — Talal Bonny, Tamer Rabie, et al. · Mathematics (2026) | TGRS Research Map | TGRS