Complementary Global–Local Feature Fusion and Ensemble Refinement for Facial-Expression Recognition on FER2013
Facial-expression recognition (FER) on FER2013 remains challenging because of low-resolution images, class imbalance, and label ambiguity. This study presents a global–local feature-fusion framework that integrates complementary representations with validation-based ensemble refinement. A frozen DINOv2 ViT-Base captures global facial semantics, while EfficientNetB3 extracts complementary local texture features. Their fused representation is used for seven-class facial-expression classification. The classification head is first trained with targeted feature-space SMOTE, and the EfficientNetB3 branch is then partially fine-tuned. Five-view test-time augmentation (TTA) is further incorporated at inference, together with an independently trained ConvNeXt-Tiny branch to provide additional architectural diversity. Ensemble weights are selected using a held-out validation set, while the test partition is reserved for final evaluation. On the 7178-image FER2013 test partition, the resulting ensemble achieves 76.51% accuracy, 75.77% macro F1, 76.32% weighted F1, 0.7158 Cohen’s kappa, and 0.9574 macro ROC AUC. The results demonstrate the potential of combining complementary global–local representations, inference augmentation, and ensemble refinement for facial-expression recognition on FER2013.
Authors
- H. M. Shahzad (ORCID: https://orcid.org/0000-0002-2452-6571)
- Hassan A. Ahmed (ORCID: https://orcid.org/0000-0002-0683-532X)
Institutions
- Cleveland State University (US)
- Superior University (PK)
Publication Details
- Journal
- Information
- Published
- 2026-10-06
- DOI
- https://doi.org/10.3390/info17100982
- Primary Topic
- Emotion and Mood Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00