Complementary Global–Local Feature Fusion and Ensemble Refinement for Facial-Expression Recognition on FER2013

Facial-expression recognition (FER) on FER2013 remains challenging because of low-resolution images, class imbalance, and label ambiguity. This study presents a global–local feature-fusion framework that integrates complementary representations with validation-based ensemble refinement. A frozen DINOv2 ViT-Base captures global facial semantics, while EfficientNetB3 extracts complementary local texture features. Their fused representation is used for seven-class facial-expression classification. The classification head is first trained with targeted feature-space SMOTE, and the EfficientNetB3 branch is then partially fine-tuned. Five-view test-time augmentation (TTA) is further incorporated at inference, together with an independently trained ConvNeXt-Tiny branch to provide additional architectural diversity. Ensemble weights are selected using a held-out validation set, while the test partition is reserved for final evaluation. On the 7178-image FER2013 test partition, the resulting ensemble achieves 76.51% accuracy, 75.77% macro F1, 76.32% weighted F1, 0.7158 Cohen’s kappa, and 0.9574 macro ROC AUC. The results demonstrate the potential of combining complementary global–local representations, inference augmentation, and ensemble refinement for facial-expression recognition on FER2013.

Authors

Institutions

Publication Details

Journal
Information
Published
2026-10-06
DOI
https://doi.org/10.3390/info17100982
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Complementary Global–Local Feature Fusion and Ensemble Refinement for Facial-Expression Recognition on FER2013

H. M. Shahzad, Hassan A. Ahmed
Information
Emotion and Mood Recognition
article

Complementary Global–Local Feature Fusion and Ensemble Refinement for Facial-Expression Recognition on FER2013

H. M. Shahzad, Hassan A. Ahmed
article en

Abstract

Facial-expression recognition (FER) on FER2013 remains challenging because of low-resolution images, class imbalance, and label ambiguity. This study presents a global–local feature-fusion framework that integrates complementary representations with validation-based ensemble refinement. A frozen DINOv2 ViT-Base captures global facial semantics, while EfficientNetB3 extracts complementary local texture features. Their fused representation is used for seven-class facial-expression classification. The classification head is first trained with targeted feature-space SMOTE, and the EfficientNetB3 branch is then partially fine-tuned. Five-view test-time augmentation (TTA) is further incorporated at inference, together with an independently trained ConvNeXt-Tiny branch to provide additional architectural diversity. Ensemble weights are selected using a held-out validation set, while the test partition is reserved for final evaluation. On the 7178-image FER2013 test partition, the resulting ensemble achieves 76.51% accuracy, 75.77% macro F1, 76.32% weighted F1, 0.7158 Cohen’s kappa, and 0.9574 macro ROC AUC. The results demonstrate the potential of combining complementary global–local representations, inference augmentation, and ensemble refinement for facial-expression recognition on FER2013.

InformationVol. 17(10)
Cleveland State University (US), Superior University (PK)
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.