Transformer-Based Multi-Task Learning with a Task-Aware Weighted Ensemble for Simultaneous Age Estimation and Gender Classification from Facial Images

Accurate age estimation and gender classification from facial images remain challenging tasks owing to substantial variations in facial appearance caused by pose, illumination, expression, occlusion, image quality, and age-related morphological changes. Although transformer-based architectures have recently demonstrated remarkable performance in facial analysis, effectively exploiting their complementary representation capabilities within a unified multi-task framework remains an open research problem. To address this limitation, this study proposes a transformer-based Multi-task Learning (MTL) framework that jointly performs age estimation and gender classification using a shared feature extraction backbone and task-specific prediction heads. Three state-of-the-art vision transformers, namely Vision Transformer (ViT), Shifted Window Transformer (Swin Transformer), and Data-efficient Image Transformer (DeiT), were independently implemented within the proposed MTL architecture to learn both global contextual information and local facial characteristics. Furthermore, two ensemble learning strategies, namely standard ensemble and Task-aware Weighted Ensemble (TAWE), were developed to exploit the complementary strengths of the individual transformer models. Experiments were conducted on the publicly available UTKFace dataset comprising 15,503 facial images, which were divided into 10,385 training, 1833 validation, and 3285 test samples, while gender balancing was applied during model training to alleviate class imbalance. Age estimation performance was evaluated using Mean Absolute Error (MAE) and Root Mean Squared Error (RMSE), whereas gender classification performance was assessed using Accuracy and F1-score. Experimental results demonstrate that the proposed TAWE achieved the best point estimates among the evaluated ensemble strategies, obtaining an MAE of 4.8905, an RMSE of 6.4674, a gender classification accuracy of 94.86%, and an F1-score of 95.17% for the dataset-defined positive class. Compared with equal-weight averaging, TAWE provided a statistically significant improvement in gender classification accuracy, while the improvement in age estimation was modest and not statistically significant.

Authors

Institutions

Publication Details

Journal
Sensors
Published
2026-09-21
DOI
https://doi.org/10.3390/s26185982
Primary Topic
Face recognition and analysis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Transformer-Based Multi-Task Learning with a Task-Aware Weighted Ensemble for Simultaneous Age Estimation and Gender Classification from Facial Images

Erdal Özbay, Sümeyye Sarıateş
Sensors
Face recognition and analysis
article

Transformer-Based Multi-Task Learning with a Task-Aware Weighted Ensemble for Simultaneous Age Estimation and Gender Classification from Facial Images

Erdal Özbay, Sümeyye Sarıateş
article en

Abstract

Accurate age estimation and gender classification from facial images remain challenging tasks owing to substantial variations in facial appearance caused by pose, illumination, expression, occlusion, image quality, and age-related morphological changes. Although transformer-based architectures have recently demonstrated remarkable performance in facial analysis, effectively exploiting their complementary representation capabilities within a unified multi-task framework remains an open research problem. To address this limitation, this study proposes a transformer-based Multi-task Learning (MTL) framework that jointly performs age estimation and gender classification using a shared feature extraction backbone and task-specific prediction heads. Three state-of-the-art vision transformers, namely Vision Transformer (ViT), Shifted Window Transformer (Swin Transformer), and Data-efficient Image Transformer (DeiT), were independently implemented within the proposed MTL architecture to learn both global contextual information and local facial characteristics. Furthermore, two ensemble learning strategies, namely standard ensemble and Task-aware Weighted Ensemble (TAWE), were developed to exploit the complementary strengths of the individual transformer models. Experiments were conducted on the publicly available UTKFace dataset comprising 15,503 facial images, which were divided into 10,385 training, 1833 validation, and 3285 test samples, while gender balancing was applied during model training to alleviate class imbalance. Age estimation performance was evaluated using Mean Absolute Error (MAE) and Root Mean Squared Error (RMSE), whereas gender classification performance was assessed using Accuracy and F1-score. Experimental results demonstrate that the proposed TAWE achieved the best point estimates among the evaluated ensemble strategies, obtaining an MAE of 4.8905, an RMSE of 6.4674, a gender classification accuracy of 94.86%, and an F1-score of 95.17% for the dataset-defined positive class. Compared with equal-weight averaging, TAWE provided a statistically significant improvement in gender classification accuracy, while the improvement in age estimation was modest and not statistically significant.

SensorsVol. 26(18)
Fırat University (TR)
Quality Education
Openalex Percentile: Top 13%
Face recognition and analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.