Task conditioned fusion of heterogeneous representation sources for fine grained affective state recognition

Automated affective-state recognition from non-intrusive video is a key enabler of Artificial Intelligence in Education, yet progress on the DAiSEE benchmark is limited by three structural challenges: single-encoder pipelines exploit only one view of the clip, the four affective dimensions are strongly correlated but typically modelled as independent tasks, and the dataset is severely class-imbalanced. We propose MAGCAF (multi-source affective gated cross-attention fusion), a task-conditioned multi-source heterogeneous-fusion framework that integrates four frozen pretrained sources of complementary representations of the same RGB clip—face-domain identity (InceptionResnetV1 / VGGFace2), self-supervised video (VideoMAE), supervised spatiotemporal video (TimeSformer), and 3D facial geometry (MediaPipe FaceMesh)—through a per-task gated cross-attention layer, an $$\Omega $$ task-correlation head regularised toward the empirical label-correlation prior, and a Kendall–Gal homoscedastic uncertainty-weighted multi-task loss. Under a unified 3-seed protocol that covers LRCN, ResNet-TCN, TimeSformer, VideoMAE, and ViBED-Net, MAGCAF reaches $$64.09\,\%$$ Top-1 accuracy, 0.286 Macro-F1, and 0.658 macro-AUC on the DAiSEE test set, exceeding the strongest single-source baseline by $$+3.24$$ pp, $$+0.013$$ , and $$+0.016$$ respectively, with the accuracy gap statistically significant under McNemar’s paired test ( $$p<0.01$$ ). Per-class analysis shows that the rarest intensity levels remain largely unrecognised; the Macro-F1 gain is therefore a relative improvement over single-source baselines rather than a resolution of the class-imbalance problem. Component and source-dependency ablations confirm that all three methodological elements contribute complementary gains and reveal a differentiated per-task source-attention pattern across the four affective dimensions.

Authors

Institutions

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-09-24
DOI
https://doi.org/10.1007/s44163-026-02285-x
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Task conditioned fusion of heterogeneous representation sources for fine grained affective state recognition

Tongcheng Geng, Lei Xun, Yubin Qu
Discover Artificial Intelligence
Emotion and Mood Recognition
article

Task conditioned fusion of heterogeneous representation sources for fine grained affective state recognition

Tongcheng Geng, Lei Xun, Yubin Qu
article en

Abstract

Automated affective-state recognition from non-intrusive video is a key enabler of Artificial Intelligence in Education, yet progress on the DAiSEE benchmark is limited by three structural challenges: single-encoder pipelines exploit only one view of the clip, the four affective dimensions are strongly correlated but typically modelled as independent tasks, and the dataset is severely class-imbalanced. We propose MAGCAF (multi-source affective gated cross-attention fusion), a task-conditioned multi-source heterogeneous-fusion framework that integrates four frozen pretrained sources of complementary representations of the same RGB clip—face-domain identity (InceptionResnetV1 / VGGFace2), self-supervised video (VideoMAE), supervised spatiotemporal video (TimeSformer), and 3D facial geometry (MediaPipe FaceMesh)—through a per-task gated cross-attention layer, an $$\Omega $$ task-correlation head regularised toward the empirical label-correlation prior, and a Kendall–Gal homoscedastic uncertainty-weighted multi-task loss. Under a unified 3-seed protocol that covers LRCN, ResNet-TCN, TimeSformer, VideoMAE, and ViBED-Net, MAGCAF reaches $$64.09\,\%$$ Top-1 accuracy, 0.286 Macro-F1, and 0.658 macro-AUC on the DAiSEE test set, exceeding the strongest single-source baseline by $$+3.24$$ pp, $$+0.013$$ , and $$+0.016$$ respectively, with the accuracy gap statistically significant under McNemar’s paired test ( $$p<0.01$$ ). Per-class analysis shows that the rarest intensity levels remain largely unrecognised; the Macro-F1 gain is therefore a relative improvement over single-source baselines rather than a resolution of the class-imbalance problem. Component and source-dependency ablations confirm that all three methodological elements contribute complementary gains and reveal a differentiated per-task source-attention pattern across the four affective dimensions.

Discover Artificial IntelligenceVol. 6(1)
China Internet Network Information Center (CN), Computer Network Information Center (CN)
Quality Education
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.