Cross-Model Attention and Multimodal Embedding Fusion for Depression Detection Using Clinical Interviews

Abstract- Automatic depression screening from clinical interviews is a challenging task because depressive symptomscan appear across language, voice, facial behavior, and questionnaire responses. Existing systems often depend on a singlemodality or simple late fusion, which may miss complementary clinical evidence. This work proposes a cross-modelattention and multimodal embedding fusion framework for depression detection using clinical interview data. Theproposed framework represents each interview through text, audio, video-derived features, and PHQ-8-guided symptomevidence. The text branch uses contextual MPNet embeddings from interview responses, the audio branch capturesspeech-related cues such as prosody and vocal behavior, and the visual branch represents facial behavior usingexpression, gaze, head-pose, and action-unit features. These modality-specific representations are combined throughmultimodal embedding fusion, while cross-model attention assigns importance to complementary model outputs so thatthe final prediction is not dominated by one unreliable modality. The system is evaluated using depression-screeninglabels derived from PHQ-8 scores and clinical interview datasets such as DAIC-WOZ and extended DAIC-style data. Theofficial locked-test model achieved 76.6% accuracy, 0.593 F1-score, and 0.868 ROC-AUC on the DAIC-WOZ test split.The PHQ-guided out-of-fold fusion achieved 83.0% accuracy and 0.826 ROC-AUC, showing that symptom-aware crossmodel fusion can improve ranking and interpretability. The framework supports depression/non-depressionclassification, risk scoring, and modality-wise evidence analysis. Overall, the proposed approach improves screeningrobustness, interpretability, and controlled multimodal evaluation while remaining a research screening aid rather than amedical diagnostic tool.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-08
DOI
https://doi.org/10.5281/zenodo.23229014
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Cross-Model Attention and Multimodal Embedding Fusion for Depression Detection Using Clinical Interviews

Prof. Ch. Satyananda Reddy Pukkala Lavanya
Zenodo (CERN European Organization for Nuclear Research)
Emotion and Mood Recognition
article

Cross-Model Attention and Multimodal Embedding Fusion for Depression Detection Using Clinical Interviews

Prof. Ch. Satyananda Reddy Pukkala Lavanya
article en

Abstract

Abstract- Automatic depression screening from clinical interviews is a challenging task because depressive symptomscan appear across language, voice, facial behavior, and questionnaire responses. Existing systems often depend on a singlemodality or simple late fusion, which may miss complementary clinical evidence. This work proposes a cross-modelattention and multimodal embedding fusion framework for depression detection using clinical interview data. Theproposed framework represents each interview through text, audio, video-derived features, and PHQ-8-guided symptomevidence. The text branch uses contextual MPNet embeddings from interview responses, the audio branch capturesspeech-related cues such as prosody and vocal behavior, and the visual branch represents facial behavior usingexpression, gaze, head-pose, and action-unit features. These modality-specific representations are combined throughmultimodal embedding fusion, while cross-model attention assigns importance to complementary model outputs so thatthe final prediction is not dominated by one unreliable modality. The system is evaluated using depression-screeninglabels derived from PHQ-8 scores and clinical interview datasets such as DAIC-WOZ and extended DAIC-style data. Theofficial locked-test model achieved 76.6% accuracy, 0.593 F1-score, and 0.868 ROC-AUC on the DAIC-WOZ test split.The PHQ-guided out-of-fold fusion achieved 83.0% accuracy and 0.826 ROC-AUC, showing that symptom-aware crossmodel fusion can improve ranking and interpretability. The framework supports depression/non-depressionclassification, risk scoring, and modality-wise evidence analysis. Overall, the proposed approach improves screeningrobustness, interpretability, and controlled multimodal evaluation while remaining a research screening aid rather than amedical diagnostic tool.

Zenodo (CERN European Organization for Nuclear Research)
Andhra University (IN)
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Cross-Model Attention and Multimodal Embedding Fusion for Depression Detection Using Clinical Interviews — Prof. Ch. Satyananda Reddy Pukkala Lavanya · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS