Beyond transfer learning: a generative self-supervised framework for fMRI-based diagnosis on small and imbalanced datasets

Abstract Diagnosing neurological and psychiatric diseases (NeuroPsyD) from functional Magnetic Resonance Imaging (fMRI) using Deep Neural Networks (DNNs) is challenging, particularly when datasets are small and imbalanced, leading to severe overfitting, instability, and biased sensitivity–specificity trade-offs. Transfer Learning (TL), including supervised and self-supervised pre-training, can partially mitigate these challenges, but its effectiveness remains limited by domain shift, annotation bias, majority-class bias, and limited gains on very small datasets. To address these limitations, we propose a purely in-domain generative self-supervised framework that does not require external pre-training, called Boundary-aware Variational Autoencoder + Self-Supervised Mixup (BVAE+SSup-Mixup). The framework integrates SSup-Mixup for label-free representation learning, multivariate VAE-based minority-class oversampling, and boundary-aware synthetic sample selection, which retains informative generated samples near the decision boundary while reducing outliers and poorly positioned synthetic samples. The framework is evaluated on five fMRI-based NeuroPsyD diagnosis tasks covering extremely and moderately small, imbalanced datasets. Compared with supervised and self-supervised TL approaches and a state-of-the-art RHVAE-based generative baseline, BVAE+SSup-Mixup improved accuracy, F1-score, and AUC across most tasks, achieving accuracies of 87.7–95.0% and AUCs of 89.3–94.7%. It also produced a more balanced sensitivity–specificity profile, with both measures mostly above 85%. These balanced gains suggest that boundary-aware augmentation provides targeted minority-class evidence that TL may not fully exploit. These findings provide a proof of concept for a TL-competitive framework in severely data-limited and imbalanced medical imaging settings.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-29
DOI
https://doi.org/10.1038/s41598-026-73268-2
Primary Topic
Functional Brain Connectivity Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Beyond transfer learning: a generative self-supervised framework for fMRI-based diagnosis on small and imbalanced datasets

Ahmad Kalhor, Hamid Soltanian‐Zadeh, Saeed Masoudnia, Ershad Hassanpour Golagani
Scientific Reports
Functional Brain Connectivity Studies
article

Beyond transfer learning: a generative self-supervised framework for fMRI-based diagnosis on small and imbalanced datasets

Ahmad Kalhor, Hamid Soltanian‐Zadeh, Saeed Masoudnia, Ershad Hassanpour Golagani
article en

Abstract

Abstract Diagnosing neurological and psychiatric diseases (NeuroPsyD) from functional Magnetic Resonance Imaging (fMRI) using Deep Neural Networks (DNNs) is challenging, particularly when datasets are small and imbalanced, leading to severe overfitting, instability, and biased sensitivity–specificity trade-offs. Transfer Learning (TL), including supervised and self-supervised pre-training, can partially mitigate these challenges, but its effectiveness remains limited by domain shift, annotation bias, majority-class bias, and limited gains on very small datasets. To address these limitations, we propose a purely in-domain generative self-supervised framework that does not require external pre-training, called Boundary-aware Variational Autoencoder + Self-Supervised Mixup (BVAE+SSup-Mixup). The framework integrates SSup-Mixup for label-free representation learning, multivariate VAE-based minority-class oversampling, and boundary-aware synthetic sample selection, which retains informative generated samples near the decision boundary while reducing outliers and poorly positioned synthetic samples. The framework is evaluated on five fMRI-based NeuroPsyD diagnosis tasks covering extremely and moderately small, imbalanced datasets. Compared with supervised and self-supervised TL approaches and a state-of-the-art RHVAE-based generative baseline, BVAE+SSup-Mixup improved accuracy, F1-score, and AUC across most tasks, achieving accuracies of 87.7–95.0% and AUCs of 89.3–94.7%. It also produced a more balanced sensitivity–specificity profile, with both measures mostly above 85%. These balanced gains suggest that boundary-aware augmentation provides targeted minority-class evidence that TL may not fully exploit. These findings provide a proof of concept for a TL-competitive framework in severely data-limited and imbalanced medical imaging settings.

Scientific Reports
Henry Ford Health System (US), University of Tehran (IR), Henry Ford Hospital (US), Institute for Research in Fundamental Sciences (IR)
Openalex Percentile: Top 10%
Functional Brain Connectivity Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.