Rigorous validation of graph-based network analysis reappraises biologically coherent ADHD-associated transcriptomic modules in peripheral blood
Background Graph attention networks (GATs) are increasingly applied to transcriptomic data because they integrate gene-network structure while producing attention weights that are often interpreted as indicators of biological importance. However, whether attention-derived explanations reliably reflect biologically meaningful signals has received little systematic evaluation, particularly in the small, heterogeneous cohorts common in psychiatric transcriptomics. Methods We systematically re-evaluated a GAT using peripheral blood RNA-sequencing data from 76 individuals (39 ADHD, 37 controls), including 16 discordant monozygotic twin pairs. To maximize analytical rigor, we implemented a leakage-aware pipeline, incorporating twin-aware group-stratified cross-validation, fold-wise gene selection, corrected transcript-to-gene mapping, and validation through repeated cross-validation, permutation testing, three complementary feature-importance methods, orthogonal differential-expression and pathway analyses, and data-quality controls. Results Across 20 repeated cross-validation splits, GAT achieved a slightly higher mean AUC than a graph convolutional network (mean AUC 0.624 vs. 0.599) and outperformed five classical machine-learning models on a prespecified split. However, GAT performance was not statistically distinguishable from a rigorously matched permutation-derived null distribution (p = 0.327), while an exploratory sample-size calculation indicated that roughly twice the current sample size would be needed to detect an effect this size at conventional power. Independent validation analyses converged on the same conclusion: attention-derived gene rankings were significantly negatively correlated with both SHAP and permutation importance, whereas differential-expression analyses---including discordant twin-pair comparisons---and pathway enrichment identified no reproducible biological signals after multiple-testing correction. Recovery of established co-expression modules and the absence of blood-cell marker differences supported pipeline validity. Conclusions Graph-based models may show modest predictive gains, but predictive performance and biological interpretation are distinct scientific questions. Our findings demonstrate that attention-derived importance should be regarded as hypothesis-generating rather than mechanistic evidence unless independently validated, illustrating why rigorous, multi-level validation is essential before biological conclusions are drawn from graph neural networks in transcriptomic research.
Authors
- Nehal M. Ali (ORCID: https://orcid.org/0000-0001-8543-6660)
Publication Details
- Journal
- PLoS ONE
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1371/journal.pone.0359656
- Primary Topic
- Bioinformatics and Genomic Networks
- Type
- article
- Field-Weighted Citation Impact
- 0.00