Enhancing interpretability in metabolomics: ranking metabolites by their impact on graph neural network predictions
Abstract A key step in the biological interpretation of untargeted metabolomics data is the ranking of compounds by importance, usually using statistical significance or impact on classification as the basis for importance assignment. However, current approaches treat metabolites as independent variables and rarely incorporate the network structure underlying biochemical relationships. This disconnect limits the ability of existing ranking methods to highlight groups of interconnected metabolites that jointly contribute to biological differences. Here we developed a new method where Formula Difference Networks are used as inputs to predictive Graph Neural Network models. After fitting, metabolites were ranked by their impact on sample class prediction probabilities. When applied to three benchmark datasets, this ranking highlighted subnetworks of metabolites, favouring connectivity as a driving factor for importance assignment. This led to an enrichment of the number of edges between the top ranked compounds. Furthermore, using datasets containing metabolites with simulated significance, we found that there was a clear bias for assigning higher importance to nodes in connected subgraphs. This new strategy for Graph Neural Network interpretability, is an alternative to common approaches based on mapping of important metabolite onto biological pathways supported by enrichment analysis.
Authors
- António E. N. Ferreira (ORCID: https://orcid.org/0000-0002-9625-8115)
- Carlos Cordeiro (ORCID: https://orcid.org/0000-0001-6327-8606)
- Francisco Traquete (ORCID: https://orcid.org/0000-0002-4081-6544)
- Marta Sousa Silva
Publication Details
- Journal
- BMC Bioinformatics
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1186/s12859-026-06633-7
- Primary Topic
- Metabolomics and Mass Spectrometry Studies
- Type
- article
- Field-Weighted Citation Impact
- 0.00