Graph Neural Network Modelling of Pre-Diagnostic Serum Proteomics for Pancreatic Cancer Detection

Background/Objectives: Pancreatic ductal adenocarcinoma (PDAC) has exceptionally high mortality, largely because of late diagnosis, while established blood biomarkers such as CA19-9 have limited sensitivity for early detection. We evaluated whether modelling serum proteomic measurements as supervised pairwise statistical graphs using synolitic graph neural networks (SGNNs) could provide a useful framework for pre-diagnostic PDAC classification. Methods: We analysed 217 pre-diagnostic serum samples from UKCTOCS, comprising 100 PDAC samples from 75 women and 117 control samples, profiled for 97 proteins. Samples were divided into a development set collected 0–1 year before diagnosis (n = 110) and a temporal holdout set collected 1–2 years before diagnosis (n = 107); 25 PDAC participants contributed samples to both windows. The final SGNN used fold-internal mutual-information selection of 30 proteins, pairwise RBF-SVM-derived edge information, minimum-connected sparsification, and a 10-member GATv2 ensemble per cross-validation fold. Conventional machine-learning comparators were tuned using development data only. Results: The SGNN achieved a mean 5-fold development ROC-AUC of 75.33 ± 10.43%, lower than the tuned Random Forest (81.67%); the other tuned comparators achieved 77.67% (XGBoost), 77.17% (logistic regression), and 73.17% (SVM). Averaging predictions across the five-fold-specific SGNN ensembles yielded a temporal-holdout ROC-AUC of 64.35% (95% CI 52.6–74.6%). In the participant-independent subset of 25 PDAC cases not represented in development and 57 controls, ROC-AUC was 66.04% (95% CI 52.56–78.88%). Conclusions: SGNNs provide a feasible graph-based representation of pre-diagnostic proteomic data but did not outperform optimised conventional machine-learning methods within development cross-validation. The temporal and participant-independent performance estimates warrant further investigation in larger fully independent cohorts. A graph-specific predictive advantage was not directly tested—no such advantage was demonstrated in this dataset—and the clinical utility has not been established.

Authors

Institutions

Publication Details

Journal
Diagnostics
Published
2026-10-09
DOI
https://doi.org/10.3390/diagnostics16203283
Primary Topic
Pancreatic and Hepatic Oncology Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Graph Neural Network Modelling of Pre-Diagnostic Serum Proteomics for Pancreatic Cancer Detection

J.G. Oganezova, Alexey A. Zaikin, Oleg B. Blyuss, Arseniy Trukhanov et al.
Diagnostics
Pancreatic and Hepatic Oncology Research
article

Graph Neural Network Modelling of Pre-Diagnostic Serum Proteomics for Pancreatic Cancer Detection

J.G. Oganezova, Alexey A. Zaikin, Oleg B. Blyuss, Arseniy Trukhanov, Aleksandra Gentry‐Maharaj, Harry J. Whitwell, Usha Menon, Sophia Apostolidou
article en

Abstract

Background/Objectives: Pancreatic ductal adenocarcinoma (PDAC) has exceptionally high mortality, largely because of late diagnosis, while established blood biomarkers such as CA19-9 have limited sensitivity for early detection. We evaluated whether modelling serum proteomic measurements as supervised pairwise statistical graphs using synolitic graph neural networks (SGNNs) could provide a useful framework for pre-diagnostic PDAC classification. Methods: We analysed 217 pre-diagnostic serum samples from UKCTOCS, comprising 100 PDAC samples from 75 women and 117 control samples, profiled for 97 proteins. Samples were divided into a development set collected 0–1 year before diagnosis (n = 110) and a temporal holdout set collected 1–2 years before diagnosis (n = 107); 25 PDAC participants contributed samples to both windows. The final SGNN used fold-internal mutual-information selection of 30 proteins, pairwise RBF-SVM-derived edge information, minimum-connected sparsification, and a 10-member GATv2 ensemble per cross-validation fold. Conventional machine-learning comparators were tuned using development data only. Results: The SGNN achieved a mean 5-fold development ROC-AUC of 75.33 ± 10.43%, lower than the tuned Random Forest (81.67%); the other tuned comparators achieved 77.67% (XGBoost), 77.17% (logistic regression), and 73.17% (SVM). Averaging predictions across the five-fold-specific SGNN ensembles yielded a temporal-holdout ROC-AUC of 64.35% (95% CI 52.6–74.6%). In the participant-independent subset of 25 PDAC cases not represented in development and 57 controls, ROC-AUC was 66.04% (95% CI 52.56–78.88%). Conclusions: SGNNs provide a feasible graph-based representation of pre-diagnostic proteomic data but did not outperform optimised conventional machine-learning methods within development cross-validation. The temporal and participant-independent performance estimates warrant further investigation in larger fully independent cohorts. A graph-specific predictive advantage was not directly tested—no such advantage was demonstrated in this dataset—and the clinical utility has not been established.

DiagnosticsVol. 16(20)
National Research University Higher School of Economics (RU), Queen Mary University of London (GB), Sechenov University (RU), Pirogov Russian National Research Medical University (RU), MRC Clinical Trials Unit at UCL (GB), University College London (GB), Imperial College London (GB), N. I. Lobachevsky State University of Nizhny Novgorod (RU)
Openalex Percentile: Top 16%
Pancreatic and Hepatic Oncology Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.