Tree-Based Classification of COVID-19 Using NanoString Whole-Blood Immune-Response Profiles: Comparison of Full-Dataset and LOOCV-Embedded Feature Selection

Background: Whole-blood transcriptomic profiling can capture systemic immune-response alterations associated with COVID-19 and may support host-response-based classification. However, evidence regarding the discriminatory value of targeted immune-gene panels remains limited, and in small, high-dimensional datasets, the timing of feature selection may substantially affect model performance and interpretation. Aim: This study aimed to evaluate whether NanoString Human Immunology Panel profiles could distinguish COVID-19 from healthy-control measurements and to compare full-dataset feature selection (FDFS) with leave-one-out cross-validation (LOOCV)-embedded feature selection (LEFS). Methods: Publicly available E-MTAB-8871 data comprising 579 genes and 32 whole-blood transcriptomic profiles were analyzed. The dataset included 22 longitudinal COVID-19 measurements obtained from three participants and 10 measurements obtained from 10 healthy controls. Elastic Net regularization was used for feature selection. Random Forest, XGBoost, and LightGBM classifiers were evaluated using sample-level LOOCV. Model performance was assessed using threshold-dependent, discrimination, and probability-based metrics. A separate exploratory LightGBM model was analyzed using SHapley Additive exPlanations (SHAP) to characterize feature contributions. Results: FDFS identified a fixed 40-gene set, whereas LEFS selected a mean of 42 genes per fold (range: 40–47). LightGBM correctly classified all 32 measurement-level profiles (derived from 13 unique participants: 10 healthy controls and three longitudinally sampled COVID-19 participants) in both frameworks, achieving area under the receiver operating characteristic curve (ROC-AUC) and area under the precision–recall curve (PR-AUC) values of 1.000 and Brier scores of 0.005 and 0.006 in the FDFS and LEFS frameworks, respectively. Random Forest achieved accuracies of 0.969 and 1.000, whereas XGBoost achieved an accuracy of 0.969 in both frameworks. SHAP analyses consistently identified AICDA as the dominant contributor to model predictions, followed by ARHGDIB. Conclusions: This exploratory analysis showed that targeted NanoString immune-response profiles contained a compact transcriptomic signal capable of distinguishing COVID-19 from healthy-control measurements within the analyzed dataset. These findings provide proof-of-concept evidence of internal measurement-level discrimination. However, because the COVID-19 profiles consisted of repeated measurements from only three participants, sample-level LOOCV did not constitute independent participant-level validation. External validation in larger cohorts comprising independently sampled participants is required. Given that the COVID-19 arm comprised only three independent participants, these biological findings should be regarded as hypothesis-generating and require validation in substantially larger independent cohorts.

Authors

Institutions

Publication Details

Journal
Viruses
Published
2026-09-13
DOI
https://doi.org/10.3390/v18091009
Primary Topic
COVID-19 Clinical Research Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Tree-Based Classification of COVID-19 Using NanoString Whole-Blood Immune-Response Profiles: Comparison of Full-Dataset and LOOCV-Embedded Feature Selection

Sami Akbulut, Zeynep Küçükakçalı, Lecturer Zeynep Burçin Yılmaz
Viruses
COVID-19 Clinical Research Studies
article

Tree-Based Classification of COVID-19 Using NanoString Whole-Blood Immune-Response Profiles: Comparison of Full-Dataset and LOOCV-Embedded Feature Selection

Sami Akbulut, Zeynep Küçükakçalı, Lecturer Zeynep Burçin Yılmaz
article en

Abstract

Background: Whole-blood transcriptomic profiling can capture systemic immune-response alterations associated with COVID-19 and may support host-response-based classification. However, evidence regarding the discriminatory value of targeted immune-gene panels remains limited, and in small, high-dimensional datasets, the timing of feature selection may substantially affect model performance and interpretation. Aim: This study aimed to evaluate whether NanoString Human Immunology Panel profiles could distinguish COVID-19 from healthy-control measurements and to compare full-dataset feature selection (FDFS) with leave-one-out cross-validation (LOOCV)-embedded feature selection (LEFS). Methods: Publicly available E-MTAB-8871 data comprising 579 genes and 32 whole-blood transcriptomic profiles were analyzed. The dataset included 22 longitudinal COVID-19 measurements obtained from three participants and 10 measurements obtained from 10 healthy controls. Elastic Net regularization was used for feature selection. Random Forest, XGBoost, and LightGBM classifiers were evaluated using sample-level LOOCV. Model performance was assessed using threshold-dependent, discrimination, and probability-based metrics. A separate exploratory LightGBM model was analyzed using SHapley Additive exPlanations (SHAP) to characterize feature contributions. Results: FDFS identified a fixed 40-gene set, whereas LEFS selected a mean of 42 genes per fold (range: 40–47). LightGBM correctly classified all 32 measurement-level profiles (derived from 13 unique participants: 10 healthy controls and three longitudinally sampled COVID-19 participants) in both frameworks, achieving area under the receiver operating characteristic curve (ROC-AUC) and area under the precision–recall curve (PR-AUC) values of 1.000 and Brier scores of 0.005 and 0.006 in the FDFS and LEFS frameworks, respectively. Random Forest achieved accuracies of 0.969 and 1.000, whereas XGBoost achieved an accuracy of 0.969 in both frameworks. SHAP analyses consistently identified AICDA as the dominant contributor to model predictions, followed by ARHGDIB. Conclusions: This exploratory analysis showed that targeted NanoString immune-response profiles contained a compact transcriptomic signal capable of distinguishing COVID-19 from healthy-control measurements within the analyzed dataset. These findings provide proof-of-concept evidence of internal measurement-level discrimination. However, because the COVID-19 profiles consisted of repeated measurements from only three participants, sample-level LOOCV did not constitute independent participant-level validation. External validation in larger cohorts comprising independently sampled participants is required. Given that the COVID-19 arm comprised only three independent participants, these biological findings should be regarded as hypothesis-generating and require validation in substantially larger independent cohorts.

VirusesVol. 18(9)
Inonu University (TR)
Reduced inequalities
Openalex Percentile: Top 11%
COVID-19 Clinical Research Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.