Tree-Based Classification of COVID-19 Using NanoString Whole-Blood Immune-Response Profiles: Comparison of Full-Dataset and LOOCV-Embedded Feature Selection
Background: Whole-blood transcriptomic profiling can capture systemic immune-response alterations associated with COVID-19 and may support host-response-based classification. However, evidence regarding the discriminatory value of targeted immune-gene panels remains limited, and in small, high-dimensional datasets, the timing of feature selection may substantially affect model performance and interpretation. Aim: This study aimed to evaluate whether NanoString Human Immunology Panel profiles could distinguish COVID-19 from healthy-control measurements and to compare full-dataset feature selection (FDFS) with leave-one-out cross-validation (LOOCV)-embedded feature selection (LEFS). Methods: Publicly available E-MTAB-8871 data comprising 579 genes and 32 whole-blood transcriptomic profiles were analyzed. The dataset included 22 longitudinal COVID-19 measurements obtained from three participants and 10 measurements obtained from 10 healthy controls. Elastic Net regularization was used for feature selection. Random Forest, XGBoost, and LightGBM classifiers were evaluated using sample-level LOOCV. Model performance was assessed using threshold-dependent, discrimination, and probability-based metrics. A separate exploratory LightGBM model was analyzed using SHapley Additive exPlanations (SHAP) to characterize feature contributions. Results: FDFS identified a fixed 40-gene set, whereas LEFS selected a mean of 42 genes per fold (range: 40–47). LightGBM correctly classified all 32 measurement-level profiles (derived from 13 unique participants: 10 healthy controls and three longitudinally sampled COVID-19 participants) in both frameworks, achieving area under the receiver operating characteristic curve (ROC-AUC) and area under the precision–recall curve (PR-AUC) values of 1.000 and Brier scores of 0.005 and 0.006 in the FDFS and LEFS frameworks, respectively. Random Forest achieved accuracies of 0.969 and 1.000, whereas XGBoost achieved an accuracy of 0.969 in both frameworks. SHAP analyses consistently identified AICDA as the dominant contributor to model predictions, followed by ARHGDIB. Conclusions: This exploratory analysis showed that targeted NanoString immune-response profiles contained a compact transcriptomic signal capable of distinguishing COVID-19 from healthy-control measurements within the analyzed dataset. These findings provide proof-of-concept evidence of internal measurement-level discrimination. However, because the COVID-19 profiles consisted of repeated measurements from only three participants, sample-level LOOCV did not constitute independent participant-level validation. External validation in larger cohorts comprising independently sampled participants is required. Given that the COVID-19 arm comprised only three independent participants, these biological findings should be regarded as hypothesis-generating and require validation in substantially larger independent cohorts.
Authors
- Sami Akbulut (ORCID: https://orcid.org/0000-0002-6864-7711)
- Zeynep Küçükakçalı (ORCID: https://orcid.org/0000-0001-7956-9272)
- Lecturer Zeynep Burçin Yılmaz (ORCID: https://orcid.org/0000-0002-6950-6013)
Institutions
- Inonu University (TR)
Publication Details
- Journal
- Viruses
- Published
- 2026-09-13
- DOI
- https://doi.org/10.3390/v18091009
- Primary Topic
- COVID-19 Clinical Research Studies
- Type
- article
- Field-Weighted Citation Impact
- 0.00