Machine Learning Classification of Heart-Disease Status Among COVID-19-Positive Patients Using Clinical and Polygenic Risk Score Data

The clinical characteristics of COVID-19-positive patients are heterogeneous, and machine learning approaches may help identify patterns associated with recorded cardiovascular comorbidity. However, class imbalance, small sample sizes, heterogeneous predictors, and potential information leakage can limit the reliability and generalizability of predictive models. Methods: We evaluated Logistic Regression (LR), Random Forest (RF), and Extreme Gradient Boosting (XGBoost) for classification of recorded heart-disease status among COVID-19-positive patients. The clinical dataset initially comprised 572 patient records collected at Yenepoya Medical College Hospital, of which 468 records were retained following data cleaning. The dataset included demographic characteristics, RT-PCR cycle-threshold measurements, IgM/IgG serology, symptoms, comorbidities, and available polygenic risk score information. The models were trained using an 80:20 stratified train–test split with five-fold stratified cross-validation on the training set. Class imbalance was addressed using balanced class weights for LR and RF and scale_pos_weight for XGBoost. Model performance was assessed using accuracy, precision, recall, F1-score, ROC-AUC, and PR-AUC. A separate GBRC genomic dataset containing 804,522 SNVs from 459 COVID-19-positive individuals was not linked to the clinical cohort and was therefore not incorporated as a predictor. Results: On the independent test set, LR achieved an accuracy of 83%, ROC-AUC of 87%, and PR-AUC of 60%, with positive-class recall and F1-score of 69% and 53%, respectively. RF achieved an accuracy of 90%, ROC-AUC of 92%, and PR-AUC of 72%, although positive-class recall was 31%. XGBoost achieved an accuracy of 89%, ROC-AUC of 89%, and PR-AUC of 68%, with positive-class recall and F1-score of 54% and 58%, respectively. The results demonstrate substantial differences between overall discrimination and minority-class detection across models. Conclusions: The findings provide a preliminary evaluation of supervised machine learning approaches for classification of recorded heart-disease status in a relatively small COVID-19-positive cohort. The limited number of positive observations, lack of external validation, and potential sensitivity to preprocessing and predictor definition constrain interpretation. Larger independently collected cohorts and leakage-resistant validation are required before clinical applicability can be assessed.

Authors

Institutions

Publication Details

Journal
COVID
Published
2026-10-09
DOI
https://doi.org/10.3390/covid6100180
Primary Topic
Artificial Intelligence in Healthcare
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Machine Learning Classification of Heart-Disease Status Among COVID-19-Positive Patients Using Clinical and Polygenic Risk Score Data

Parameshwar R. Hegde, Diptiman Choudhury, Ranajit Das, Sangita Roy et al.
COVID
Artificial Intelligence in Healthcare
article

Machine Learning Classification of Heart-Disease Status Among COVID-19-Positive Patients Using Clinical and Polygenic Risk Score Data

Parameshwar R. Hegde, Diptiman Choudhury, Ranajit Das, Sangita Roy, Deepthi Vikram, Shivank Bhatia
article en

Abstract

The clinical characteristics of COVID-19-positive patients are heterogeneous, and machine learning approaches may help identify patterns associated with recorded cardiovascular comorbidity. However, class imbalance, small sample sizes, heterogeneous predictors, and potential information leakage can limit the reliability and generalizability of predictive models. Methods: We evaluated Logistic Regression (LR), Random Forest (RF), and Extreme Gradient Boosting (XGBoost) for classification of recorded heart-disease status among COVID-19-positive patients. The clinical dataset initially comprised 572 patient records collected at Yenepoya Medical College Hospital, of which 468 records were retained following data cleaning. The dataset included demographic characteristics, RT-PCR cycle-threshold measurements, IgM/IgG serology, symptoms, comorbidities, and available polygenic risk score information. The models were trained using an 80:20 stratified train–test split with five-fold stratified cross-validation on the training set. Class imbalance was addressed using balanced class weights for LR and RF and scale_pos_weight for XGBoost. Model performance was assessed using accuracy, precision, recall, F1-score, ROC-AUC, and PR-AUC. A separate GBRC genomic dataset containing 804,522 SNVs from 459 COVID-19-positive individuals was not linked to the clinical cohort and was therefore not incorporated as a predictor. Results: On the independent test set, LR achieved an accuracy of 83%, ROC-AUC of 87%, and PR-AUC of 60%, with positive-class recall and F1-score of 69% and 53%, respectively. RF achieved an accuracy of 90%, ROC-AUC of 92%, and PR-AUC of 72%, although positive-class recall was 31%. XGBoost achieved an accuracy of 89%, ROC-AUC of 89%, and PR-AUC of 68%, with positive-class recall and F1-score of 54% and 58%, respectively. The results demonstrate substantial differences between overall discrimination and minority-class detection across models. Conclusions: The findings provide a preliminary evaluation of supervised machine learning approaches for classification of recorded heart-disease status in a relatively small COVID-19-positive cohort. The limited number of positive observations, lack of external validation, and potential sensitivity to preprocessing and predictor definition constrain interpretation. Larger independently collected cohorts and leakage-resistant validation are required before clinical applicability can be assessed.

COVIDVol. 6(10)
Yenepoya University (IN), Thapar Institute of Engineering & Technology (IN)
Openalex Percentile: Top 8%
Artificial Intelligence in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.