Machine Learning Classification of Heart-Disease Status Among COVID-19-Positive Patients Using Clinical and Polygenic Risk Score Data
The clinical characteristics of COVID-19-positive patients are heterogeneous, and machine learning approaches may help identify patterns associated with recorded cardiovascular comorbidity. However, class imbalance, small sample sizes, heterogeneous predictors, and potential information leakage can limit the reliability and generalizability of predictive models. Methods: We evaluated Logistic Regression (LR), Random Forest (RF), and Extreme Gradient Boosting (XGBoost) for classification of recorded heart-disease status among COVID-19-positive patients. The clinical dataset initially comprised 572 patient records collected at Yenepoya Medical College Hospital, of which 468 records were retained following data cleaning. The dataset included demographic characteristics, RT-PCR cycle-threshold measurements, IgM/IgG serology, symptoms, comorbidities, and available polygenic risk score information. The models were trained using an 80:20 stratified train–test split with five-fold stratified cross-validation on the training set. Class imbalance was addressed using balanced class weights for LR and RF and scale_pos_weight for XGBoost. Model performance was assessed using accuracy, precision, recall, F1-score, ROC-AUC, and PR-AUC. A separate GBRC genomic dataset containing 804,522 SNVs from 459 COVID-19-positive individuals was not linked to the clinical cohort and was therefore not incorporated as a predictor. Results: On the independent test set, LR achieved an accuracy of 83%, ROC-AUC of 87%, and PR-AUC of 60%, with positive-class recall and F1-score of 69% and 53%, respectively. RF achieved an accuracy of 90%, ROC-AUC of 92%, and PR-AUC of 72%, although positive-class recall was 31%. XGBoost achieved an accuracy of 89%, ROC-AUC of 89%, and PR-AUC of 68%, with positive-class recall and F1-score of 54% and 58%, respectively. The results demonstrate substantial differences between overall discrimination and minority-class detection across models. Conclusions: The findings provide a preliminary evaluation of supervised machine learning approaches for classification of recorded heart-disease status in a relatively small COVID-19-positive cohort. The limited number of positive observations, lack of external validation, and potential sensitivity to preprocessing and predictor definition constrain interpretation. Larger independently collected cohorts and leakage-resistant validation are required before clinical applicability can be assessed.
Authors
- Parameshwar R. Hegde (ORCID: https://orcid.org/0000-0003-0689-7667)
- Diptiman Choudhury (ORCID: https://orcid.org/0000-0003-1080-4558)
- Ranajit Das (ORCID: https://orcid.org/0000-0001-5308-3477)
- Sangita Roy (ORCID: https://orcid.org/0000-0002-8898-0183)
- Deepthi Vikram (ORCID: https://orcid.org/0009-0001-6809-8919)
- Shivank Bhatia
Institutions
- Yenepoya University (IN)
- Thapar Institute of Engineering & Technology (IN)
Publication Details
- Journal
- COVID
- Published
- 2026-10-09
- DOI
- https://doi.org/10.3390/covid6100180
- Primary Topic
- Artificial Intelligence in Healthcare
- Type
- article
- Field-Weighted Citation Impact
- 0.00