Enhancing the performance of machine learning models for the reliable diagnosis of hepatitis C virus
Early and reliable detection of hepatitis C virus infection is essential for timely treatment and for reducing the risk of progressive liver damage and long-term complications. This study presents a methodological benchmarking framework to systematically evaluate the performance and stability of machine learning classifiers for HCV detection using routine biochemical attributes from a public benchmark dataset. A diagnostic framework was developed to address the problems associated with real-world medical datasets, particularly the problem of class imbalance and outliers, which can reduce model generalization. Feature importance was assessed using a decision tree-based recursive feature elimination approach to retain only the necessary predictors. Multiple machine learning classifiers including k-Nearest Neighbors, support vector machine, logistic regression, random forest, naive bayes, gradient boosting, extreme gradient boosting, and LightGBM, were trained and evaluated under a consistent experimental setting. The parameters for various models were optimized prior to testing using hyperparameter tuning. Performance was evaluated on an independent test cohort, and model generalizability was further verified using a stratified 5-fold cross-validation protocol. On the independent test cohort, the random forest model achieved the best diagnostic performance (98.38% accuracy, 100% precision, 93.02% F1-score). The subsequent 5-fold cross-validation confirmed high framework stability, with LightGBM delivering a top mean accuracy of \(95.77 \pm 1.45\%\) and an AUC of \(0.97 \pm 0.02\) . Other classifiers, especially support vector machine and logistic regression, also showed stable and competitive results, suggesting consistent generalization on unseen data. When compared to previous studies, the proposed models showed improved diagnostic sensitivity and stronger overall robustness, particularly when handling data imbalance and outliers. These findings confirm the value of data balancing, structured feature selection and careful model optimization for enhancing HCV diagnosis from clinical data. This benchmarking study demonstrates that integrating structured data balancing, feature elimination, and hyperparameter tuning significantly stabilizes classifier performance on routine laboratory data, establishing a reliable baseline methodology for future computer-aided screening research.
Authors
- Dawar Awan (ORCID: https://orcid.org/0009-0005-8843-259X)
- Fouzia Idrees (ORCID: https://orcid.org/0000-0002-5936-1842)
- Muhammad Uzair Khan (ORCID: https://orcid.org/0000-0002-0557-2900)
- Fazal Muhammad (ORCID: https://orcid.org/0000-0003-0405-0083)
- Shahid Khan
- Amal Al-Rasheed
- Jalal Khan
Institutions
- Princess Nourah bint Abdulrahman University (SA)
- Shaheed Benazir Bhutto Women University Peshawar (PK)
- COMSATS University Islamabad (PK)
- Abdul Wali Khan University Mardan (PK)
- Northern University (PK)
Publication Details
- Journal
- BMC Infectious Diseases
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1186/s12879-026-14179-5
- Primary Topic
- Artificial Intelligence in Healthcare
- Type
- article
- Field-Weighted Citation Impact
- 0.00