Enhancing the performance of machine learning models for the reliable diagnosis of hepatitis C virus

Early and reliable detection of hepatitis C virus infection is essential for timely treatment and for reducing the risk of progressive liver damage and long-term complications. This study presents a methodological benchmarking framework to systematically evaluate the performance and stability of machine learning classifiers for HCV detection using routine biochemical attributes from a public benchmark dataset. A diagnostic framework was developed to address the problems associated with real-world medical datasets, particularly the problem of class imbalance and outliers, which can reduce model generalization. Feature importance was assessed using a decision tree-based recursive feature elimination approach to retain only the necessary predictors. Multiple machine learning classifiers including k-Nearest Neighbors, support vector machine, logistic regression, random forest, naive bayes, gradient boosting, extreme gradient boosting, and LightGBM, were trained and evaluated under a consistent experimental setting. The parameters for various models were optimized prior to testing using hyperparameter tuning. Performance was evaluated on an independent test cohort, and model generalizability was further verified using a stratified 5-fold cross-validation protocol. On the independent test cohort, the random forest model achieved the best diagnostic performance (98.38% accuracy, 100% precision, 93.02% F1-score). The subsequent 5-fold cross-validation confirmed high framework stability, with LightGBM delivering a top mean accuracy of \(95.77 \pm 1.45\%\) and an AUC of \(0.97 \pm 0.02\) . Other classifiers, especially support vector machine and logistic regression, also showed stable and competitive results, suggesting consistent generalization on unseen data. When compared to previous studies, the proposed models showed improved diagnostic sensitivity and stronger overall robustness, particularly when handling data imbalance and outliers. These findings confirm the value of data balancing, structured feature selection and careful model optimization for enhancing HCV diagnosis from clinical data. This benchmarking study demonstrates that integrating structured data balancing, feature elimination, and hyperparameter tuning significantly stabilizes classifier performance on routine laboratory data, establishing a reliable baseline methodology for future computer-aided screening research.

Authors

Institutions

Publication Details

Journal
BMC Infectious Diseases
Published
2026-10-07
DOI
https://doi.org/10.1186/s12879-026-14179-5
Primary Topic
Artificial Intelligence in Healthcare
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Enhancing the performance of machine learning models for the reliable diagnosis of hepatitis C virus

Dawar Awan, Fouzia Idrees, Muhammad Uzair Khan, Fazal Muhammad et al.
BMC Infectious Diseases
Artificial Intelligence in Healthcare
article

Enhancing the performance of machine learning models for the reliable diagnosis of hepatitis C virus

Dawar Awan, Fouzia Idrees, Muhammad Uzair Khan, Fazal Muhammad, Shahid Khan, Amal Al-Rasheed, Jalal Khan
article en

Abstract

Early and reliable detection of hepatitis C virus infection is essential for timely treatment and for reducing the risk of progressive liver damage and long-term complications. This study presents a methodological benchmarking framework to systematically evaluate the performance and stability of machine learning classifiers for HCV detection using routine biochemical attributes from a public benchmark dataset. A diagnostic framework was developed to address the problems associated with real-world medical datasets, particularly the problem of class imbalance and outliers, which can reduce model generalization. Feature importance was assessed using a decision tree-based recursive feature elimination approach to retain only the necessary predictors. Multiple machine learning classifiers including k-Nearest Neighbors, support vector machine, logistic regression, random forest, naive bayes, gradient boosting, extreme gradient boosting, and LightGBM, were trained and evaluated under a consistent experimental setting. The parameters for various models were optimized prior to testing using hyperparameter tuning. Performance was evaluated on an independent test cohort, and model generalizability was further verified using a stratified 5-fold cross-validation protocol. On the independent test cohort, the random forest model achieved the best diagnostic performance (98.38% accuracy, 100% precision, 93.02% F1-score). The subsequent 5-fold cross-validation confirmed high framework stability, with LightGBM delivering a top mean accuracy of \(95.77 \pm 1.45\%\) and an AUC of \(0.97 \pm 0.02\) . Other classifiers, especially support vector machine and logistic regression, also showed stable and competitive results, suggesting consistent generalization on unseen data. When compared to previous studies, the proposed models showed improved diagnostic sensitivity and stronger overall robustness, particularly when handling data imbalance and outliers. These findings confirm the value of data balancing, structured feature selection and careful model optimization for enhancing HCV diagnosis from clinical data. This benchmarking study demonstrates that integrating structured data balancing, feature elimination, and hyperparameter tuning significantly stabilizes classifier performance on routine laboratory data, establishing a reliable baseline methodology for future computer-aided screening research.

BMC Infectious Diseases
Princess Nourah bint Abdulrahman University (SA), Shaheed Benazir Bhutto Women University Peshawar (PK), COMSATS University Islamabad (PK), Abdul Wali Khan University Mardan (PK), Northern University (PK)
Openalex Percentile: Top 7%
Artificial Intelligence in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.