Classification of Synthetic Biosensor Data with Quantum Simulations Using Feature Optimization and Explainable Machine Learning Supported by RFECV

Background/Aim: Accurate detection of trace amounts of heavy metal toxins in environmental samples requires sensitive and reliable diagnostic systems. This study presents a machine learning framework to classify synthetic biosensor signals—generated via quantum simulation—representing five toxin classes: Control, Lead, Mercury, Cadmium, and Arsenic.Methods: Methods: A total of 75 pre-calculated temporal, spectral, entropy, nonlinear, and wavelet features were utilized. To mitigate the risk of data leakage and obtain reliable performance estimates, the Recursive Feature Elimination with Cross-Validation (RFECV) method was integrated into a GroupKFold-based validation framework. The number of features selected across the five outer folds was 47, 73, 55, 47, and 41, respectively, resulting in an average of 52.6 features. Six ensemble learning algorithms were evaluated using a fixed 47-feature structure. Model interpretability was examined via SHAP analysis, while the contributions of feature groups were assessed through ablation and Wilcoxon signed-rank analyses.Results: Results: Configurations with 75 and 41 features were evaluated in a separate direct comparison experiment, yielding average accuracies of 84.06% and 83.73%, respectively. The main evaluation was conducted on the 47-feature configuration. Among the six ensemble learning algorithms, HistGradientBoosting demonstrated the highest performance, achieving an average accuracy of 85.27% and a weighted F1-score of 85.25%. The Random Forest model exhibited the highest discriminative performance, with a Macro-AUC of 0.9799. The Friedman test indicated that performance differences among the algorithms were statistically significant (p = 0.0186), while Nemenyi post-hoc analysis assessed the relative performance differences between the models. The 47-feature configuration was used for the ablation analysis. The most significant drop in performance occurred upon the removal of frequency-domain features, followed by the removal of non-linear features. The Wilcoxon signed-rank test, applied to evaluate the contribution of non-linear features, demonstrated that this group provided a positive trend and contribution to model performance (p = 0.0625).Conclusion: SHAP-based analyses have demonstrated that features such as Hjorth Mobility, D1 Energy, D1 Standard Deviation, Petrosian Fractal Dimension, and Spectral Entropy contribute significantly to the model's decisions. The results indicate that the combined use of GroupKFold-based validation, RFECV, and explainable AI methods facilitates the reliable and interpretable classification of biosensor data.

Authors

Institutions

Publication Details

Journal
Erciyes Üniversitesi Fen Bilimleri Enstitüsü Fen Bilimleri Dergisi
Published
2026-09-16
DOI
https://doi.org/10.65520/erciyesfen.1994336
Primary Topic
Machine Learning in Bioinformatics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Classification of Synthetic Biosensor Data with Quantum Simulations Using Feature Optimization and Explainable Machine Learning Supported by RFECV

Kübra Keser, Halil İbrahim Sarı
Erciyes Üniversitesi Fen Bilimleri Enstitüsü Fen Bilimleri Dergisi
Machine Learning in Bioinformatics
article

Classification of Synthetic Biosensor Data with Quantum Simulations Using Feature Optimization and Explainable Machine Learning Supported by RFECV

Kübra Keser, Halil İbrahim Sarı
article en

Abstract

Background/Aim: Accurate detection of trace amounts of heavy metal toxins in environmental samples requires sensitive and reliable diagnostic systems. This study presents a machine learning framework to classify synthetic biosensor signals—generated via quantum simulation—representing five toxin classes: Control, Lead, Mercury, Cadmium, and Arsenic.Methods: Methods: A total of 75 pre-calculated temporal, spectral, entropy, nonlinear, and wavelet features were utilized. To mitigate the risk of data leakage and obtain reliable performance estimates, the Recursive Feature Elimination with Cross-Validation (RFECV) method was integrated into a GroupKFold-based validation framework. The number of features selected across the five outer folds was 47, 73, 55, 47, and 41, respectively, resulting in an average of 52.6 features. Six ensemble learning algorithms were evaluated using a fixed 47-feature structure. Model interpretability was examined via SHAP analysis, while the contributions of feature groups were assessed through ablation and Wilcoxon signed-rank analyses.Results: Results: Configurations with 75 and 41 features were evaluated in a separate direct comparison experiment, yielding average accuracies of 84.06% and 83.73%, respectively. The main evaluation was conducted on the 47-feature configuration. Among the six ensemble learning algorithms, HistGradientBoosting demonstrated the highest performance, achieving an average accuracy of 85.27% and a weighted F1-score of 85.25%. The Random Forest model exhibited the highest discriminative performance, with a Macro-AUC of 0.9799. The Friedman test indicated that performance differences among the algorithms were statistically significant (p = 0.0186), while Nemenyi post-hoc analysis assessed the relative performance differences between the models. The 47-feature configuration was used for the ablation analysis. The most significant drop in performance occurred upon the removal of frequency-domain features, followed by the removal of non-linear features. The Wilcoxon signed-rank test, applied to evaluate the contribution of non-linear features, demonstrated that this group provided a positive trend and contribution to model performance (p = 0.0625).Conclusion: SHAP-based analyses have demonstrated that features such as Hjorth Mobility, D1 Energy, D1 Standard Deviation, Petrosian Fractal Dimension, and Spectral Entropy contribute significantly to the model's decisions. The results indicate that the combined use of GroupKFold-based validation, RFECV, and explainable AI methods facilitates the reliable and interpretable classification of biosensor data.

Erciyes Üniversitesi Fen Bilimleri Enstitüsü Fen Bilimleri DergisiVol. 42(3)
Kütahya Dumlupınar Üniversitesi (TR)
Reduced inequalities
Openalex Percentile: Top 18%
Machine Learning in Bioinformatics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.