When Accuracy Is Not Enough: A Validation Framework for Machine Learning in Early Parkinson's Disease Research

Abstract Machine learning studies of Parkinson's disease often produce a single measure of predictive accuracy, although their clinical questions differ substantially. Recognizing established disease in a selected dataset, estimating future diagnosis among adults without Parkinson's disease, and tracking change within a known patient require different evaluation designs. This methodological perspective proposes a practical framework for separating those questions and making early-detection claims auditable. A limited narrative selection of primary studies and original methodological guidance informed the framework; no systematic review or pooled performance estimate was undertaken. The proposed evaluation sequence begins with an explicit target population, prediction time, outcome horizon, and intended research decision. It then links participant separation, temporal ordering, site independence, preprocessing isolation, calibration, and decision analysis to that specification. Two hypothetical calculations illustrate why high sensitivity and specificity can coexist with a low positive predictive value, and why a seemingly useful classifier may offer limited net benefit at a particular referral threshold. A compact evidence table connects common analytical failures to the artifacts required to investigate them. For NeuralCipher, the framework defines future validation work and does not establish performance of the current platform. The central recommendation is to publish a reproducible chain from the research question to the denominator of every reported estimate. Independent validation should assess whether that chain survives new participants, different settings, incomplete measurements, and realistic outcome frequencies before any clinical early-detection claim is considered. Article details Authors: NeuralCipherai; Kadir Tamrak; Salih Yaldız; Feride Yaldız; Ömer Ağyol; Yavuz Selim Silay; Hasan Randa Publisher: neluracipher.ai DOI: 10.5281/zenodo.22779034 Version: 1.0Language: English Project website: https://neuralcipher.ai References Bossuyt, P. M., Reitsma, J. B., Bruns, D. E., Gatsonis, C. A., Glasziou, P. P., Irwig, L., Lijmer, J. G., Moher, D., Rennie, D., de Vet, H. C. W., Kressel, H. Y., Rifai, N., Golub, R. M., Altman, D. G., Hooft, L., Korevaar, D. A., Cohen, J. F., & for the STARD Group. (2015). STARD 2015: An updated list of essential items for reporting diagnostic accuracy studies. BMJ, 351, Article h5527. https://doi.org/10.1136/bmj.h5527 Bot, B. M., Suver, C., Neto, E. C., Kellen, M., Klein, A., Bare, C., Doerr, M., Pratap, A., Wilbanks, J., Dorsey, E. R., Friend, S. H., & Trister, A. D. (2016). The mPower study, Parkinson disease mobile data collected using ResearchKit. Scientific Data, 3(1), Article 160011. https://doi.org/10.1038/sdata.2016.11 Collins, G. S., Moons, K. G. M., Dhiman, P., Riley, R. D., Beam, A. L., Van Calster, B., Ghassemi, M., Liu, X., Reitsma, J. B., van Smeden, M., Boulesteix, A.-L., Camaradou, J. C., Celi, L. A., Denaxas, S., Denniston, A. K., Glocker, B., Golub, R. M., Harvey, H., Heinze, G., . . . Logullo, P. (2024). TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ, 385, Article e078378. https://doi.org/10.1136/bmj-2023-078378 Heinzel, S., Berg, D., Gasser, T., Chen, H., Yao, C., Postuma, R. B., & the MDS Task Force on the Definition of Parkinson's Disease. (2019). Update of the MDS research criteria for prodromal Parkinson's disease. Movement Disorders, 34(10), 1464–1470. https://doi.org/10.1002/mds.27802 Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), Article 100804. https://doi.org/10.1016/j.patter.2023.100804 Moons, K. G. M., Damen, J. A. A., Kaul, T., Hooft, L., Andaur Navarro, C., Dhiman, P., Beam, A. L., Van Calster, B., Celi, L. A., Denaxas, S., Denniston, A. K., Ghassemi, M., Heinze, G., Kengne, A. P., Maier-Hein, L., Liu, X., Logullo, P., McCradden, M. D., Liu, N., . . . van Smeden, M. (2025). PROBAST+AI: An updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ, 388, Article e082505. https://doi.org/10.1136/bmj-2024-082505 Postuma, R. B., Berg, D., Stern, M., Poewe, W., Olanow, C. W., Oertel, W., Obeso, J., Marek, K., Litvan, I., Lang, A. E., Halliday, G., Goetz, C. G., Gasser, T., Dubois, B., Chan, P., Bloem, B. R., Adler, C. H., & Deuschl, G. (2015). MDS clinical diagnostic criteria for Parkinson's disease. Movement Disorders, 30(12), 1591–1601. https://doi.org/10.1002/mds.26424 Riley, R. D., Debray, T. P. A., Collins, G. S., Archer, L., Ensor, J., van Smeden, M., & Snell, K. I. E. (2021). Minimum sample size for external validation of a clinical prediction model with a binary outcome. Statistics in Medicine, 40(19), 4230–4251. https://doi.org/10.1002/sim.9025 Saeb, S., Lonini, L., Jayaraman, A., Mohr, D. C., & Kording, K. P. (2017). The need to approximate the use-case in clinical machine learning. GigaScience, 6(5), Article gix019. https://doi.org/10.1093/gigascience/gix019 Schalkamp, A.-K., Peall, K. J., Harrison, N. A., & Sandor, C. (2023). Wearable movement-tracking data identify Parkinson’s disease years before clinical diagnosis. Nature Medicine, 29(8), 2048–2056. https://doi.org/10.1038/s41591-023-02440-2 Van Calster, B., McLernon, D. J., van Smeden, M., Wynants, L., Steyerberg, E. W., & On behalf of Topic Group ‘Evaluating diagnostic tests and prediction models’ of the STRATOS initiative. (2019). Calibration: The Achilles heel of predictive analytics. BMC Medicine, 17(1), Article 230. https://doi.org/10.1186/s12916-019-1466-7 Varoquaux, G. (2018). Cross-validation failure: Small sample sizes lead to large error bars. NeuroImage, 180, 68–77. https://doi.org/10.1016/j.neuroimage.2017.06.061 Vickers, A. J., & Elkin, E. B. (2006). Decision curve analysis: A novel method for evaluating prediction models. Medical Decision Making, 26(6), 565–574. https://doi.org/10.1177/0272989x06295361 von Elm, E., Altman, D. G., Egger, M., Pocock, S. J., Gøtzsche, P. C., Vandenbroucke, J. P., & for the STROBE Initiative. (2007). The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: Guidelines for reporting observational studies. PLoS Medicine, 4(10), Article e296. https://doi.org/10.1371/journal.pmed.0040296 Yang, Y., Yuan, Y., Zhang, G., Wang, H., Chen, Y.-C., Liu, Y., Tarolli, C. G., Crepeau, D., Bukartyk, J., Junna, M. R., Videnovic, A., Ellis, T. D., Lipford, M. C., Dorsey, R., & Katabi, D. (2022). Artificial intelligence-enabled detection and assessment of Parkinson’s disease using nocturnal breathing signals. Nature Medicine, 28(10), 2207–2215. https://doi.org/10.1038/s41591-022-01932-x

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-16
DOI
https://doi.org/10.5281/zenodo.22779034
Primary Topic
Parkinson's Disease Mechanisms and Treatments
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

When Accuracy Is Not Enough: A Validation Framework for Machine Learning in Early Parkinson's Disease Research

Feride Yaldiz, Yavuz Selim Sılay, Kadir Tamrak, Salih Yaldız et al.
Zenodo (CERN European Organization for Nuclear Research)
Parkinson's Disease Mechanisms and Treatments
preprint

When Accuracy Is Not Enough: A Validation Framework for Machine Learning in Early Parkinson's Disease Research

Feride Yaldiz, Yavuz Selim Sılay, Kadir Tamrak, Salih Yaldız, NeuralCipherai, Hasan Randa, Ömer Ağyol
preprint en

Abstract

Abstract Machine learning studies of Parkinson's disease often produce a single measure of predictive accuracy, although their clinical questions differ substantially. Recognizing established disease in a selected dataset, estimating future diagnosis among adults without Parkinson's disease, and tracking change within a known patient require different evaluation designs. This methodological perspective proposes a practical framework for separating those questions and making early-detection claims auditable. A limited narrative selection of primary studies and original methodological guidance informed the framework; no systematic review or pooled performance estimate was undertaken. The proposed evaluation sequence begins with an explicit target population, prediction time, outcome horizon, and intended research decision. It then links participant separation, temporal ordering, site independence, preprocessing isolation, calibration, and decision analysis to that specification. Two hypothetical calculations illustrate why high sensitivity and specificity can coexist with a low positive predictive value, and why a seemingly useful classifier may offer limited net benefit at a particular referral threshold. A compact evidence table connects common analytical failures to the artifacts required to investigate them. For NeuralCipher, the framework defines future validation work and does not establish performance of the current platform. The central recommendation is to publish a reproducible chain from the research question to the denominator of every reported estimate. Independent validation should assess whether that chain survives new participants, different settings, incomplete measurements, and realistic outcome frequencies before any clinical early-detection claim is considered. Article details Authors: NeuralCipherai; Kadir Tamrak; Salih Yaldız; Feride Yaldız; Ömer Ağyol; Yavuz Selim Silay; Hasan Randa Publisher: neluracipher.ai DOI: 10.5281/zenodo.22779034 Version: 1.0Language: English Project website: https://neuralcipher.ai References Bossuyt, P. M., Reitsma, J. B., Bruns, D. E., Gatsonis, C. A., Glasziou, P. P., Irwig, L., Lijmer, J. G., Moher, D., Rennie, D., de Vet, H. C. W., Kressel, H. Y., Rifai, N., Golub, R. M., Altman, D. G., Hooft, L., Korevaar, D. A., Cohen, J. F., & for the STARD Group. (2015). STARD 2015: An updated list of essential items for reporting diagnostic accuracy studies. BMJ, 351, Article h5527. https://doi.org/10.1136/bmj.h5527 Bot, B. M., Suver, C., Neto, E. C., Kellen, M., Klein, A., Bare, C., Doerr, M., Pratap, A., Wilbanks, J., Dorsey, E. R., Friend, S. H., & Trister, A. D. (2016). The mPower study, Parkinson disease mobile data collected using ResearchKit. Scientific Data, 3(1), Article 160011. https://doi.org/10.1038/sdata.2016.11 Collins, G. S., Moons, K. G. M., Dhiman, P., Riley, R. D., Beam, A. L., Van Calster, B., Ghassemi, M., Liu, X., Reitsma, J. B., van Smeden, M., Boulesteix, A.-L., Camaradou, J. C., Celi, L. A., Denaxas, S., Denniston, A. K., Glocker, B., Golub, R. M., Harvey, H., Heinze, G., . . . Logullo, P. (2024). TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ, 385, Article e078378. https://doi.org/10.1136/bmj-2023-078378 Heinzel, S., Berg, D., Gasser, T., Chen, H., Yao, C., Postuma, R. B., & the MDS Task Force on the Definition of Parkinson's Disease. (2019). Update of the MDS research criteria for prodromal Parkinson's disease. Movement Disorders, 34(10), 1464–1470. https://doi.org/10.1002/mds.27802 Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), Article 100804. https://doi.org/10.1016/j.patter.2023.100804 Moons, K. G. M., Damen, J. A. A., Kaul, T., Hooft, L., Andaur Navarro, C., Dhiman, P., Beam, A. L., Van Calster, B., Celi, L. A., Denaxas, S., Denniston, A. K., Ghassemi, M., Heinze, G., Kengne, A. P., Maier-Hein, L., Liu, X., Logullo, P., McCradden, M. D., Liu, N., . . . van Smeden, M. (2025). PROBAST+AI: An updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ, 388, Article e082505. https://doi.org/10.1136/bmj-2024-082505 Postuma, R. B., Berg, D., Stern, M., Poewe, W., Olanow, C. W., Oertel, W., Obeso, J., Marek, K., Litvan, I., Lang, A. E., Halliday, G., Goetz, C. G., Gasser, T., Dubois, B., Chan, P., Bloem, B. R., Adler, C. H., & Deuschl, G. (2015). MDS clinical diagnostic criteria for Parkinson's disease. Movement Disorders, 30(12), 1591–1601. https://doi.org/10.1002/mds.26424 Riley, R. D., Debray, T. P. A., Collins, G. S., Archer, L., Ensor, J., van Smeden, M., & Snell, K. I. E. (2021). Minimum sample size for external validation of a clinical prediction model with a binary outcome. Statistics in Medicine, 40(19), 4230–4251. https://doi.org/10.1002/sim.9025 Saeb, S., Lonini, L., Jayaraman, A., Mohr, D. C., & Kording, K. P. (2017). The need to approximate the use-case in clinical machine learning. GigaScience, 6(5), Article gix019. https://doi.org/10.1093/gigascience/gix019 Schalkamp, A.-K., Peall, K. J., Harrison, N. A., & Sandor, C. (2023). Wearable movement-tracking data identify Parkinson’s disease years before clinical diagnosis. Nature Medicine, 29(8), 2048–2056. https://doi.org/10.1038/s41591-023-02440-2 Van Calster, B., McLernon, D. J., van Smeden, M., Wynants, L., Steyerberg, E. W., & On behalf of Topic Group ‘Evaluating diagnostic tests and prediction models’ of the STRATOS initiative. (2019). Calibration: The Achilles heel of predictive analytics. BMC Medicine, 17(1), Article 230. https://doi.org/10.1186/s12916-019-1466-7 Varoquaux, G. (2018). Cross-validation failure: Small sample sizes lead to large error bars. NeuroImage, 180, 68–77. https://doi.org/10.1016/j.neuroimage.2017.06.061 Vickers, A. J., & Elkin, E. B. (2006). Decision curve analysis: A novel method for evaluating prediction models. Medical Decision Making, 26(6), 565–574. https://doi.org/10.1177/0272989x06295361 von Elm, E., Altman, D. G., Egger, M., Pocock, S. J., Gøtzsche, P. C., Vandenbroucke, J. P., & for the STROBE Initiative. (2007). The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: Guidelines for reporting observational studies. PLoS Medicine, 4(10), Article e296. https://doi.org/10.1371/journal.pmed.0040296 Yang, Y., Yuan, Y., Zhang, G., Wang, H., Chen, Y.-C., Liu, Y., Tarolli, C. G., Crepeau, D., Bukartyk, J., Junna, M. R., Videnovic, A., Ellis, T. D., Lipford, M. C., Dorsey, R., & Katabi, D. (2022). Artificial intelligence-enabled detection and assessment of Parkinson’s disease using nocturnal breathing signals. Nature Medicine, 28(10), 2207–2215. https://doi.org/10.1038/s41591-022-01932-x

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Parkinson's Disease Mechanisms and Treatments
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.