CleanSurvival: automated data preprocessing for time-to-event models using reinforcement learning

Abstract Background Data preprocessing is often paid little attention in machine learning, despite its potentially significant impact on model performance. While automated machine learning pipelines are starting to recognise and integrate data preprocessing into their solutions for classification and regression tasks, this integration is lacking for more specialised tasks like time-to-event models for censored data. As a result, survival analysis not only faces the general challenges of data preprocessing but also suffers from the lack of tailored, automated solutions in this area. Method To address this gap, this paper presents , a reinforcement-learning-based solution for optimizing preprocessing pipelines, extended specifically for survival analysis. The framework can handle continuous and categorical variables. It builds upon Learn2Clean’s $$Q$$ -learning to select which combination of data imputation, outlier detection and feature extraction techniques achieves optimal performance for a Cox, random forest, neural network or user-supplied time-to-event model. The Python package is available on GitHub: https://github.com/datasciapps/CleanSurvival . Results Experimental benchmarks on real-world datasets show that the $$Q$$ -learning-based data preprocessing can improve predictive performance relative to simple baselines, while runtime behaviour is condition-dependent and most clearly interpretable in the best-covered benchmark cells. Furthermore, a simulation study demonstrates effectiveness across different types and levels of missingness and noise. Conclusion With an increase in the use of machine learning, it becomes important to generalise AutoML pipelines to a variety of models now present, including survival analysis. Tools like , which integrate preprocessing for survival analysis, can make survival studies faster and easier to perform, while also yielding more robust results.

Authors

Publication Details

Journal
BMC Medical Informatics and Decision Making
Published
2026-09-24
DOI
https://doi.org/10.1186/s12911-026-03786-6
Primary Topic
Simulation Techniques and Applications
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

CleanSurvival: automated data preprocessing for time-to-event models using reinforcement learning

Gerrit Großmann, David Selby, Yousef Koka
BMC Medical Informatics and Decision Making
Simulation Techniques and Applications
article

CleanSurvival: automated data preprocessing for time-to-event models using reinforcement learning

Gerrit Großmann, David Selby, Yousef Koka
article en

Abstract

Abstract Background Data preprocessing is often paid little attention in machine learning, despite its potentially significant impact on model performance. While automated machine learning pipelines are starting to recognise and integrate data preprocessing into their solutions for classification and regression tasks, this integration is lacking for more specialised tasks like time-to-event models for censored data. As a result, survival analysis not only faces the general challenges of data preprocessing but also suffers from the lack of tailored, automated solutions in this area. Method To address this gap, this paper presents , a reinforcement-learning-based solution for optimizing preprocessing pipelines, extended specifically for survival analysis. The framework can handle continuous and categorical variables. It builds upon Learn2Clean’s $$Q$$ -learning to select which combination of data imputation, outlier detection and feature extraction techniques achieves optimal performance for a Cox, random forest, neural network or user-supplied time-to-event model. The Python package is available on GitHub: https://github.com/datasciapps/CleanSurvival . Results Experimental benchmarks on real-world datasets show that the $$Q$$ -learning-based data preprocessing can improve predictive performance relative to simple baselines, while runtime behaviour is condition-dependent and most clearly interpretable in the best-covered benchmark cells. Furthermore, a simulation study demonstrates effectiveness across different types and levels of missingness and noise. Conclusion With an increase in the use of machine learning, it becomes important to generalise AutoML pipelines to a variety of models now present, including survival analysis. Tools like , which integrate preprocessing for survival analysis, can make survival studies faster and easier to perform, while also yielding more robust results.

BMC Medical Informatics and Decision MakingVol. 26(1)
Bundesministerium für Bildung und Forschung
Openalex Percentile: Top 99%
Simulation Techniques and Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

CleanSurvival: automated data preprocessing for time-to-event models using reinforcement learning — Gerrit Großmann, David Selby, et al. · BMC Medical Informatics and Decision Making (2026) | TGRS Research Map | TGRS