ChemAIFlow: a scalable and interpretable applied-AI framework for high-dimensional data analytics in health informatics
Purpose The study addresses operational demands for scalable artificial intelligence frameworks in high-dimensional affinity prediction. Current computational solutions prioritize architectural complexity over interpretability. Architectural complexity creates operational barriers for deployment in decision-support systems. Design/methodology/approach The ChemAIFlow pipeline processed 68,443 bioactive interaction pairs across 500 targets utilizing a hybrid descriptor framework. Methodological implementation encompasses structural deduplication yielding 55,535 unique chemical compounds, hyperparameter optimization, and Y-scrambling validation. Evaluation protocols utilized 5-fold cross-validation and residual analysis for applicability domain definition. Findings Gradient Boosted Trees functioned as the primary inference engine. The model achieved an R2 of 0.992, an RMSE of 0.127, and an MAE of 0.070. Interpretability analysis identified Ligand-Lipophilicity Efficiency and Binding Efficiency Index as primary predictive drivers. Applicability domain assessment confirmed 98% of predictions within confidence thresholds. Prospective execution demonstrated a screening reduction to the top 0.018% of the available chemical space. Research limitations/implications Predictive capability remains dependent on data quality within public biochemical repositories. Structural data leakage inherent to random dataset partitioning represents a recognized methodology limitation. Future research requires selective enrichment of feature representations and cluster-based partitioning protocols. Practical implications System implementation on non-specialized hardware lowers operational barriers for virtual screening. The framework facilitates lead prioritization within standard health informatics infrastructures. Dedicated computing clusters remain unnecessary for system execution. Originality/value The research presents a lightweight engineering architecture capable of large-scale pharmacological profiling without resource-intensive computing. Integration of gradient boosting techniques with feature attribution establishes a reproducible workflow for high-throughput chemical triage.
Authors
- Achmad Nizar Hidayanto (ORCID: https://orcid.org/0000-0002-5793-9460)
- Chairote Yaiprasert (ORCID: https://orcid.org/0000-0001-7012-1324)
- Natthanan Na Prasit
Institutions
- University of Indonesia (ID)
- Walailak University (TH)
Publication Details
- Journal
- Applied Computing and Informatics
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1108/aci-12-2025-0542
- Primary Topic
- Computational Drug Discovery Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00