ChemAIFlow: a scalable and interpretable applied-AI framework for high-dimensional data analytics in health informatics

Purpose The study addresses operational demands for scalable artificial intelligence frameworks in high-dimensional affinity prediction. Current computational solutions prioritize architectural complexity over interpretability. Architectural complexity creates operational barriers for deployment in decision-support systems. Design/methodology/approach The ChemAIFlow pipeline processed 68,443 bioactive interaction pairs across 500 targets utilizing a hybrid descriptor framework. Methodological implementation encompasses structural deduplication yielding 55,535 unique chemical compounds, hyperparameter optimization, and Y-scrambling validation. Evaluation protocols utilized 5-fold cross-validation and residual analysis for applicability domain definition. Findings Gradient Boosted Trees functioned as the primary inference engine. The model achieved an R2 of 0.992, an RMSE of 0.127, and an MAE of 0.070. Interpretability analysis identified Ligand-Lipophilicity Efficiency and Binding Efficiency Index as primary predictive drivers. Applicability domain assessment confirmed 98% of predictions within confidence thresholds. Prospective execution demonstrated a screening reduction to the top 0.018% of the available chemical space. Research limitations/implications Predictive capability remains dependent on data quality within public biochemical repositories. Structural data leakage inherent to random dataset partitioning represents a recognized methodology limitation. Future research requires selective enrichment of feature representations and cluster-based partitioning protocols. Practical implications System implementation on non-specialized hardware lowers operational barriers for virtual screening. The framework facilitates lead prioritization within standard health informatics infrastructures. Dedicated computing clusters remain unnecessary for system execution. Originality/value The research presents a lightweight engineering architecture capable of large-scale pharmacological profiling without resource-intensive computing. Integration of gradient boosting techniques with feature attribution establishes a reproducible workflow for high-throughput chemical triage.

Authors

Institutions

Publication Details

Journal
Applied Computing and Informatics
Published
2026-10-05
DOI
https://doi.org/10.1108/aci-12-2025-0542
Primary Topic
Computational Drug Discovery Methods
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

ChemAIFlow: a scalable and interpretable applied-AI framework for high-dimensional data analytics in health informatics

Achmad Nizar Hidayanto, Chairote Yaiprasert, Natthanan Na Prasit
Applied Computing and Informatics
Computational Drug Discovery Methods
article

ChemAIFlow: a scalable and interpretable applied-AI framework for high-dimensional data analytics in health informatics

Achmad Nizar Hidayanto, Chairote Yaiprasert, Natthanan Na Prasit
article en

Abstract

Purpose The study addresses operational demands for scalable artificial intelligence frameworks in high-dimensional affinity prediction. Current computational solutions prioritize architectural complexity over interpretability. Architectural complexity creates operational barriers for deployment in decision-support systems. Design/methodology/approach The ChemAIFlow pipeline processed 68,443 bioactive interaction pairs across 500 targets utilizing a hybrid descriptor framework. Methodological implementation encompasses structural deduplication yielding 55,535 unique chemical compounds, hyperparameter optimization, and Y-scrambling validation. Evaluation protocols utilized 5-fold cross-validation and residual analysis for applicability domain definition. Findings Gradient Boosted Trees functioned as the primary inference engine. The model achieved an R2 of 0.992, an RMSE of 0.127, and an MAE of 0.070. Interpretability analysis identified Ligand-Lipophilicity Efficiency and Binding Efficiency Index as primary predictive drivers. Applicability domain assessment confirmed 98% of predictions within confidence thresholds. Prospective execution demonstrated a screening reduction to the top 0.018% of the available chemical space. Research limitations/implications Predictive capability remains dependent on data quality within public biochemical repositories. Structural data leakage inherent to random dataset partitioning represents a recognized methodology limitation. Future research requires selective enrichment of feature representations and cluster-based partitioning protocols. Practical implications System implementation on non-specialized hardware lowers operational barriers for virtual screening. The framework facilitates lead prioritization within standard health informatics infrastructures. Dedicated computing clusters remain unnecessary for system execution. Originality/value The research presents a lightweight engineering architecture capable of large-scale pharmacological profiling without resource-intensive computing. Integration of gradient boosting techniques with feature attribution establishes a reproducible workflow for high-throughput chemical triage.

Applied Computing and Informatics
University of Indonesia (ID), Walailak University (TH)
Good health and well-being
Openalex Percentile: Top 31%
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.