Domain-Aware Feature Engineering for Prediction of Molecular Optical Interference in High-Throughput Screening

Abstract Optical interference from small molecules─compounds that absorb UV/visible light or exhibit intrinsic fluorescence─is a significant source of false positives and false negatives in fluorescence- and absorbance-based high-throughput screening assays. Computational prediction of optical activity from molecular structure offers a promising route to early triage of such compounds, yet the problem is challenging due to its quantum mechanical underpinnings and the extreme class imbalance characteristic of diverse drug-like libraries. Here I describe a domain-aware feature engineering and automated hyperparameter optimization approach developed for the EUOS25 Challenge, a benchmark comprising approximately 100,000 small organic compounds with experimentally measured UV/visible absorbance and fluorescence spectra. Beyond standard RDKit physicochemical descriptors and molecular fingerprints, I developed four domain-specific descriptor groups that encode prior chemical knowledge about chromophoric systems: count fingerprints of Murcko scaffolds derived from known dyes and fluorophores, binary indicators for manually curated chromophoric substructures, semiempirical quantum-mechanical descriptors from GFN2-xTB calculations, and ML-predicted quantum properties trained on the QM9 data set. Meta-learning features - probability outputs of baseline models trained on structurally related, easier-to-model end points - were included as cross-task inputs. Hyperparameter optimization was performed using the FLAML automated machine learning framework. The best model for each end point was selected from three complementary strategies: single best model, ensemble consensus, and semisupervised pseudolabeling. Postchallenge feature importance and ablation analyses showed that cross-task meta-features were the dominant driver of predictive accuracy for physically related end points. Semiempirical xTB descriptors─in particular─HOMO–LUMO gap and dipole/quadrupole terms - rank among the top-weighted CatBoost features and provide a physically interpretable account of model behavior. The dye-scaffold and chromophore-indicator descriptors, by contrast, carry genuine but redundant signal: used in isolation they classify well above chance for the two well-predicted end points, yet their removal from the full feature set produces no statistically significant change for any end point, indicating that the fingerprint/2D backbone already encodes this information. These results demonstrate that combining a standard cheminformatics backbone with cross-task meta-learning and automated hyperparameter optimization yields highly competitive predictive models on this benchmark.

Authors

Institutions

Publication Details

Journal
Journal of Chemical Information and Modeling
Published
2026-10-06
DOI
https://doi.org/10.1021/acs.jcim.6c02217
Primary Topic
Computational Drug Discovery Methods
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Domain-Aware Feature Engineering for Prediction of Molecular Optical Interference in High-Throughput Screening

Filip Stefaniak
Journal of Chemical Information and Modeling
Computational Drug Discovery Methods
article

Domain-Aware Feature Engineering for Prediction of Molecular Optical Interference in High-Throughput Screening

Filip Stefaniak
article en

Abstract

Abstract Optical interference from small molecules─compounds that absorb UV/visible light or exhibit intrinsic fluorescence─is a significant source of false positives and false negatives in fluorescence- and absorbance-based high-throughput screening assays. Computational prediction of optical activity from molecular structure offers a promising route to early triage of such compounds, yet the problem is challenging due to its quantum mechanical underpinnings and the extreme class imbalance characteristic of diverse drug-like libraries. Here I describe a domain-aware feature engineering and automated hyperparameter optimization approach developed for the EUOS25 Challenge, a benchmark comprising approximately 100,000 small organic compounds with experimentally measured UV/visible absorbance and fluorescence spectra. Beyond standard RDKit physicochemical descriptors and molecular fingerprints, I developed four domain-specific descriptor groups that encode prior chemical knowledge about chromophoric systems: count fingerprints of Murcko scaffolds derived from known dyes and fluorophores, binary indicators for manually curated chromophoric substructures, semiempirical quantum-mechanical descriptors from GFN2-xTB calculations, and ML-predicted quantum properties trained on the QM9 data set. Meta-learning features - probability outputs of baseline models trained on structurally related, easier-to-model end points - were included as cross-task inputs. Hyperparameter optimization was performed using the FLAML automated machine learning framework. The best model for each end point was selected from three complementary strategies: single best model, ensemble consensus, and semisupervised pseudolabeling. Postchallenge feature importance and ablation analyses showed that cross-task meta-features were the dominant driver of predictive accuracy for physically related end points. Semiempirical xTB descriptors─in particular─HOMO–LUMO gap and dipole/quadrupole terms - rank among the top-weighted CatBoost features and provide a physically interpretable account of model behavior. The dye-scaffold and chromophore-indicator descriptors, by contrast, carry genuine but redundant signal: used in isolation they classify well above chance for the two well-predicted end points, yet their removal from the full feature set produces no statistically significant change for any end point, indicating that the fingerprint/2D backbone already encodes this information. These results demonstrate that combining a standard cheminformatics backbone with cross-task meta-learning and automated hyperparameter optimization yields highly competitive predictive models on this benchmark.

Journal of Chemical Information and Modeling
International Institute of Molecular and Cell Biology (PL)
Openalex Percentile: Top 12%
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.