Domain-Aware Feature Engineering for Prediction of Molecular Optical Interference in High-Throughput Screening
Abstract Optical interference from small molecules─compounds that absorb UV/visible light or exhibit intrinsic fluorescence─is a significant source of false positives and false negatives in fluorescence- and absorbance-based high-throughput screening assays. Computational prediction of optical activity from molecular structure offers a promising route to early triage of such compounds, yet the problem is challenging due to its quantum mechanical underpinnings and the extreme class imbalance characteristic of diverse drug-like libraries. Here I describe a domain-aware feature engineering and automated hyperparameter optimization approach developed for the EUOS25 Challenge, a benchmark comprising approximately 100,000 small organic compounds with experimentally measured UV/visible absorbance and fluorescence spectra. Beyond standard RDKit physicochemical descriptors and molecular fingerprints, I developed four domain-specific descriptor groups that encode prior chemical knowledge about chromophoric systems: count fingerprints of Murcko scaffolds derived from known dyes and fluorophores, binary indicators for manually curated chromophoric substructures, semiempirical quantum-mechanical descriptors from GFN2-xTB calculations, and ML-predicted quantum properties trained on the QM9 data set. Meta-learning features - probability outputs of baseline models trained on structurally related, easier-to-model end points - were included as cross-task inputs. Hyperparameter optimization was performed using the FLAML automated machine learning framework. The best model for each end point was selected from three complementary strategies: single best model, ensemble consensus, and semisupervised pseudolabeling. Postchallenge feature importance and ablation analyses showed that cross-task meta-features were the dominant driver of predictive accuracy for physically related end points. Semiempirical xTB descriptors─in particular─HOMO–LUMO gap and dipole/quadrupole terms - rank among the top-weighted CatBoost features and provide a physically interpretable account of model behavior. The dye-scaffold and chromophore-indicator descriptors, by contrast, carry genuine but redundant signal: used in isolation they classify well above chance for the two well-predicted end points, yet their removal from the full feature set produces no statistically significant change for any end point, indicating that the fingerprint/2D backbone already encodes this information. These results demonstrate that combining a standard cheminformatics backbone with cross-task meta-learning and automated hyperparameter optimization yields highly competitive predictive models on this benchmark.
Authors
- Filip Stefaniak (ORCID: https://orcid.org/0000-0001-5758-9416)
Institutions
- International Institute of Molecular and Cell Biology (PL)
Publication Details
- Journal
- Journal of Chemical Information and Modeling
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1021/acs.jcim.6c02217
- Primary Topic
- Computational Drug Discovery Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00