Scalable molecular representations enabled by multimodal fusion and sequence distillation
Molecular prediction depends on how chemical structures are represented, yet descriptors capture only partial aspects of chemical information. Here, we show that the Chemical Dice Integrator combines six complementary molecular views spanning physicochemical properties, molecular topology, two-dimensional structural images, bioactivity profiles, quantum properties, and molecular language into a unified latent space. This multimodal representation is distilled into a sequence-based model that generates the embedding directly from molecular strings. Across classification and regression benchmarks, the representation outperformed classical fusion methods, matched or improved established molecular descriptors, retained complete embedding coverage when individual feature-generation pipelines failed, and supported efficient inference. Scaffold-based and low-data evaluations showed stable generalization across chemical space and improved utility under limited data. The framework also prioritized compounds predicted to protect genome stability, leading to experimental validation of isoeugenol and eugenyl acetate in a yeast damage-response assay. These findings establish a scalable framework for molecular prediction and discovery. This study integrates six complementary molecular representations into a distilled molecular string embedding, enabling scalable property prediction and prioritization of compounds that reduce DNA-damage phenotypes in yeast.
Authors
- Subhadeep Duari (ORCID: https://orcid.org/0000-0001-9805-7635)
- Sonam Chauhan
- Abhinav Kumar Sharma (ORCID: https://orcid.org/0000-0002-4774-0669)
- Sanjay Kumar Mohanty (ORCID: https://orcid.org/0000-0002-1375-2223)
- Shiva Satija (ORCID: https://orcid.org/0000-0001-8768-1539)
- Saveena Solanki (ORCID: https://orcid.org/0000-0003-1079-9349)
- Gaurav Ahuja (ORCID: https://orcid.org/0000-0002-2837-9361)
- Vishakha Gautam (ORCID: https://orcid.org/0000-0002-4755-0291)
- Debarka Sengupta (ORCID: https://orcid.org/0000-0002-6353-5411)
- Suvendu Kumar (ORCID: https://orcid.org/0000-0002-4405-8402)
- Aayushi Mittal (ORCID: https://orcid.org/0000-0002-1973-8553)
- Adnan Raza (ORCID: https://orcid.org/0009-0006-6398-3747)
- Mudit Gupta (ORCID: https://orcid.org/0000-0002-7812-1473)
- Sourav Sinha (ORCID: https://orcid.org/0000-0002-3327-8215)
- Syed Yasser Ali
- Raidhani Shome
- Sakshi Arora
- Arushi Sharma
- Natarajan Arul Murugan
Institutions
- Indraprastha Institute of Information Technology Delhi (IN)
- Indian Institute of Technology Delhi (IN)
Publication Details
- Journal
- Nature Communications
- Published
- 2026-09-19
- DOI
- https://doi.org/10.1038/s41467-026-77700-z
- Primary Topic
- Computational Drug Discovery Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00