Exploration and validation of large language models as tools for molecular optimization
Large Language Models (LLMs) are transforming the process of drug discovery through their ability to operate on chemical data, generate novel chemical structures, and enable natural language interaction between human experts and computational tools. A key question is how to effectively integrate LLMs into drug design projects to generate biologically active drug-like molecules. Here, we introduce a many-shot in-context learning methodology for generating drug-like molecules, demonstrating how LLM-based design can be effectively guided through prompt engineering for targeted chemical space exploration. Our evaluation demonstrates that both the structure-activity relationship (SAR) data provided in-context and the instructions given in the prompt have a significant impact on design outcomes: prompts without explicit structural constraints explore chemical space around the provided SAR examples, while prompts with explicit instructions—including property requirements, required substructures, scaffolds, or functional groups—enable navigation toward desired chemical cores while preserving and transferring SAR information. In two prospective studies, we validate this approach experimentally, achieving remarkable potency improvements in single optimization steps for inhibitors against Cathepsin A and RIPK1. These results establish LLMs as powerful tools for drug design, capable of transferring SAR knowledge between chemical series and accelerating hit-to-lead optimization through intuitive, language-based interaction to the design engine. The effective integration of large language models into drug design projects to generate biologically active drug-like molecules is a promising but challenging avenue of research. Here, the authors present a many-shot in-context learning approach using prompt engineering to guide LLMs, achieving significant potency improvements in drug design and demonstrating the potential of LLMs to accelerate hit-to-lead optimization.
Authors
- Jiří Vymětal (ORCID: https://orcid.org/0000-0002-0165-8707)
- Christoph Grebner (ORCID: https://orcid.org/0000-0001-5301-1078)
- Elisabeth Speckmeier (ORCID: https://orcid.org/0000-0003-2917-3741)
- Gerhard Heßler (ORCID: https://orcid.org/0000-0001-5602-0965)
- Jan H. Griwatz (ORCID: https://orcid.org/0000-0003-2028-0938)
- Christian Buning
- Thorsten Sadowski (ORCID: https://orcid.org/0000-0003-2250-5512)
- Alejandro Corrochano-Navarro
- Lorenzo Kogler-Anele
- Saeed Moayedpour (ORCID: https://orcid.org/0009-0007-6375-3084)
- Sven Jager
- María Méndez
- Hans Matter
- Ziv Bar-Joseph
- Sven Ruf
Institutions
- Sanofi (United States) (US)
- Sanofi (Germany) (DE)
Publication Details
- Journal
- Communications Chemistry
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1038/s42004-026-02193-2
- Primary Topic
- Machine Learning in Materials Science
- Type
- article
- Field-Weighted Citation Impact
- 0.00