Exploration and validation of large language models as tools for molecular optimization

Large Language Models (LLMs) are transforming the process of drug discovery through their ability to operate on chemical data, generate novel chemical structures, and enable natural language interaction between human experts and computational tools. A key question is how to effectively integrate LLMs into drug design projects to generate biologically active drug-like molecules. Here, we introduce a many-shot in-context learning methodology for generating drug-like molecules, demonstrating how LLM-based design can be effectively guided through prompt engineering for targeted chemical space exploration. Our evaluation demonstrates that both the structure-activity relationship (SAR) data provided in-context and the instructions given in the prompt have a significant impact on design outcomes: prompts without explicit structural constraints explore chemical space around the provided SAR examples, while prompts with explicit instructions—including property requirements, required substructures, scaffolds, or functional groups—enable navigation toward desired chemical cores while preserving and transferring SAR information. In two prospective studies, we validate this approach experimentally, achieving remarkable potency improvements in single optimization steps for inhibitors against Cathepsin A and RIPK1. These results establish LLMs as powerful tools for drug design, capable of transferring SAR knowledge between chemical series and accelerating hit-to-lead optimization through intuitive, language-based interaction to the design engine. The effective integration of large language models into drug design projects to generate biologically active drug-like molecules is a promising but challenging avenue of research. Here, the authors present a many-shot in-context learning approach using prompt engineering to guide LLMs, achieving significant potency improvements in drug design and demonstrating the potential of LLMs to accelerate hit-to-lead optimization.

Authors

Institutions

Publication Details

Journal
Communications Chemistry
Published
2026-09-15
DOI
https://doi.org/10.1038/s42004-026-02193-2
Primary Topic
Machine Learning in Materials Science
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Exploration and validation of large language models as tools for molecular optimization

Jiří Vymětal, Christoph Grebner, Elisabeth Speckmeier, Gerhard Heßler et al.
Communications Chemistry
Machine Learning in Materials Science
article

Exploration and validation of large language models as tools for molecular optimization

Jiří Vymětal, Christoph Grebner, Elisabeth Speckmeier, Gerhard Heßler, Jan H. Griwatz, Christian Buning, Thorsten Sadowski, Alejandro Corrochano-Navarro, Lorenzo Kogler-Anele, Saeed Moayedpour, Sven Jager, María Méndez, Hans Matter, Ziv Bar-Joseph, Sven Ruf
article en

Abstract

Large Language Models (LLMs) are transforming the process of drug discovery through their ability to operate on chemical data, generate novel chemical structures, and enable natural language interaction between human experts and computational tools. A key question is how to effectively integrate LLMs into drug design projects to generate biologically active drug-like molecules. Here, we introduce a many-shot in-context learning methodology for generating drug-like molecules, demonstrating how LLM-based design can be effectively guided through prompt engineering for targeted chemical space exploration. Our evaluation demonstrates that both the structure-activity relationship (SAR) data provided in-context and the instructions given in the prompt have a significant impact on design outcomes: prompts without explicit structural constraints explore chemical space around the provided SAR examples, while prompts with explicit instructions—including property requirements, required substructures, scaffolds, or functional groups—enable navigation toward desired chemical cores while preserving and transferring SAR information. In two prospective studies, we validate this approach experimentally, achieving remarkable potency improvements in single optimization steps for inhibitors against Cathepsin A and RIPK1. These results establish LLMs as powerful tools for drug design, capable of transferring SAR knowledge between chemical series and accelerating hit-to-lead optimization through intuitive, language-based interaction to the design engine. The effective integration of large language models into drug design projects to generate biologically active drug-like molecules is a promising but challenging avenue of research. Here, the authors present a many-shot in-context learning approach using prompt engineering to guide LLMs, achieving significant potency improvements in drug design and demonstrating the potential of LLMs to accelerate hit-to-lead optimization.

Communications Chemistry
Sanofi (United States) (US), Sanofi (Germany) (DE)
Quality Education
Openalex Percentile: Top 24%
Machine Learning in Materials Science
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.