Synergistic integration of large language models and knowledge graphs for intelligent metabolic pathway design in food synthetic biology
Metabolic pathway design in food synthetic biology requires navigating vast, heterogeneous biochemical data and complex enzymatic networks, yet existing computational approaches struggle with incomplete databases, inefficient reasoning over large search spaces, and insufficient predictive accuracy for thermodynamic and kinetic feasibility. Here we present an integrated framework that synergistically couples large language models (LLMs) with a domain-specific knowledge graph (KG) to enable intelligent pathway inference and optimization. The framework comprises three core modules: (i) a food metabolic pathway knowledge graph consolidating over 66,000 entities and 185,000 relations from KEGG, MetaCyc, BRENDA, UniProt, and literature mining; (ii) an LLM-driven reasoning engine that fuses graph-structured retrieval with learned semantic inference through a cross-modal attention mechanism, enabling both topological traversal and gap-filling over incomplete reaction networks; and (iii) a multi-objective optimization model, solved via a hybrid genetic algorithm with LLM-mediated repair, that jointly maximizes theoretical yield and thermodynamic driving force while minimizing heterologous enzyme burden. Benchmark experiments on 32 experimentally validated biosynthetic pathways demonstrate that the integrated system achieves an 87.5% pathway hit rate and 82.3% enzyme prediction accuracy, surpassing the best single-method baseline by 18.7 and 5.9 percentage points, respectively. Ablation studies confirm that the LLM and KG components address complementary failure modes—the graph suppresses hallucinated reactions while the language model bridges missing edges—producing gains that are complementary rather than merely additive. The hybrid optimizer attains a hypervolume indicator of 0.861, representing a 15.9% improvement over standard evolutionary search. On the eight pathways withheld from the graph the hit rate is 75.0%, so the advantage does not rest on graph coverage alone. We frame the contribution as a domain adaptation rather than a new computational paradigm: GraphRAG principles are specialized here for food-relevant biosynthesis, and the constraint layers establish computational plausibility rather than experimental feasibility, leaving enzyme specificity, kinetics, regulation and host physiology to laboratory testing. The implementation, the knowledge graph and the evaluation data are released with the paper.
Authors
- Yuan Cao (ORCID: https://orcid.org/0000-0002-3779-9982)
- Jinyang Zhou
- Nan Cheng
- Guangxin Zhu
Institutions
- National Library of China (CN)
- China Agricultural University (CN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-08
- DOI
- https://doi.org/10.1038/s41598-026-70322-x
- Primary Topic
- Microbial Metabolic Engineering and Bioproduction
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- China Agricultural University