Genolator enables protein function interpretation using a multimodal large language model fusing genomic and structural interpretation with natural language interaction
Abstract Background Decoding the genetic code to unveil its genome functionality is a monumental task which would greatly advance the understanding of disease mechanisms and development of targeted treatments. Although large language models (LLMs) have transformed natural language processing across diverse domains, translating the complex language of DNA into human-readable form remains challenging due to genomic data complexity and unexplored regions of the human genome. Current language models either are capable of processing natural language or the genomic code. Models fusing both aspects are largely lacking. Results Here we present Genolator, a multimodal large language model that integrates embeddings from DNA sequences, amino acid sequences, and protein structures with natural language queries. Fine-tuned on over 365,000 question–answer pairs generated using abstracted Gene-Ontology (GO) terms, Genolator effectively answers queries regarding protein subcellular localization, molecular function, and biological processes. Evaluation demonstrates high accuracy in confirming or denying protein function associations, outperforming baseline models such as openly available allrounder LLMs like GPT 4.1 as well as smaller domain-specific models integrating knowledge from a protein structure transformer. Explorations of Genolator’s hidden states unveil a biologically and linguistically plausible organization of its learned representations. Analysis of the attention heads of the underlying language model and an ablation study provide evidence for a benefit of the multi-modal approach. Conclusion Genolator enhances accessibility to genomic information by enabling natural language interaction with protein data, facilitating biological discovery, and clinical research. It represents a step towards bridging genomic code and human language through the integration of a multimodal LLM.
Authors
- Jeremias Krause (ORCID: https://orcid.org/0000-0001-9915-7400)
- Tanhim Islam (ORCID: https://orcid.org/0000-0003-3182-1138)
- Martin Danner (ORCID: https://orcid.org/0000-0002-5459-4452)
- Matthias Begemann (ORCID: https://orcid.org/0000-0002-4659-8437)
- Miriam Elbracht (ORCID: https://orcid.org/0000-0001-5088-1369)
- Ingo Kurth
- Florian Kraft
Institutions
- Universitätsklinikum Aachen (DE)
- Institute for Resource Efficiency and Energy Strategies (DE)
- RWTH Aachen University (DE)
Publication Details
- Journal
- Genome biology
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1186/s13059-026-04274-w
- Primary Topic
- Genomics and Rare Diseases
- Type
- article
- Field-Weighted Citation Impact
- 0.00