Evidence-Grounded Clinical Pharmacogenomics Decision Support: Integrating Fine-Tuned Large Language Models with Hybrid Retrieval-Augmented Generation

Pharmacogenomics (PGx) play a key role in personalized medicine by guiding drug and dosage selection based on genetics. However, the volume and complexity of PGx data hinder clinical decision-making. This research proposes a data-driven clinical decision support framework that combines large language models (LLMs) with hybrid retrieval-augmented generation (RAG) to improve answers to PGx queries. The proposed framework evaluates Meta-LLaMA-3.1-8B-Instruct and Qwen3-8B across various configurations, including base models, Low-Rank Adaptation (LoRA) fine-tuning, and hybrid RAG methods. To build a robust dataset, structured data from the Clinical Pharmacogenetics Implementation Consortium (CPIC) and clinical guideline content from ClinPGx are prepared as JSON Lines (JSONL) resources, with CPIC-derived records used for instruction tuning and structured retrieval and ClinPGx guideline text used as a separate retrieval resource. The hybrid retrieval pipeline pairs lexical filtering with dense semantic similarity via sentence embeddings to maximize factual grounding. Evaluation relies on both automated metrics and human clinical review for correctness, relevance, completeness, and clarity. Results indicate that Meta-LLaMA-3.1-8B-Instruct benefits most consistently from the combined RAG and LoRA configuration, while Qwen3-8B shows more modest action-level classification performance but improved evidence-grounded text-generation quality when retrieval is added. Fine-tuning alone proved insufficient, highlighting the limitations of purely parametric knowledge. This study shows that combining retrieval methods with parameter-efficient fine-tuning enhances LLM reliability in clinical settings. The proposed methodology offers a scalable, trustworthy framework for AI-driven decision support in pharmacogenomics and broader healthcare applications.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-09-15
DOI
https://doi.org/10.3390/app16189148
Primary Topic
Biomedical Text Mining and Ontologies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evidence-Grounded Clinical Pharmacogenomics Decision Support: Integrating Fine-Tuned Large Language Models with Hybrid Retrieval-Augmented Generation

Abedalrhman Alkhateeb, Protiva Arafin, Md Moniruzzaman
Applied Sciences
Biomedical Text Mining and Ontologies
article

Evidence-Grounded Clinical Pharmacogenomics Decision Support: Integrating Fine-Tuned Large Language Models with Hybrid Retrieval-Augmented Generation

Abedalrhman Alkhateeb, Protiva Arafin, Md Moniruzzaman
article en

Abstract

Pharmacogenomics (PGx) play a key role in personalized medicine by guiding drug and dosage selection based on genetics. However, the volume and complexity of PGx data hinder clinical decision-making. This research proposes a data-driven clinical decision support framework that combines large language models (LLMs) with hybrid retrieval-augmented generation (RAG) to improve answers to PGx queries. The proposed framework evaluates Meta-LLaMA-3.1-8B-Instruct and Qwen3-8B across various configurations, including base models, Low-Rank Adaptation (LoRA) fine-tuning, and hybrid RAG methods. To build a robust dataset, structured data from the Clinical Pharmacogenetics Implementation Consortium (CPIC) and clinical guideline content from ClinPGx are prepared as JSON Lines (JSONL) resources, with CPIC-derived records used for instruction tuning and structured retrieval and ClinPGx guideline text used as a separate retrieval resource. The hybrid retrieval pipeline pairs lexical filtering with dense semantic similarity via sentence embeddings to maximize factual grounding. Evaluation relies on both automated metrics and human clinical review for correctness, relevance, completeness, and clarity. Results indicate that Meta-LLaMA-3.1-8B-Instruct benefits most consistently from the combined RAG and LoRA configuration, while Qwen3-8B shows more modest action-level classification performance but improved evidence-grounded text-generation quality when retrieval is added. Fine-tuning alone proved insufficient, highlighting the limitations of purely parametric knowledge. This study shows that combining retrieval methods with parameter-efficient fine-tuning enhances LLM reliability in clinical settings. The proposed methodology offers a scalable, trustworthy framework for AI-driven decision support in pharmacogenomics and broader healthcare applications.

Applied SciencesVol. 16(18)
Thompson Rivers University (CA), Lakehead University (CA)
Peace, Justice and strong institutions
Openalex Percentile: Top 18%
Biomedical Text Mining and Ontologies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Evidence-Grounded Clinical Pharmacogenomics Decision Support: Integrating Fine-Tuned Large Language Models with Hybrid Retrieval-Augmented Generation — Abedalrhman Alkhateeb, Protiva Arafin, et al. · Applied Sciences (2026) | TGRS Research Map | TGRS