A structured exploratory multidisciplinary evaluation of four large language models responding to frequently asked patient questions about pregabalin

Large Language Models (LLMs) are increasingly consulted by millions of patients seeking pharmaceutical information, yet their reliability and safety in providing medication-related advice remain inadequately evaluated. This study assessed the performance of four leading LLMs in responding to frequently asked questions about pregabalin, a widely prescribed antiepileptic and analgesic agent. We conducted a prospective comparative observational study from 15 June to 30 September 2025, evaluating ChatGPT-4o, Copilot, Gemini 2.0 Flash, and Claude 3.5 Sonnet. Fifteen frequently asked questions about pregabalin, stratified by thematic category and complexity level, were selected from National Health Service resources. A multidisciplinary panel comprising a rheumatologist, neurologist, and pharmacologist independently assessed responses using a multidimensional framework encompassing warning compliance, accuracy, completeness, content safety, readability, and clinical relevance. Inter-rater reliability was calculated using Fleiss’ kappa and intraclass correlation coefficient ICC(2,k), and error patterns were systematically categorized. Performance varied across models and dimensions without consistent superiority of any single LLM. ChatGPT-4o, Gemini 2.0 Flash, and Copilot showed identical accuracy scores (0.933), clinical relevance scores ranged from 0.867 to 1.000 across models, with overlapping confidence intervals throughout. No model showed consistent superiority across dimensions. All models demonstrated incomplete compliance with the explicit warning instruction, with scores ranging from 0.467 to 0.733. Error analysis revealed oversimplification as the predominant failure mode (56.3%), followed by omission of essential information (34.8%) and hallucinations (8.9%). Inter-rater agreement point estimates fell in the moderate range across all dimensions (κ = 0.44–0.52; ICC = 0.56–0.58; all p < 0.01), though lower confidence interval bounds extended into the fair to slight range, indicating residual uncertainty in agreement estimation consistent with the modest sample size of 15 questions. While LLMs demonstrated accuracy point estimates of 0.87–0.93 for selected basic pharmaceutical questions, significant limitations persist in safety warning compliance, response comprehensiveness, and hallucination rate.

Authors

Institutions

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-09-30
DOI
https://doi.org/10.1007/s44163-026-02394-7
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A structured exploratory multidisciplinary evaluation of four large language models responding to frequently asked patient questions about pregabalin

Abdoul Aziz, Dieu‐Donné Ouedraogo, Ismaël Ayouba Tinni, Wendlassida Martin Nacanabo et al.
Discover Artificial Intelligence
Artificial Intelligence in Healthcare and Education
article

A structured exploratory multidisciplinary evaluation of four large language models responding to frequently asked patient questions about pregabalin

Abdoul Aziz, Dieu‐Donné Ouedraogo, Ismaël Ayouba Tinni, Wendlassida Martin Nacanabo, Abdoul Nassir Porgo, Yannick Laurent Tchenadoyo Bayala, Fulgence Kaboré, Wendlassida Joëlle Stéphanie Zabsonré Tiendrébeogo, Abdoul Djalidou Traoré, Thierry Boris Wend-Yam Yaméogo
article en

Abstract

Large Language Models (LLMs) are increasingly consulted by millions of patients seeking pharmaceutical information, yet their reliability and safety in providing medication-related advice remain inadequately evaluated. This study assessed the performance of four leading LLMs in responding to frequently asked questions about pregabalin, a widely prescribed antiepileptic and analgesic agent. We conducted a prospective comparative observational study from 15 June to 30 September 2025, evaluating ChatGPT-4o, Copilot, Gemini 2.0 Flash, and Claude 3.5 Sonnet. Fifteen frequently asked questions about pregabalin, stratified by thematic category and complexity level, were selected from National Health Service resources. A multidisciplinary panel comprising a rheumatologist, neurologist, and pharmacologist independently assessed responses using a multidimensional framework encompassing warning compliance, accuracy, completeness, content safety, readability, and clinical relevance. Inter-rater reliability was calculated using Fleiss’ kappa and intraclass correlation coefficient ICC(2,k), and error patterns were systematically categorized. Performance varied across models and dimensions without consistent superiority of any single LLM. ChatGPT-4o, Gemini 2.0 Flash, and Copilot showed identical accuracy scores (0.933), clinical relevance scores ranged from 0.867 to 1.000 across models, with overlapping confidence intervals throughout. No model showed consistent superiority across dimensions. All models demonstrated incomplete compliance with the explicit warning instruction, with scores ranging from 0.467 to 0.733. Error analysis revealed oversimplification as the predominant failure mode (56.3%), followed by omission of essential information (34.8%) and hallucinations (8.9%). Inter-rater agreement point estimates fell in the moderate range across all dimensions (κ = 0.44–0.52; ICC = 0.56–0.58; all p < 0.01), though lower confidence interval bounds extended into the fair to slight range, indicating residual uncertainty in agreement estimation consistent with the modest sample size of 15 questions. While LLMs demonstrated accuracy point estimates of 0.87–0.93 for selected basic pharmaceutical questions, significant limitations persist in safety warning compliance, response comprehensiveness, and hallucination rate.

Discover Artificial IntelligenceVol. 6(1)
Université Joseph Ki-Zerbo (BF), National Hospital Niamey (NE), CHU de Bogodogo (BF)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.