Performance comparison of a neuro-symbolic large language model system versus conventional AI models and human experts in cholangitis management

BACKGROUND: Large language models (LLMs) have shown promising results in medical decision support; Background: Large language models (LLMs) have demonstrated promising outcomes in medical decision support; however, their efficacy in managing complex hepatobiliary conditions remains insufficiently examined. We have developed a genetic neuro-symbolic LLM system that integrates multiple AI agents with neural-symbolic reasoning for the management of cholangitis, and we have compared its performance to that of conventional LLMs and human experts.genetic neuro-symbolic LLM system integrating multiple AI agents with neural-symbolic reasoning for cholangitis management and compared its performance against conventional LLMs and human experts. METHODS: This multi-center cross-sectional study included 30 case-based questions from American Board of Internal Medicine (ABIM) gastroenterology subspecialty examinations covering acute cholangitis. Questions were categorized into diagnosis (n = 10), treatment (n = 10), and complications/prognosis (n = 10). Performance of a genetic neuro-symbolic LLM system orchestrated via LangGraph was compared against Claude 4.5 Sonnet, ChatGPT 5.2, Gemini 2.0 Flash, 10 gastroenterology specialists, and 4 emergency medicine physicians from four tertiary centers in Turkey. RESULTS: The genetic neuro-symbolic system achieved the highest overall accuracy (100%, 30/30), significantly outperforming Claude 4.5 Sonnet (90.0%), ChatGPT 5.2 (60.0%), Gemini 2.0 Flash (63.3%), gastroenterology experts (mean 95.7% ± 3.2%), and emergency medicine physicians (mean 84.2% ± 8.8%). The neuro-symbolic system demonstrated superior performance across all categories and cholangitis subtypes. Among human participants, gastroenterologists outperformed emergency physicians in treatment decisions (p = 0.012) and showed non-inferior performance to Gemini 2.0 Flash overall (p = 0.034). CONCLUSIONS: The genetic neuro-symbolic LLM system demonstrated superior accuracy in cholangitis management compared to all conventional AI models and human experts. This proof-of-concept study suggests that multi-agent architectures with neural-symbolic reasoning may offer a promising direction for AI-assisted clinical decision support in complex hepatobiliary conditions, although prospective clinical validation is required before broader implementation claims can be warranted.

Authors

Institutions

Publication Details

Journal
BMC Medical Informatics and Decision Making
Published
2026-06-01
DOI
https://doi.org/10.1186/s12911-026-03593-z
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Performance comparison of a neuro-symbolic large language model system versus conventional AI models and human experts in cholangitis management

Evren Ekingen, Mete Ucdal
BMC Medical Informatics and Decision Making
Artificial Intelligence in Healthcare and Education
article

Performance comparison of a neuro-symbolic large language model system versus conventional AI models and human experts in cholangitis management

Evren Ekingen, Mete Ucdal
article en

Abstract

BACKGROUND: Large language models (LLMs) have shown promising results in medical decision support; Background: Large language models (LLMs) have demonstrated promising outcomes in medical decision support; however, their efficacy in managing complex hepatobiliary conditions remains insufficiently examined. We have developed a genetic neuro-symbolic LLM system that integrates multiple AI agents with neural-symbolic reasoning for the management of cholangitis, and we have compared its performance to that of conventional LLMs and human experts.genetic neuro-symbolic LLM system integrating multiple AI agents with neural-symbolic reasoning for cholangitis management and compared its performance against conventional LLMs and human experts. METHODS: This multi-center cross-sectional study included 30 case-based questions from American Board of Internal Medicine (ABIM) gastroenterology subspecialty examinations covering acute cholangitis. Questions were categorized into diagnosis (n = 10), treatment (n = 10), and complications/prognosis (n = 10). Performance of a genetic neuro-symbolic LLM system orchestrated via LangGraph was compared against Claude 4.5 Sonnet, ChatGPT 5.2, Gemini 2.0 Flash, 10 gastroenterology specialists, and 4 emergency medicine physicians from four tertiary centers in Turkey. RESULTS: The genetic neuro-symbolic system achieved the highest overall accuracy (100%, 30/30), significantly outperforming Claude 4.5 Sonnet (90.0%), ChatGPT 5.2 (60.0%), Gemini 2.0 Flash (63.3%), gastroenterology experts (mean 95.7% ± 3.2%), and emergency medicine physicians (mean 84.2% ± 8.8%). The neuro-symbolic system demonstrated superior performance across all categories and cholangitis subtypes. Among human participants, gastroenterologists outperformed emergency physicians in treatment decisions (p = 0.012) and showed non-inferior performance to Gemini 2.0 Flash overall (p = 0.034). CONCLUSIONS: The genetic neuro-symbolic LLM system demonstrated superior accuracy in cholangitis management compared to all conventional AI models and human experts. This proof-of-concept study suggests that multi-agent architectures with neural-symbolic reasoning may offer a promising direction for AI-assisted clinical decision support in complex hepatobiliary conditions, although prospective clinical validation is required before broader implementation claims can be warranted.

BMC Medical Informatics and Decision Making
Hacettepe University Hospital (TR), Etimesgut Asker Hastanesi (TR)
Peace, Justice and strong institutions
Openalex Percentile: Top 9%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.