Comparative Evaluation of ChatGPT, DeepSeek, and Physician-Generated Acupuncture Treatment Protocols Across Chinese and English Language Settings: Cross-Sectional Evaluation Study

BACKGROUND Traditional Chinese medicine (TCM) has received increasing attention in evidence-based medicine. However, its largely experience-based and practice-oriented nature has slowed modernization and standardization. Advances in AI provide new opportunities for integrating AI with TCM. OBJECTIVE This study evaluates the application of large language models (LLMs), including ChatGPT (OpenAI) and DeepSeek, in generating acupuncture treatment protocols. We compare the validity and clinical effectiveness of physician-generated protocols with those produced by LLMs, examine differences in outputs across language corpora (Chinese and English), and establish a standardized assessment framework for systematically assessing the effectiveness of acupuncture point selection. Ultimately, this research aims to support the standardization of acupoint selection and promote greater consistency, reliability, and evidence-based development in acupuncture practice. METHODS Qualified acupuncture treatment cases published in Acupuncture in Medicine were identified and translated from English into Chinese to create a bilingual case dataset. Standardized case prompts were independently input into DeepSeek and ChatGPT to generate acupuncture treatment protocols. Ten senior TCM physicians evaluated the AI-generated and physician-generated protocols using 7 criteria: local point selection, distal point selection, syndrome-based point selection, meridian-tracing point selection, neuroanatomical point selection, core acupoint selection, and synergistic effects. Repeated-measures ANOVA with sphericity assessment was used to examine differences across LLMs, language corpora, and protocol groups. RESULTS Blinded expert evaluations revealed significant differences among the 5 acupuncture groups for local point selection (P=.006), syndrome-based point selection (P=.005), neuroanatomical point selection (P<.001), and core acupoints coverage principle (P=.03). For local point selection, DeepSeek Chinese output received higher scores than the other evaluated outputs, while ChatGPT English output scored higher than ChatGPT Chinese output. For syndrome-based point selection, the journal physician-generated protocols received lower scores than the DeepSeek Chinese, DeepSeek English, and ChatGPT Chinese outputs. For neuroanatomical point selection, DeepSeek English output achieved the highest scores among the 5 groups. For core acupoints coverage, the journal physician-generated protocols received lower scores than the DeepSeek Chinese, DeepSeek English, and ChatGPT English outputs. No significant differences were observed among the 5 groups for distal point selection, meridian-tracing point selection, or therapeutic synergy. Overall, DeepSeek performed better in the Chinese-language setting, whereas ChatGPT performed better in the English-language setting. CONCLUSIONS AI has demonstrated substantial progress in encoding explicit acupuncture knowledge and performing systematic acupoint selection. However, challenges remain in individualized treatment and consistency across linguistic contexts. Enhancing bilingual or multilingual training corpora and developing syndrome-specific reasoning modules may improve the clinical applicability of AI-assisted TCM systems. CLINICALTRIAL

Authors

Publication Details

Journal
JMIR Formative Research
Published
2026-10-08
DOI
https://doi.org/10.2196/95851
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Comparative Evaluation of ChatGPT, DeepSeek, and Physician-Generated Acupuncture Treatment Protocols Across Chinese and English Language Settings: Cross-Sectional Evaluation Study

Xin Liu, Zining Guo
JMIR Formative Research
Artificial Intelligence in Healthcare and Education
article

Comparative Evaluation of ChatGPT, DeepSeek, and Physician-Generated Acupuncture Treatment Protocols Across Chinese and English Language Settings: Cross-Sectional Evaluation Study

Xin Liu, Zining Guo
article en

Abstract

BACKGROUND Traditional Chinese medicine (TCM) has received increasing attention in evidence-based medicine. However, its largely experience-based and practice-oriented nature has slowed modernization and standardization. Advances in AI provide new opportunities for integrating AI with TCM. OBJECTIVE This study evaluates the application of large language models (LLMs), including ChatGPT (OpenAI) and DeepSeek, in generating acupuncture treatment protocols. We compare the validity and clinical effectiveness of physician-generated protocols with those produced by LLMs, examine differences in outputs across language corpora (Chinese and English), and establish a standardized assessment framework for systematically assessing the effectiveness of acupuncture point selection. Ultimately, this research aims to support the standardization of acupoint selection and promote greater consistency, reliability, and evidence-based development in acupuncture practice. METHODS Qualified acupuncture treatment cases published in Acupuncture in Medicine were identified and translated from English into Chinese to create a bilingual case dataset. Standardized case prompts were independently input into DeepSeek and ChatGPT to generate acupuncture treatment protocols. Ten senior TCM physicians evaluated the AI-generated and physician-generated protocols using 7 criteria: local point selection, distal point selection, syndrome-based point selection, meridian-tracing point selection, neuroanatomical point selection, core acupoint selection, and synergistic effects. Repeated-measures ANOVA with sphericity assessment was used to examine differences across LLMs, language corpora, and protocol groups. RESULTS Blinded expert evaluations revealed significant differences among the 5 acupuncture groups for local point selection (P=.006), syndrome-based point selection (P=.005), neuroanatomical point selection (P<.001), and core acupoints coverage principle (P=.03). For local point selection, DeepSeek Chinese output received higher scores than the other evaluated outputs, while ChatGPT English output scored higher than ChatGPT Chinese output. For syndrome-based point selection, the journal physician-generated protocols received lower scores than the DeepSeek Chinese, DeepSeek English, and ChatGPT Chinese outputs. For neuroanatomical point selection, DeepSeek English output achieved the highest scores among the 5 groups. For core acupoints coverage, the journal physician-generated protocols received lower scores than the DeepSeek Chinese, DeepSeek English, and ChatGPT English outputs. No significant differences were observed among the 5 groups for distal point selection, meridian-tracing point selection, or therapeutic synergy. Overall, DeepSeek performed better in the Chinese-language setting, whereas ChatGPT performed better in the English-language setting. CONCLUSIONS AI has demonstrated substantial progress in encoding explicit acupuncture knowledge and performing systematic acupoint selection. However, challenges remain in individualized treatment and consistency across linguistic contexts. Enhancing bilingual or multilingual training corpora and developing syndrome-specific reasoning modules may improve the clinical applicability of AI-assisted TCM systems. CLINICALTRIAL

JMIR Formative ResearchVol. 10
Openalex Percentile: Top 19%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.