Symptom Terminology Normalization in Traditional Chinese Medicine: Development and Evaluation of a 2-Stage Deep Learning Framework Based on Fine-Grained Semantic Classification
Abstract Background Due to the heterogeneity of symptom terminology and the lack of industry standards, the same symptom is often described using multiple expressions. Current normalization approaches struggle to comprehensively retrieve standard terms when a raw term maps to multiple symptoms. Objective This study aimed to address the lack of industry standards for traditional Chinese medicine (TCM) symptom terminology. This study proposed the split-then-concatenate normalization framework (STC-NF), a novel approach based on fine-grained semantic classification and a 2-stage deep learning architecture that uses electronic medical records (EMRs) as the data source. Methods This study proposed a 2-stage deep learning framework, “split-then-concatenate.” In the splitting stage, TCM symptom entities were categorized into 12 fine-grained semantic labels, and 3 named entity recognition (NER) models were trained to extract TCM symptom terminology from EMRs. In the concatenation stage, standard terms with the same concept as raw terms were identified using a Bidirectional Encoder Representations from Transformers (BERT)–based binary classification model. The standard terms with specific semantic labels were concatenated and reordered according to predefined rules to output structured text, thereby normalizing TCM symptom terminology. Results The proposed STC-NF model achieved an accuracy of 91.4% (180/197) and an F 1 -score of 360 out of 389 (92.5%) on the single-implication test set. For multi-implication terms, STC-NF achieved an accuracy of 84.3% (311/369) and an F 1 -score of 1958 out of 2316 (84.5%), outperforming sequence generation in accuracy by 33.1 percentage points. On the mixed test set containing both single- and multi-implication terms, STC-NF achieved an accuracy of 88.1% (990/1124) and an F 1 -score of 3862 out of 4385 (88.1%), exceeding the best-performing baseline model, multi-task candidate generator (MTCG), by 16.7 and 22.2 percentage points in accuracy and F 1 -score, respectively. Conclusions In this study, we verified that the fine-grained semantic classification and the 2-stage “split-then-concatenate” framework effectively improved performance of named entity recognition and entity alignment, providing an improved approach to normalizing TCM symptom terminology.
Authors
- Hui Ye (ORCID: https://orcid.org/0000-0003-0193-780X)
- Dong Cao (ORCID: https://orcid.org/0000-0003-1563-1983)
- Yuzhu Gao (ORCID: https://orcid.org/0009-0003-3467-1662)
- Chuangan Zhou (ORCID: https://orcid.org/0009-0006-4619-2038)
- Jun Yi (ORCID: https://orcid.org/0009-0003-2093-9779)
- Junyu Yao (ORCID: https://orcid.org/0009-0009-5273-3410)
- Siqi Wang (ORCID: https://orcid.org/0009-0008-1648-1669)
- Jing Tian (ORCID: https://orcid.org/0009-0002-9047-5650)
- Wei Lai (ORCID: https://orcid.org/0009-0000-9491-4249)
- Xingyue Gou (ORCID: https://orcid.org/0009-0007-6380-0231)
Publication Details
- Journal
- JMIR Medical Informatics
- Published
- 2026-09-25
- DOI
- https://doi.org/10.2196/85825
- Primary Topic
- Traditional Chinese Medicine Studies
- Type
- article
- Field-Weighted Citation Impact
- 0.00