Improving clinical data standardization in surgical notes through biomedical entity linking with contextual evidence from large language models

Surgical notes contain essential clinical information for postoperative care, yet free-text format and institutional variability limit their use for standardized data representation and secondary analysis. Biomedical entity linking enables mapping of heterogeneous clinical expressions to standardized ontologies such as SNOMED-CT, supporting semantic interoperability. However, existing approaches often rely on predefined mention spans through named entity recognition (NER), which is labor-intensive and may introduce errors. We analyzed 9,051 gastric cancer surgical notes from Seoul National University Hospital. We developed a framework that leverages an open-source large language model (LLM; LLaMA-3.1-8B) to identify contextually relevant text segments, termed evidence spans, which provide cues for ontology-based entity linking. These spans were explicitly marked and used to fine-tune SapBERT, a pretrained embedding-based biomedical encoder. We compared multiple input variants against conventional pipelines and LLM-based approaches, including in-context learning and re-ranking. Incorporating evidence spans improved entity linking performance across metrics, with gains of +2.7 in Recall@1 and +2.2 in mean Average Precision at 3 (mAP@3) compared to raw text inputs. Evidence-guided models outperformed other LLM-based approaches, with additional gains when using the evidence marker token as the pooled representation. Attention analysis indicated that explicit evidence span marking reinforced the model’s focus on ontology-relevant context while reducing attention to irrelevant text. Leveraging LLM-derived contextual evidence improves ontology-based representation of clinical text by enhancing biomedical entity linking. This approach provides a practical strategy for standardizing unstructured surgical notes and supports more reliable secondary use of clinical data in real-world healthcare settings. More broadly, the framework supports mapping of unstructured clinical text to standardized ontologies, contributing to semantic interoperability and enabling downstream secondary use of clinical data.

Authors

Institutions

Publication Details

Journal
BMC Medical Informatics and Decision Making
Published
2026-10-03
DOI
https://doi.org/10.1186/s12911-026-03877-4
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Improving clinical data standardization in surgical notes through biomedical entity linking with contextual evidence from large language models

Siun Kim, Hahyun You, Chaiho Shin, Dareen Eom et al.
BMC Medical Informatics and Decision Making
Topic Modeling
article

Improving clinical data standardization in surgical notes through biomedical entity linking with contextual evidence from large language models

Siun Kim, Hahyun You, Chaiho Shin, Dareen Eom, Hyung-Jin Yoon, Kwangsoo Kim
article en

Abstract

Surgical notes contain essential clinical information for postoperative care, yet free-text format and institutional variability limit their use for standardized data representation and secondary analysis. Biomedical entity linking enables mapping of heterogeneous clinical expressions to standardized ontologies such as SNOMED-CT, supporting semantic interoperability. However, existing approaches often rely on predefined mention spans through named entity recognition (NER), which is labor-intensive and may introduce errors. We analyzed 9,051 gastric cancer surgical notes from Seoul National University Hospital. We developed a framework that leverages an open-source large language model (LLM; LLaMA-3.1-8B) to identify contextually relevant text segments, termed evidence spans, which provide cues for ontology-based entity linking. These spans were explicitly marked and used to fine-tune SapBERT, a pretrained embedding-based biomedical encoder. We compared multiple input variants against conventional pipelines and LLM-based approaches, including in-context learning and re-ranking. Incorporating evidence spans improved entity linking performance across metrics, with gains of +2.7 in Recall@1 and +2.2 in mean Average Precision at 3 (mAP@3) compared to raw text inputs. Evidence-guided models outperformed other LLM-based approaches, with additional gains when using the evidence marker token as the pooled representation. Attention analysis indicated that explicit evidence span marking reinforced the model’s focus on ontology-relevant context while reducing attention to irrelevant text. Leveraging LLM-derived contextual evidence improves ontology-based representation of clinical text by enhancing biomedical entity linking. This approach provides a practical strategy for standardizing unstructured surgical notes and supports more reliable secondary use of clinical data in real-world healthcare settings. More broadly, the framework supports mapping of unstructured clinical text to standardized ontologies, contributing to semantic interoperability and enabling downstream secondary use of clinical data.

BMC Medical Informatics and Decision Making
New Generation University College (ET), Seoul National University Hospital (KR), Kangbuk Samsung Hospital (KR)
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.