Knowledge-enhanced LLMs for multilingual biomedical concept normalization: a multilingual benchmarking and behavioral analysis

Abstract Biomedical concept normalization (BCN)—mapping clinical terms to standardized concepts in biomedical terminologies—is a key bottleneck for semantic interoperability across health information systems, particularly in multilingual settings and when terms appear with minimal surrounding context, as is common in coded fields of real-world electronic health record databases. Large language models (LLMs) offer new opportunities to automate BCN, yet their behavior under these conditions remains poorly understood. This study introduces MedLexAlign, a multilingual benchmark comprising 52,011 unique terms mapped to 22,787 biomedical concepts across five European languages (English, French, German, Spanish, and Turkish). It proposes a knowledge-enhanced retrieve-then-rerank pipeline employing discriminative LLMs as dense retrievers and generative LLMs as rerankers, augmented with structured knowledge from the Unified Medical Language System (UMLS). We evaluate eight discriminative LLMs as dense retrievers and six generative LLMs (8B–32B parameters; across different reasoning modes) as rerankers, examining abbreviation handling, synonym sensitivity, and positional bias. The best-performing dense retriever, nemotron, achieved R@1 of 0.57 and R@10 of 0.80, outperforming kalm-gemma (R@1 = 0.56, R@10 = 0.78; p -value = 0.02). For reranking, UMLS-based knowledge enrichment yielded consistent gains: qwen3-32b-r achieved ΔR@1 of +0.10 with enrichment versus +0.05 without ( p -value < 0.001). However, LLMs exhibited sensitivity to lexical surface forms, handled synonymous representations inconsistently, and showed strong primacy bias toward earlier-ranked candidates. These findings demonstrate that structured knowledge enrichment is critical for effective LLM-based multilingual concept normalization, while surface-form sensitivity and positional biases remain important challenges for fully automated clinical pipelines.

Authors

Publication Details

Journal
npj Digital Medicine
Published
2026-09-16
DOI
https://doi.org/10.1038/s41746-026-03224-x
Primary Topic
Biomedical Text Mining and Ontologies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Knowledge-enhanced LLMs for multilingual biomedical concept normalization: a multilingual benchmarking and behavioral analysis

Andreas Walter, Matthias Hüser, David Vicente Alvarez, Anthony Yazdani et al.
npj Digital Medicine
Biomedical Text Mining and Ontologies
article

Knowledge-enhanced LLMs for multilingual biomedical concept normalization: a multilingual benchmarking and behavioral analysis

Andreas Walter, Matthias Hüser, David Vicente Alvarez, Anthony Yazdani, Douglas Teodoro, Hossein Rouhizadeh, Alexandre Vanobberghen, Rui Yang, Huitao Li, Boya Zhang, Nan Liu
article en

Abstract

Abstract Biomedical concept normalization (BCN)—mapping clinical terms to standardized concepts in biomedical terminologies—is a key bottleneck for semantic interoperability across health information systems, particularly in multilingual settings and when terms appear with minimal surrounding context, as is common in coded fields of real-world electronic health record databases. Large language models (LLMs) offer new opportunities to automate BCN, yet their behavior under these conditions remains poorly understood. This study introduces MedLexAlign, a multilingual benchmark comprising 52,011 unique terms mapped to 22,787 biomedical concepts across five European languages (English, French, German, Spanish, and Turkish). It proposes a knowledge-enhanced retrieve-then-rerank pipeline employing discriminative LLMs as dense retrievers and generative LLMs as rerankers, augmented with structured knowledge from the Unified Medical Language System (UMLS). We evaluate eight discriminative LLMs as dense retrievers and six generative LLMs (8B–32B parameters; across different reasoning modes) as rerankers, examining abbreviation handling, synonym sensitivity, and positional bias. The best-performing dense retriever, nemotron, achieved R@1 of 0.57 and R@10 of 0.80, outperforming kalm-gemma (R@1 = 0.56, R@10 = 0.78; p -value = 0.02). For reranking, UMLS-based knowledge enrichment yielded consistent gains: qwen3-32b-r achieved ΔR@1 of +0.10 with enrichment versus +0.05 without ( p -value < 0.001). However, LLMs exhibited sensitivity to lexical surface forms, handled synonymous representations inconsistently, and showed strong primacy bias toward earlier-ranked candidates. These findings demonstrate that structured knowledge enrichment is critical for effective LLM-based multilingual concept normalization, while surface-form sensitivity and positional biases remain important challenges for fully automated clinical pipelines.

npj Digital Medicine
Reduced inequalities
Openalex Percentile: Top 18%
Biomedical Text Mining and Ontologies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.