Automating clinical information retrieval from Finnish electronic health records using large language models

Abstract Clinicians often need to search for patient-specific information within electronic health records (EHRs) containing years of heterogeneous longitudinal documentation, a task that can be time-consuming, cognitively demanding, and prone to oversight. We evaluated 14 locally deployable open-source large language models for structured patient-specific clinical information retrieval using a dataset comprising 1664 expert-annotated question–answer pairs from EHR-derived records of 183 Finnish patients undergoing breast cancer screening, diagnosis, or treatment under fully offline conditions. The records contained predominantly Finnish clinical text with occasional English and Latin terminology. The best-performing model, Llama-3.1-70B, achieved 95.3% accuracy and 97.3% consistency across semantically equivalent question formulations, while Qwen3-30B-A3B-2507 achieved comparable performance with lower measured memory use and latency under the standardized Transformers configuration. For the strongest models, accuracy was similar with and without predefined answer options in the prompt. Calibration varied across architectures, and 4-bit quantization reduced memory requirements while largely preserving accuracy. Medical-domain or Finnish-specialized models showed no systematic advantage over generalist models. Clinical review identified clinically significant errors in 2.9% of outputs, and semantically equivalent question formulations occasionally produced divergent clinical safety outcomes. These findings indicate that locally hosted open-source LLMs may support structured clinical information retrieval from longitudinal EHR-derived records, but clinically significant errors and sensitivity to question formulation remain limiting factors. Source attribution, prospective workflow validation, and human oversight are needed before clinical deployment.

Authors

Institutions

Publication Details

Journal
npj Digital Medicine
Published
2026-09-28
DOI
https://doi.org/10.1038/s41746-026-03282-1
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Automating clinical information retrieval from Finnish electronic health records using large language models

Otso Arponen, Nicole Hernández, KIMMO KASKI, Jaakko Sahlsten et al.
npj Digital Medicine
Topic Modeling
article

Automating clinical information retrieval from Finnish electronic health records using large language models

Otso Arponen, Nicole Hernández, KIMMO KASKI, Jaakko Sahlsten, Mikko Saukkoriipi
article en

Abstract

Abstract Clinicians often need to search for patient-specific information within electronic health records (EHRs) containing years of heterogeneous longitudinal documentation, a task that can be time-consuming, cognitively demanding, and prone to oversight. We evaluated 14 locally deployable open-source large language models for structured patient-specific clinical information retrieval using a dataset comprising 1664 expert-annotated question–answer pairs from EHR-derived records of 183 Finnish patients undergoing breast cancer screening, diagnosis, or treatment under fully offline conditions. The records contained predominantly Finnish clinical text with occasional English and Latin terminology. The best-performing model, Llama-3.1-70B, achieved 95.3% accuracy and 97.3% consistency across semantically equivalent question formulations, while Qwen3-30B-A3B-2507 achieved comparable performance with lower measured memory use and latency under the standardized Transformers configuration. For the strongest models, accuracy was similar with and without predefined answer options in the prompt. Calibration varied across architectures, and 4-bit quantization reduced memory requirements while largely preserving accuracy. Medical-domain or Finnish-specialized models showed no systematic advantage over generalist models. Clinical review identified clinically significant errors in 2.9% of outputs, and semantically equivalent question formulations occasionally produced divergent clinical safety outcomes. These findings indicate that locally hosted open-source LLMs may support structured clinical information retrieval from longitudinal EHR-derived records, but clinically significant errors and sensitivity to question formulation remain limiting factors. Source attribution, prospective workflow validation, and human oversight are needed before clinical deployment.

npj Digital Medicine
Tampere University (FI), Tampere University Hospital (FI), Tampere University (FI), University of Tampere (FI), Aalto University (FI)
Quality Education
Openalex Percentile: Top 76%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.