AI-Based Structured Information Extraction From Synthetic Nursing Handover Transcripts: Comparative Evaluation of Large Language Models

Abstract Background Clinical handover is the process during which responsibility and accountability for care are transferred between clinicians. AI has the potential to improve the reliability and completeness of clinical handover by helping clinicians detect predefined content areas that have been communicated, identify explicit information gaps, and prompt clarification before responsibility is transferred. Objective This study evaluated the performance of several large language models and prompt optimization strategies for structured information extraction of synthetic nursing handover transcripts. Methods Two registered nurses independently annotated a dataset of 203 synthetic handover transcripts to produce consensus labels for information extraction tasks. Tasks included (1) labeling spans of text into SBAR (Situation, Background, Assessment, Recommendation) categories, (2) content detection to determine if specific pieces of information were communicated, and (3) labeling spans of text that communicated information using uncertain terms that included a subtask for identifying unknown facts. Baseline and Genetic-Pareto (GEPA)–optimized prompts were compared for the GPT-5.2, GPT-5-nano, and MedGemma 27B large language models. Additionally, the LangExtract framework was evaluated for span-extraction tasks. Results The GPT-5.2–optimized model achieved a micro –F 1 -score of 0.85 (95% CI 0.83‐0.88) for content detection, an absolute improvement of +0.08 compared with the matched baseline. GPT-5-nano also performed better after optimization for content detection (micro –F 1 -score 0.81, 95% CI 0.78‐0.84), suggesting that this structured task was not limited to the highest-capacity model. For SBAR span extraction, GPT-5.2 with prompt optimization achieved a micro –F 1 -score of 0.76 (95% CI 0.72‐0.79), improving by +0.24 compared with baseline and exceeding LangExtract; GPT-5-nano also improved to a micro –F 1 -score of 0.69 (95% CI 0.66‐0.72). Broad uncertainty-span extraction remained comparatively weak despite prompt optimization (micro –F 1 -score 0.41, 95% CI 0.33‐0.48; absolute improvement +0.06). In contrast, explicit unknown-fact extraction was more accurate with GPT-5.2 (micro –F 1 -score 0.84, 95% CI 0.63‐1.00), GPT-5-nano (micro –F 1 -score, 0.84 95% CI 0.63‐1.00), and MedGemma 27B (micro –F 1 -score 0.80, 95% CI 0.63‐1.00). Genetic-Pareto–optimized prompts outperformed the LangExtract approach across each span-extraction task. Conclusions Prompt optimization improved matched-model point estimates, with the highest performance observed for predefined content detection and SBAR span extraction. Broad uncertainty extraction remained less accurate than the narrower unknown-fact task. These technical results do not establish clinical effectiveness, safety, or readiness for real-time use. Validation using authentic nursing handover communication and prospective evaluation in clinical workflows are required before clinical application.

Authors

Publication Details

Journal
JMIR Nursing
Published
2026-10-05
DOI
https://doi.org/10.2196/106133
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

AI-Based Structured Information Extraction From Synthetic Nursing Handover Transcripts: Comparative Evaluation of Large Language Models

Aaron Conway, Jessica Schluter, Tim Miller, Adriana Hada et al.
JMIR Nursing
Artificial Intelligence in Healthcare and Education
article

AI-Based Structured Information Extraction From Synthetic Nursing Handover Transcripts: Comparative Evaluation of Large Language Models

Aaron Conway, Jessica Schluter, Tim Miller, Adriana Hada, Andrew Teodorczuk, Hui Grace Xu, Ken Donald, Dan Lowden
article en

Abstract

Abstract Background Clinical handover is the process during which responsibility and accountability for care are transferred between clinicians. AI has the potential to improve the reliability and completeness of clinical handover by helping clinicians detect predefined content areas that have been communicated, identify explicit information gaps, and prompt clarification before responsibility is transferred. Objective This study evaluated the performance of several large language models and prompt optimization strategies for structured information extraction of synthetic nursing handover transcripts. Methods Two registered nurses independently annotated a dataset of 203 synthetic handover transcripts to produce consensus labels for information extraction tasks. Tasks included (1) labeling spans of text into SBAR (Situation, Background, Assessment, Recommendation) categories, (2) content detection to determine if specific pieces of information were communicated, and (3) labeling spans of text that communicated information using uncertain terms that included a subtask for identifying unknown facts. Baseline and Genetic-Pareto (GEPA)–optimized prompts were compared for the GPT-5.2, GPT-5-nano, and MedGemma 27B large language models. Additionally, the LangExtract framework was evaluated for span-extraction tasks. Results The GPT-5.2–optimized model achieved a micro –F 1 -score of 0.85 (95% CI 0.83‐0.88) for content detection, an absolute improvement of +0.08 compared with the matched baseline. GPT-5-nano also performed better after optimization for content detection (micro –F 1 -score 0.81, 95% CI 0.78‐0.84), suggesting that this structured task was not limited to the highest-capacity model. For SBAR span extraction, GPT-5.2 with prompt optimization achieved a micro –F 1 -score of 0.76 (95% CI 0.72‐0.79), improving by +0.24 compared with baseline and exceeding LangExtract; GPT-5-nano also improved to a micro –F 1 -score of 0.69 (95% CI 0.66‐0.72). Broad uncertainty-span extraction remained comparatively weak despite prompt optimization (micro –F 1 -score 0.41, 95% CI 0.33‐0.48; absolute improvement +0.06). In contrast, explicit unknown-fact extraction was more accurate with GPT-5.2 (micro –F 1 -score 0.84, 95% CI 0.63‐1.00), GPT-5-nano (micro –F 1 -score, 0.84 95% CI 0.63‐1.00), and MedGemma 27B (micro –F 1 -score 0.80, 95% CI 0.63‐1.00). Genetic-Pareto–optimized prompts outperformed the LangExtract approach across each span-extraction task. Conclusions Prompt optimization improved matched-model point estimates, with the highest performance observed for predefined content detection and SBAR span extraction. Broad uncertainty extraction remained less accurate than the narrower unknown-fact task. These technical results do not establish clinical effectiveness, safety, or readiness for real-time use. Validation using authentic nursing handover communication and prospective evaluation in clinical workflows are required before clinical application.

JMIR NursingVol. 9
Openalex Percentile: Top 18%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.