Keyphrase identification using minimal labeled data with hierarchical contexts and transfer learning

Interoperable clinical decision support system (CDSS) rules provide a pathway to interoperability, a well-recognized challenge in health information technology. Building an ontology facilitates creating interoperable CDSS rules, and identifying the keyphrases (KP) from the existing literature can be a first step for building an ontology. Ontology construction requires curation by human domain experts (HDE) traditionally, and modern natural language processing (NLP) techniques can be a critical complementary component nevertheless requires human proficiency, consensus, and contextual understanding for data labeling. We present a semi-supervised KP identification framework using a hierarchical-attention BiLSTM-CRF (Hier-Attn-BiLSTM-CRF) with word-, sentence-, and document-level attention. A domain-adapted sciSpacy model was used to generate synthetic labels for bootstrap training, followed by fine-tuning with minimal HDE-labeled data. We then evaluated robustness through component ablation, comparison with fine-tuned biomedical transformers (BioBERT, SciBERT, PubMedBERT) with bootstrap confidence intervals, train-split sensitivity, multi-seed variance analysis, error analysis, and benchmarking against public corpora (KPBioMed, PubMedAKE). The Hier-Attn-BiLSTM-CRF is competitive compared with fine-tuned transformer baselines on the HDE labeled dataset (GS42: ~44 vs. ~46 F1; GS91: 61.3 vs. ~54 F1), through an explicit hierarchical inductive bias over word-, sentence-, and document-level representations, complementing the implicit contextual modeling of the transformer. Controlled ablation identifies gold-standard fine-tuning as the dominant performance lever. Mixing HDE and synthetic labels in 2:100–4:100 ratios improved performance without exhausting the human-labeled set too quickly. Models trained on sparser public corpora transferred poorly to CDSS across all architectures, underscoring the value of in-domain synthetic labels. This feasibility study demonstrates a practical, resource-efficient framework for CDSS KP identification under limited HDE annotation. The contribution lies in integrating established components—domain-adapted synthetic labels, hierarchical attention, and minimal gold-standard fine-tuning—for the CDSS sub-domain. A full downstream evaluation of the role of the NLP pipeline for CDSS ontology curation is the primary next step.

Authors

Institutions

Publication Details

Journal
BMC Medical Informatics and Decision Making
Published
2026-09-18
DOI
https://doi.org/10.1186/s12911-026-03820-7
Primary Topic
Advanced Text Analysis Techniques
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Keyphrase identification using minimal labeled data with hierarchical contexts and transfer learning

Dean F. Sittig, Christian Nøhr, Timothy Law, Hua Min et al.
BMC Medical Informatics and Decision Making
Advanced Text Analysis Techniques
article

Keyphrase identification using minimal labeled data with hierarchical contexts and transfer learning

Dean F. Sittig, Christian Nøhr, Timothy Law, Hua Min, Yang Gong, Ronald W. Gimbel, Aneesa Weaver, Arild Faxvaag, Xia Jing, Rohan Goli, Shailesh Alluri, Keerthana Komatineni, Lior Rennert, Nina Hubig, Paul Biondich, Adam Wright, David Robinson
article en

Abstract

Interoperable clinical decision support system (CDSS) rules provide a pathway to interoperability, a well-recognized challenge in health information technology. Building an ontology facilitates creating interoperable CDSS rules, and identifying the keyphrases (KP) from the existing literature can be a first step for building an ontology. Ontology construction requires curation by human domain experts (HDE) traditionally, and modern natural language processing (NLP) techniques can be a critical complementary component nevertheless requires human proficiency, consensus, and contextual understanding for data labeling. We present a semi-supervised KP identification framework using a hierarchical-attention BiLSTM-CRF (Hier-Attn-BiLSTM-CRF) with word-, sentence-, and document-level attention. A domain-adapted sciSpacy model was used to generate synthetic labels for bootstrap training, followed by fine-tuning with minimal HDE-labeled data. We then evaluated robustness through component ablation, comparison with fine-tuned biomedical transformers (BioBERT, SciBERT, PubMedBERT) with bootstrap confidence intervals, train-split sensitivity, multi-seed variance analysis, error analysis, and benchmarking against public corpora (KPBioMed, PubMedAKE). The Hier-Attn-BiLSTM-CRF is competitive compared with fine-tuned transformer baselines on the HDE labeled dataset (GS42: ~44 vs. ~46 F1; GS91: 61.3 vs. ~54 F1), through an explicit hierarchical inductive bias over word-, sentence-, and document-level representations, complementing the implicit contextual modeling of the transformer. Controlled ablation identifies gold-standard fine-tuning as the dominant performance lever. Mixing HDE and synthetic labels in 2:100–4:100 ratios improved performance without exhausting the human-labeled set too quickly. Models trained on sparser public corpora transferred poorly to CDSS across all architectures, underscoring the value of in-domain synthetic labels. This feasibility study demonstrates a practical, resource-efficient framework for CDSS KP identification under limited HDE annotation. The contribution lies in integrating established components—domain-adapted synthetic labels, hierarchical attention, and minimal gold-standard fine-tuning—for the CDSS sub-domain. A full downstream evaluation of the role of the NLP pipeline for CDSS ontology curation is the primary next step.

BMC Medical Informatics and Decision Making
Regenstrief Institute (US), George Mason University (US), Norwegian University of Science and Technology (NO), Ohio University (US), Indiana University School of Medicine, Indiana University – Purdue University Indianapolis (US), Clemson University (US), Aalborg University (DK), Vanderbilt University Medical Center (US), The University of Texas Health Science Center at Houston (US), University of Cumbria (GB)
National Institute of General Medical Sciences, U.S. National Library of Medicine
Openalex Percentile: Top 9%
Advanced Text Analysis Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.