Keyphrase identification using minimal labeled data with hierarchical contexts and transfer learning
Interoperable clinical decision support system (CDSS) rules provide a pathway to interoperability, a well-recognized challenge in health information technology. Building an ontology facilitates creating interoperable CDSS rules, and identifying the keyphrases (KP) from the existing literature can be a first step for building an ontology. Ontology construction requires curation by human domain experts (HDE) traditionally, and modern natural language processing (NLP) techniques can be a critical complementary component nevertheless requires human proficiency, consensus, and contextual understanding for data labeling. We present a semi-supervised KP identification framework using a hierarchical-attention BiLSTM-CRF (Hier-Attn-BiLSTM-CRF) with word-, sentence-, and document-level attention. A domain-adapted sciSpacy model was used to generate synthetic labels for bootstrap training, followed by fine-tuning with minimal HDE-labeled data. We then evaluated robustness through component ablation, comparison with fine-tuned biomedical transformers (BioBERT, SciBERT, PubMedBERT) with bootstrap confidence intervals, train-split sensitivity, multi-seed variance analysis, error analysis, and benchmarking against public corpora (KPBioMed, PubMedAKE). The Hier-Attn-BiLSTM-CRF is competitive compared with fine-tuned transformer baselines on the HDE labeled dataset (GS42: ~44 vs. ~46 F1; GS91: 61.3 vs. ~54 F1), through an explicit hierarchical inductive bias over word-, sentence-, and document-level representations, complementing the implicit contextual modeling of the transformer. Controlled ablation identifies gold-standard fine-tuning as the dominant performance lever. Mixing HDE and synthetic labels in 2:100–4:100 ratios improved performance without exhausting the human-labeled set too quickly. Models trained on sparser public corpora transferred poorly to CDSS across all architectures, underscoring the value of in-domain synthetic labels. This feasibility study demonstrates a practical, resource-efficient framework for CDSS KP identification under limited HDE annotation. The contribution lies in integrating established components—domain-adapted synthetic labels, hierarchical attention, and minimal gold-standard fine-tuning—for the CDSS sub-domain. A full downstream evaluation of the role of the NLP pipeline for CDSS ontology curation is the primary next step.
Authors
- Dean F. Sittig (ORCID: https://orcid.org/0000-0001-5811-8915)
- Christian Nøhr (ORCID: https://orcid.org/0000-0003-1299-7365)
- Timothy Law (ORCID: https://orcid.org/0000-0002-3899-072X)
- Hua Min (ORCID: https://orcid.org/0000-0003-2422-0043)
- Yang Gong (ORCID: https://orcid.org/0000-0002-0864-8368)
- Ronald W. Gimbel (ORCID: https://orcid.org/0000-0001-8185-4013)
- Aneesa Weaver
- Arild Faxvaag (ORCID: https://orcid.org/0000-0002-6510-7306)
- Xia Jing (ORCID: https://orcid.org/0000-0002-1916-4588)
- Rohan Goli (ORCID: https://orcid.org/0000-0002-2354-4938)
- Shailesh Alluri
- Keerthana Komatineni
- Lior Rennert
- Nina Hubig
- Paul Biondich
- Adam Wright
- David Robinson
Institutions
- Regenstrief Institute (US)
- George Mason University (US)
- Norwegian University of Science and Technology (NO)
- Ohio University (US)
- Indiana University School of Medicine
- Indiana University – Purdue University Indianapolis (US)
- Clemson University (US)
- Aalborg University (DK)
- Vanderbilt University Medical Center (US)
- The University of Texas Health Science Center at Houston (US)
- University of Cumbria (GB)
Publication Details
- Journal
- BMC Medical Informatics and Decision Making
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1186/s12911-026-03820-7
- Primary Topic
- Advanced Text Analysis Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Institute of General Medical Sciences
- U.S. National Library of Medicine