Hierarchical conditional memory for cross-vocabulary medical entity linking

Objective: To introduce HyCoM (Hierarchical Conditional Memory), a neuro-symbolic framework that decouples static ontology retrieval from dynamic contextual reasoning for extremescale cross-vocabulary medical entity linking, with emphasis on inference efficiency and ontological consistency.Methods: HyCoM integrates three components: (1) a hash-based Conditional Memory module providing 𝑂(1) retrieval over 142,679 UMLS concepts; (2) a Context-Aware Gating mechanism fusing neural query representations with memory-retrieved embeddings; and (3) a two-phase Logic Tensor Network (LTN) training protocol that resolves the cross-entropy/LTN objective conflict by pre-training embedding geometry with LTN-only loss before joint finetuning. Evaluation was conducted on the MedPath-MIMIC-III benchmark (4,486 queries, 142,679 concepts) and the MedMentions ST21pv public benchmark (5,904 in-vocabulary test mentions), with BM25, SBERT+FAISS, and SapBERT+FAISS baselines.Results: On the widely used MedMentions ST21pv benchmark, HyCoM achieves P@1=72.6% and MRR=76.1%, outperforming the SapBERT+FAISS baseline (P@1=63.7%, MRR=70.7%) with high precision. Crucially, the two-phase LTN training protocol successfully resolved underlying structural conflicts, reducing the test Hierarchy Violation Rate from 69.1% to1.2%. To examine cross-domain generalization, we conducted a zero-shot stress test on the MedPath-MIMIC-III dataset. While exact-match metrics were uniformly low across all tested algorithms due to distribution mismatch (synthetic templates vs. real clinical fragments), HyCoM successfully retrieved the target or its hierarchical neighbor (H@5) in 52.3% of cases at 40.1ms CPU latency. Following in-domain MIMIC-III adaptation, H@5 accuracy reached96.77% on the holdout set. Finally, a controlled seen/unseen concept split confirms only a −2.7 percentage point P@1 gap when hash entries are withheld, demonstrating robust generalization to truly novel concepts.Conclusion: HyCoM demonstrates that decoupling memorization from reasoning yields an efficient, hierarchically consistent foundation for medical ontology retrieval. 𝑂(1) hash retrieval scales to vocabulary sizes that challenge FAISS-based dense retrieval, while two-phase LTN training simultaneously achieves high precision and ontological constraint satisfaction.

Authors

Institutions

Publication Details

Journal
Informatics in Medicine Unlocked
Published
2026-09-07
DOI
https://doi.org/10.1016/j.imu.2026.101809
Primary Topic
Machine Learning in Healthcare
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Hierarchical conditional memory for cross-vocabulary medical entity linking

Yunguo Yu
Informatics in Medicine Unlocked
Machine Learning in Healthcare
article

Hierarchical conditional memory for cross-vocabulary medical entity linking

Yunguo Yu
article en

Abstract

Objective: To introduce HyCoM (Hierarchical Conditional Memory), a neuro-symbolic framework that decouples static ontology retrieval from dynamic contextual reasoning for extremescale cross-vocabulary medical entity linking, with emphasis on inference efficiency and ontological consistency.Methods: HyCoM integrates three components: (1) a hash-based Conditional Memory module providing 𝑂(1) retrieval over 142,679 UMLS concepts; (2) a Context-Aware Gating mechanism fusing neural query representations with memory-retrieved embeddings; and (3) a two-phase Logic Tensor Network (LTN) training protocol that resolves the cross-entropy/LTN objective conflict by pre-training embedding geometry with LTN-only loss before joint finetuning. Evaluation was conducted on the MedPath-MIMIC-III benchmark (4,486 queries, 142,679 concepts) and the MedMentions ST21pv public benchmark (5,904 in-vocabulary test mentions), with BM25, SBERT+FAISS, and SapBERT+FAISS baselines.Results: On the widely used MedMentions ST21pv benchmark, HyCoM achieves P@1=72.6% and MRR=76.1%, outperforming the SapBERT+FAISS baseline (P@1=63.7%, MRR=70.7%) with high precision. Crucially, the two-phase LTN training protocol successfully resolved underlying structural conflicts, reducing the test Hierarchy Violation Rate from 69.1% to1.2%. To examine cross-domain generalization, we conducted a zero-shot stress test on the MedPath-MIMIC-III dataset. While exact-match metrics were uniformly low across all tested algorithms due to distribution mismatch (synthetic templates vs. real clinical fragments), HyCoM successfully retrieved the target or its hierarchical neighbor (H@5) in 52.3% of cases at 40.1ms CPU latency. Following in-domain MIMIC-III adaptation, H@5 accuracy reached96.77% on the holdout set. Finally, a controlled seen/unseen concept split confirms only a −2.7 percentage point P@1 gap when hash entries are withheld, demonstrating robust generalization to truly novel concepts.Conclusion: HyCoM demonstrates that decoupling memorization from reasoning yields an efficient, hierarchically consistent foundation for medical ontology retrieval. 𝑂(1) hash retrieval scales to vocabulary sizes that challenge FAISS-based dense retrieval, while two-phase LTN training simultaneously achieves high precision and ontological constraint satisfaction.

Informatics in Medicine UnlockedVol. 66
Prototypes (United States) (US)
Peace, Justice and strong institutions
Openalex Percentile: Top 70%
Machine Learning in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.