Keeping it real: a hybrid approach combining de-identification and synthetic text generation for privacy preserving clinical data sharing
Clinical notes are a foundational resource for healthcare research and innovation, yet their dissemination is constrained by stringent privacy regulations such as HIPAA and GDPR. Conventional de-identification pipelines aim to remove protected health information (PHI), but they remain imperfect. In contrast, fully synthetic clinical notes generation struggles to preserve clinical detail and comprehensive content coverage. To address these complementary limitations, we introduce a hybrid framework that combines both approaches to generate privacy-aware and clinically informative notes. The proposed hybrid framework retains de-identification as a preprocessing step and introduces a clinically aware filtration stage that preserves not only general non-sensitive tokens, but also biomedical entities, clinically meaningful quantities, and UMLS-grounded concept spans. This richer retained context is then provided to a generative model to restore narrative coherence and improve clinical fidelity. We evaluated two hybrid variants against two standalone de-identification baselines by measuring residual PHI leakage. To assess the linguistic quality and content fidelity of the generated clinical notes, we used a combination of established natural language generation metrics, including one content-retention measure (percentage of words retained), two lexical-overlap metrics (BLEU and ROUGE), and a semantic-similarity measure (cosine similarity). The two standalone de-identification baselines left 855 and 193 residual PHI occurrences, corresponding to removal rates of 94.9% and 98.9%, respectively. In contrast, the first hybrid variant reduced leakage to only 12 occurrences, achieving 99.9% recall and eliminating 93.8% to 98.6% of the post-de-identification PHI leakage across both de-identification baselines, while retaining approximately 50% of the original content. The second hybrid variant, augmented with UMLS-based concept retention, further improved fidelity by increasing content retention by approximately 10% to about 60% and achieved a cosine semantic similarity score of 0.835, while eliminating 79.8% to 92.9% of the post-de-identification PHI leakage. Across experiments, the hybrid variants substantially improved information retention and content fidelity relative to synthetic note generation. The proposed approach offers substantial privacy improvement over standalone de-identification while preserving more clinically meaningful content than fully synthetic approaches. These findings indicate that hybrid clinical note generation is a promising direction for privacy-conscious and utility-preserving clinical text release. Further research is needed to eliminate residual PHI and to assess compliance under formal regulatory standards.
Authors
- Atiquer Rahman Sarkar (ORCID: https://orcid.org/0000-0002-6909-2930)
- Fatima Jahan Sarmin
- Xiaoqian Jiang (ORCID: https://orcid.org/0000-0001-9933-2205)
- Yao-Shun Chuang (ORCID: https://orcid.org/0000-0001-9582-9222)
- Noman Mohammed (ORCID: https://orcid.org/0000-0001-8547-9951)
Institutions
- University of Toronto (CA)
- The University of Texas Health Science Center (US)
- University of Manitoba (CA)
- The University of Texas Health Science Center at Houston (US)
Publication Details
- Journal
- BMC Medical Informatics and Decision Making
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1186/s12911-026-03811-8
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00