Keeping it real: a hybrid approach combining de-identification and synthetic text generation for privacy preserving clinical data sharing

Clinical notes are a foundational resource for healthcare research and innovation, yet their dissemination is constrained by stringent privacy regulations such as HIPAA and GDPR. Conventional de-identification pipelines aim to remove protected health information (PHI), but they remain imperfect. In contrast, fully synthetic clinical notes generation struggles to preserve clinical detail and comprehensive content coverage. To address these complementary limitations, we introduce a hybrid framework that combines both approaches to generate privacy-aware and clinically informative notes. The proposed hybrid framework retains de-identification as a preprocessing step and introduces a clinically aware filtration stage that preserves not only general non-sensitive tokens, but also biomedical entities, clinically meaningful quantities, and UMLS-grounded concept spans. This richer retained context is then provided to a generative model to restore narrative coherence and improve clinical fidelity. We evaluated two hybrid variants against two standalone de-identification baselines by measuring residual PHI leakage. To assess the linguistic quality and content fidelity of the generated clinical notes, we used a combination of established natural language generation metrics, including one content-retention measure (percentage of words retained), two lexical-overlap metrics (BLEU and ROUGE), and a semantic-similarity measure (cosine similarity). The two standalone de-identification baselines left 855 and 193 residual PHI occurrences, corresponding to removal rates of 94.9% and 98.9%, respectively. In contrast, the first hybrid variant reduced leakage to only 12 occurrences, achieving 99.9% recall and eliminating 93.8% to 98.6% of the post-de-identification PHI leakage across both de-identification baselines, while retaining approximately 50% of the original content. The second hybrid variant, augmented with UMLS-based concept retention, further improved fidelity by increasing content retention by approximately 10% to about 60% and achieved a cosine semantic similarity score of 0.835, while eliminating 79.8% to 92.9% of the post-de-identification PHI leakage. Across experiments, the hybrid variants substantially improved information retention and content fidelity relative to synthetic note generation. The proposed approach offers substantial privacy improvement over standalone de-identification while preserving more clinically meaningful content than fully synthetic approaches. These findings indicate that hybrid clinical note generation is a promising direction for privacy-conscious and utility-preserving clinical text release. Further research is needed to eliminate residual PHI and to assess compliance under formal regulatory standards.

Authors

Institutions

Publication Details

Journal
BMC Medical Informatics and Decision Making
Published
2026-09-15
DOI
https://doi.org/10.1186/s12911-026-03811-8
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Keeping it real: a hybrid approach combining de-identification and synthetic text generation for privacy preserving clinical data sharing

Atiquer Rahman Sarkar, Fatima Jahan Sarmin, Xiaoqian Jiang, Yao-Shun Chuang et al.
BMC Medical Informatics and Decision Making
Topic Modeling
article

Keeping it real: a hybrid approach combining de-identification and synthetic text generation for privacy preserving clinical data sharing

Atiquer Rahman Sarkar, Fatima Jahan Sarmin, Xiaoqian Jiang, Yao-Shun Chuang, Noman Mohammed
article en

Abstract

Clinical notes are a foundational resource for healthcare research and innovation, yet their dissemination is constrained by stringent privacy regulations such as HIPAA and GDPR. Conventional de-identification pipelines aim to remove protected health information (PHI), but they remain imperfect. In contrast, fully synthetic clinical notes generation struggles to preserve clinical detail and comprehensive content coverage. To address these complementary limitations, we introduce a hybrid framework that combines both approaches to generate privacy-aware and clinically informative notes. The proposed hybrid framework retains de-identification as a preprocessing step and introduces a clinically aware filtration stage that preserves not only general non-sensitive tokens, but also biomedical entities, clinically meaningful quantities, and UMLS-grounded concept spans. This richer retained context is then provided to a generative model to restore narrative coherence and improve clinical fidelity. We evaluated two hybrid variants against two standalone de-identification baselines by measuring residual PHI leakage. To assess the linguistic quality and content fidelity of the generated clinical notes, we used a combination of established natural language generation metrics, including one content-retention measure (percentage of words retained), two lexical-overlap metrics (BLEU and ROUGE), and a semantic-similarity measure (cosine similarity). The two standalone de-identification baselines left 855 and 193 residual PHI occurrences, corresponding to removal rates of 94.9% and 98.9%, respectively. In contrast, the first hybrid variant reduced leakage to only 12 occurrences, achieving 99.9% recall and eliminating 93.8% to 98.6% of the post-de-identification PHI leakage across both de-identification baselines, while retaining approximately 50% of the original content. The second hybrid variant, augmented with UMLS-based concept retention, further improved fidelity by increasing content retention by approximately 10% to about 60% and achieved a cosine semantic similarity score of 0.835, while eliminating 79.8% to 92.9% of the post-de-identification PHI leakage. Across experiments, the hybrid variants substantially improved information retention and content fidelity relative to synthetic note generation. The proposed approach offers substantial privacy improvement over standalone de-identification while preserving more clinically meaningful content than fully synthetic approaches. These findings indicate that hybrid clinical note generation is a promising direction for privacy-conscious and utility-preserving clinical text release. Further research is needed to eliminate residual PHI and to assess compliance under formal regulatory standards.

BMC Medical Informatics and Decision Making
University of Toronto (CA), The University of Texas Health Science Center (US), University of Manitoba (CA), The University of Texas Health Science Center at Houston (US)
Industry, innovation and infrastructure
Openalex Percentile: Top 8%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.