Performance of pseudonymization methods applied to hospital documents in the context of health data warehouses

Introduction The development of health data warehouses (HDWs) in France raises concerns about personal data protection. Regulations require health research to use anonymized or pseudonymized data. This study evaluates a document pseudonymization tool developed by Assistance Publique - Hopitaux de Paris (AP-HP), which integrates named entity recognition methods, within the NOVA Health cooperation group's project to establish an HDW at Limoges University Hospital. Methods This single-center observational study assessed the performance of the AP-HP pseudonymization algorithm on documents from Limoges University Hospital. The corpus included documents extracted from the Crossway clinical information system software. The study aimed to identify reidentifying labels revealing patient identity. Precision, recall, and F1-score were calculated by document type, hospital department, and patient. Results The AP-HP algorithm achieved an overall F1-score of 79%, with performance varying from 15% to 96% among label types, document formats, and departments. Most false positives and negatives occurred in last names (25% and 21%), first names (23% and 18%), and dates (13% and 14%). The algorithm performed best for emails and phone numbers (0.5% and 0.5%; 4% and 3% for false positives and negatives, respectively). By document type, letters had the best performance (5% false positive, 6% false negative). Per patient, performance was good for up to six documents (80%) but dropped significantly for more than six (−25%). Conclusion While the AP-HP algorithm shows strong label identification, detailed error analysis and contextual adaptation of regular expressions at Limoges University Hospital are necessary to optimize accuracy and ensure full General Data Protection Regulation compliance in HDW data use.

Authors

Institutions

Publication Details

Journal
Journal of Epidemiology and Population Health
Published
2026-09-21
DOI
https://doi.org/10.1016/j.jeph.2026.203790
Primary Topic
Data Quality and Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Performance of pseudonymization methods applied to hospital documents in the context of health data warehouses

Marion Vergonjeanne, Julien Magné, Romain Griffier, Dounia Meddad et al.
Journal of Epidemiology and Population Health
Data Quality and Management
article

Performance of pseudonymization methods applied to hospital documents in the context of health data warehouses

Marion Vergonjeanne, Julien Magné, Romain Griffier, Dounia Meddad, Arnold Didier Ghomsi Talla, Erij Ameur, Caroline Adou-Le Bris, Welba Danwe Danwang, Guillaume Coulaud, Fabien Rastello
article en

Abstract

Introduction The development of health data warehouses (HDWs) in France raises concerns about personal data protection. Regulations require health research to use anonymized or pseudonymized data. This study evaluates a document pseudonymization tool developed by Assistance Publique - Hopitaux de Paris (AP-HP), which integrates named entity recognition methods, within the NOVA Health cooperation group's project to establish an HDW at Limoges University Hospital. Methods This single-center observational study assessed the performance of the AP-HP pseudonymization algorithm on documents from Limoges University Hospital. The corpus included documents extracted from the Crossway clinical information system software. The study aimed to identify reidentifying labels revealing patient identity. Precision, recall, and F1-score were calculated by document type, hospital department, and patient. Results The AP-HP algorithm achieved an overall F1-score of 79%, with performance varying from 15% to 96% among label types, document formats, and departments. Most false positives and negatives occurred in last names (25% and 21%), first names (23% and 18%), and dates (13% and 14%). The algorithm performed best for emails and phone numbers (0.5% and 0.5%; 4% and 3% for false positives and negatives, respectively). By document type, letters had the best performance (5% false positive, 6% false negative). Per patient, performance was good for up to six documents (80%) but dropped significantly for more than six (−25%). Conclusion While the AP-HP algorithm shows strong label identification, detailed error analysis and contextual adaptation of regular expressions at Limoges University Hospital are necessary to optimize accuracy and ensure full General Data Protection Regulation compliance in HDW data use.

Journal of Epidemiology and Population HealthVol. 74(6)
Inserm (FR), Bordeaux Population Health (FR), Hôpital Pellegrin (FR), Institut de Recherche pour le Développement (FR), Université de Limoges (FR)
Partnerships for the goals
Openalex Percentile: Top 7%
Data Quality and Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.