Real-World Use of Controlled Terminologies, Ontologies, and Vocabularies for Evidence Generation Across a Large International Observational Network: Challenges and Lessons Learned From a Mixed Method Study

Abstract Background Large-scale international real-world evidence generation benefits from terminology harmonization. Despite widespread adoption of standardized vocabularies, their effective use and long-term sustainability at scale remain poorly understood. Objective This study aimed to examine real-world terminology and code use in data across a federated network of observational data sources. Methods We conducted a 2-part survey of researchers and data owners within the Observational Health Data Sciences and Informatics community on their terminology use, challenges, and needs, accompanied by the analysis of code use across a subset of real-world data sources. Results The survey covered 144 institutions across the United States, the United Kingdom, Europe, Asia, and Africa. Data on terminology use covered 60 sources, including 22 data sources that provided detailed code-level use information. We observed significant variations in terminology use, with 61 out of 89 terminologies used in the data present in less than 10% of the data sources. Code use even after data harmonization was also highly variable: less than 1% (95/742,337) of codes were found in all data sources. Mapping and hierarchy completeness, terminology coverage, versioning, and terminology changes were among the most common challenges. We outlined several of our subsequent process improvements: community contribution and stewardship pipelines, metadata for relationships, and informatics tools for assessment of the impact of terminology change. Conclusions Terminology and coding inconsistencies across observational data sources require a standardized terminology system. Such a system is complex and time-consuming and needs community contribution and informatics solutions for harmonization to be scalable and sustainable. Even with a common reference standard, high heterogeneity of terminology and code use across different observational data sources remains.

Authors

Publication Details

Journal
JMIR Medical Informatics
Published
2026-09-15
DOI
https://doi.org/10.2196/92727
Primary Topic
Biomedical Text Mining and Ontologies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Real-World Use of Controlled Terminologies, Ontologies, and Vocabularies for Evidence Generation Across a Large International Observational Network: Challenges and Lessons Learned From a Mixed Method Study

Dmitry Dymshyts, Christian Reich, Anna Ostropolets, Oleg Zhuk et al.
JMIR Medical Informatics
Biomedical Text Mining and Ontologies
article

Real-World Use of Controlled Terminologies, Ontologies, and Vocabularies for Evidence Generation Across a Large International Observational Network: Challenges and Lessons Learned From a Mixed Method Study

Dmitry Dymshyts, Christian Reich, Anna Ostropolets, Oleg Zhuk, George Hripcsak, Patrick Ryan, Tatsiana Skuhareuskaya, Alexander Davydov, Vlad Korsik, Maryia Khitrun
article en

Abstract

Abstract Background Large-scale international real-world evidence generation benefits from terminology harmonization. Despite widespread adoption of standardized vocabularies, their effective use and long-term sustainability at scale remain poorly understood. Objective This study aimed to examine real-world terminology and code use in data across a federated network of observational data sources. Methods We conducted a 2-part survey of researchers and data owners within the Observational Health Data Sciences and Informatics community on their terminology use, challenges, and needs, accompanied by the analysis of code use across a subset of real-world data sources. Results The survey covered 144 institutions across the United States, the United Kingdom, Europe, Asia, and Africa. Data on terminology use covered 60 sources, including 22 data sources that provided detailed code-level use information. We observed significant variations in terminology use, with 61 out of 89 terminologies used in the data present in less than 10% of the data sources. Code use even after data harmonization was also highly variable: less than 1% (95/742,337) of codes were found in all data sources. Mapping and hierarchy completeness, terminology coverage, versioning, and terminology changes were among the most common challenges. We outlined several of our subsequent process improvements: community contribution and stewardship pipelines, metadata for relationships, and informatics tools for assessment of the impact of terminology change. Conclusions Terminology and coding inconsistencies across observational data sources require a standardized terminology system. Such a system is complex and time-consuming and needs community contribution and informatics solutions for harmonization to be scalable and sustainable. Even with a common reference standard, high heterogeneity of terminology and code use across different observational data sources remains.

JMIR Medical InformaticsVol. 14
Openalex Percentile: Top 18%
Biomedical Text Mining and Ontologies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.