Real-World Use of Controlled Terminologies, Ontologies, and Vocabularies for Evidence Generation Across a Large International Observational Network: Challenges and Lessons Learned From a Mixed Method Study
Abstract Background Large-scale international real-world evidence generation benefits from terminology harmonization. Despite widespread adoption of standardized vocabularies, their effective use and long-term sustainability at scale remain poorly understood. Objective This study aimed to examine real-world terminology and code use in data across a federated network of observational data sources. Methods We conducted a 2-part survey of researchers and data owners within the Observational Health Data Sciences and Informatics community on their terminology use, challenges, and needs, accompanied by the analysis of code use across a subset of real-world data sources. Results The survey covered 144 institutions across the United States, the United Kingdom, Europe, Asia, and Africa. Data on terminology use covered 60 sources, including 22 data sources that provided detailed code-level use information. We observed significant variations in terminology use, with 61 out of 89 terminologies used in the data present in less than 10% of the data sources. Code use even after data harmonization was also highly variable: less than 1% (95/742,337) of codes were found in all data sources. Mapping and hierarchy completeness, terminology coverage, versioning, and terminology changes were among the most common challenges. We outlined several of our subsequent process improvements: community contribution and stewardship pipelines, metadata for relationships, and informatics tools for assessment of the impact of terminology change. Conclusions Terminology and coding inconsistencies across observational data sources require a standardized terminology system. Such a system is complex and time-consuming and needs community contribution and informatics solutions for harmonization to be scalable and sustainable. Even with a common reference standard, high heterogeneity of terminology and code use across different observational data sources remains.
Authors
- Dmitry Dymshyts (ORCID: https://orcid.org/0000-0001-8718-0013)
- Christian Reich (ORCID: https://orcid.org/0000-0002-3641-055X)
- Anna Ostropolets (ORCID: https://orcid.org/0000-0002-0847-6682)
- Oleg Zhuk (ORCID: https://orcid.org/0000-0001-9016-7602)
- George Hripcsak (ORCID: https://orcid.org/0000-0003-2664-7614)
- Patrick Ryan (ORCID: https://orcid.org/0000-0002-9727-2138)
- Tatsiana Skuhareuskaya (ORCID: https://orcid.org/0000-0002-4591-6085)
- Alexander Davydov
- Vlad Korsik (ORCID: https://orcid.org/0000-0002-0543-4815)
- Maryia Khitrun (ORCID: https://orcid.org/0009-0007-6232-2311)
Publication Details
- Journal
- JMIR Medical Informatics
- Published
- 2026-09-15
- DOI
- https://doi.org/10.2196/92727
- Primary Topic
- Biomedical Text Mining and Ontologies
- Type
- article
- Field-Weighted Citation Impact
- 0.00