FAIRness assessment of the metadata of omics datasets in online repositories
Abstract Background Current omics repositories aim to facilitate the findability of different types of omics data in order to improve data reuse. Data reuse requires sufficient metadata about how the data was produced and processed. This is part of the “rich metadata” for which the FAIR principles advocate. Adoption of these FAIR principles will make the ever-increasing amount of omics data more Findable, Accessible, Interoperable and Reusable for machines and humans. This is especially important in the rare disease domain, which often relies on the reuse of data for analysis, given the small number of patients with a specific rare disease. We investigated the findability and reusability of omics datasets in two ways; with a generic, automated tool that evaluates the metadata of a resource for compliance with the FAIR principles and with a use-case driven assessment specific to the rare-disease domain. We chose 7 repositories to test for human and machine findability and reusability. Results The metadata of omics datasets presented in the various webpages of repositories on average passed only 2 of the 10 tests done by the generic FAIR Evaluator. The main reason for the failed tests was the lack of embedded metadata on the webpages of the datasets. The second method, where we searched for rare disease datasets, showed that the results were mostly false positives over all repositories. This reduces the precision of the search, which on average was 0.09 with a maximum of 0.3 and a minimum of 0 over the 7 repositories tested. Conclusions FAIRifying metadata on omics repositories will allow users to find and reuse the datasets in the repositories efficiently. To achieve this, based on our analyses, we recommend 5 actions repositories could take. Steps like providing machine-readable metadata using semantic web technologies like RDF, using a fully semantic model and proper use of controlled terminologies would vastly improve the Findability and Reusability of datasets.
Authors
- Ronald Cornet (ORCID: https://orcid.org/0000-0002-1704-5980)
- Eleni Mina (ORCID: https://orcid.org/0000-0002-8972-9206)
- Nirupama Benis (ORCID: https://orcid.org/0000-0002-2101-6154)
Institutions
- Leiden University Medical Center (NL)
- GGD Amsterdam (NL)
- Amsterdam University Medical Centers (NL)
- University of Amsterdam (NL)
Publication Details
- Journal
- Journal of Biomedical Semantics
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1186/s13326-026-00368-3
- Primary Topic
- Research Data Management Practices
- Type
- article
- Field-Weighted Citation Impact
- 0.00