FAIRness assessment of the metadata of omics datasets in online repositories

Abstract Background Current omics repositories aim to facilitate the findability of different types of omics data in order to improve data reuse. Data reuse requires sufficient metadata about how the data was produced and processed. This is part of the “rich metadata” for which the FAIR principles advocate. Adoption of these FAIR principles will make the ever-increasing amount of omics data more Findable, Accessible, Interoperable and Reusable for machines and humans. This is especially important in the rare disease domain, which often relies on the reuse of data for analysis, given the small number of patients with a specific rare disease. We investigated the findability and reusability of omics datasets in two ways; with a generic, automated tool that evaluates the metadata of a resource for compliance with the FAIR principles and with a use-case driven assessment specific to the rare-disease domain. We chose 7 repositories to test for human and machine findability and reusability. Results The metadata of omics datasets presented in the various webpages of repositories on average passed only 2 of the 10 tests done by the generic FAIR Evaluator. The main reason for the failed tests was the lack of embedded metadata on the webpages of the datasets. The second method, where we searched for rare disease datasets, showed that the results were mostly false positives over all repositories. This reduces the precision of the search, which on average was 0.09 with a maximum of 0.3 and a minimum of 0 over the 7 repositories tested. Conclusions FAIRifying metadata on omics repositories will allow users to find and reuse the datasets in the repositories efficiently. To achieve this, based on our analyses, we recommend 5 actions repositories could take. Steps like providing machine-readable metadata using semantic web technologies like RDF, using a fully semantic model and proper use of controlled terminologies would vastly improve the Findability and Reusability of datasets.

Authors

Institutions

Publication Details

Journal
Journal of Biomedical Semantics
Published
2026-09-15
DOI
https://doi.org/10.1186/s13326-026-00368-3
Primary Topic
Research Data Management Practices
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

FAIRness assessment of the metadata of omics datasets in online repositories

Ronald Cornet, Eleni Mina, Nirupama Benis
Journal of Biomedical Semantics
Research Data Management Practices
article

FAIRness assessment of the metadata of omics datasets in online repositories

Ronald Cornet, Eleni Mina, Nirupama Benis
article en

Abstract

Abstract Background Current omics repositories aim to facilitate the findability of different types of omics data in order to improve data reuse. Data reuse requires sufficient metadata about how the data was produced and processed. This is part of the “rich metadata” for which the FAIR principles advocate. Adoption of these FAIR principles will make the ever-increasing amount of omics data more Findable, Accessible, Interoperable and Reusable for machines and humans. This is especially important in the rare disease domain, which often relies on the reuse of data for analysis, given the small number of patients with a specific rare disease. We investigated the findability and reusability of omics datasets in two ways; with a generic, automated tool that evaluates the metadata of a resource for compliance with the FAIR principles and with a use-case driven assessment specific to the rare-disease domain. We chose 7 repositories to test for human and machine findability and reusability. Results The metadata of omics datasets presented in the various webpages of repositories on average passed only 2 of the 10 tests done by the generic FAIR Evaluator. The main reason for the failed tests was the lack of embedded metadata on the webpages of the datasets. The second method, where we searched for rare disease datasets, showed that the results were mostly false positives over all repositories. This reduces the precision of the search, which on average was 0.09 with a maximum of 0.3 and a minimum of 0 over the 7 repositories tested. Conclusions FAIRifying metadata on omics repositories will allow users to find and reuse the datasets in the repositories efficiently. To achieve this, based on our analyses, we recommend 5 actions repositories could take. Steps like providing machine-readable metadata using semantic web technologies like RDF, using a fully semantic model and proper use of controlled terminologies would vastly improve the Findability and Reusability of datasets.

Journal of Biomedical Semantics
Leiden University Medical Center (NL), GGD Amsterdam (NL), Amsterdam University Medical Centers (NL), University of Amsterdam (NL)
Openalex Percentile: Top 4%
Research Data Management Practices
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.