Open but Not Whole: A Review of the Value and Limitations of Public Databases in Research
A narrative review for early-career researchers, students and practitioners who build their work on data they did not collect themselves. It opens by sorting public databases into five kinds (primary repositories, aggregators, cohort and clinical resources, registries, and official and survey statistics). It then shows what they have made possible: the Protein Data Bank's role in AlphaFold, real-time genomic surveillance during COVID-19, reproducible analysis, and lower-cost research in low-resource settings. Next it sets out eight ways a database can differ systematically from the reality it appears to describe, each illustrated with published cases: GBIF records clustering near roads and cities, UK Biobank's healthier-than-average participants, fungal sequences filed under the wrong species, gene names silently converted to dates by spreadsheets, research data that disappear as papers age, re-identification attacks that force deliberate gaps in privacy-protected releases, and the retracted Surgisphere-based COVID-19 study. The paper ends with ten practical checks for people who use public data and a set of practices for those who contribute to it.
Authors
- Raunak Sharma (ORCID: https://orcid.org/0000-0003-4173-7714)
- Sreemoyee Chakraborty (ORCID: https://orcid.org/0000-0001-5180-156X)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23181391
- Primary Topic
- Research Data Management Practices
- Type
- preprint