Sequence‐identical microbial genomes with disparate taxonomy: A large‐scale redundancy analysis in NCBI genome database

Public microbial genome repositories (such as NCBI) contain thousands of sequence-identical genome assemblies assigned to different taxonomic lineages, potentially inflating microbial diversity and introducing bias into comparative genomics, metagenomics, and pathogen surveillance. By systematically screening 729,559 NCBI prokaryotic genomes using assembly metadata filtering followed by MD5 hash-based sequence validation, we identified 3334 sets of completely identical genome assemblies with taxonomic inconsistencies. We propose a simple redundancy-aware genome submission workflow that combines assembly-level screening with cryptographic sequence identity checks to flag identical submissions, harmonize metadata, and strengthen the accuracy, transparency, and reliability of public genomic databases.

Authors

Institutions

Publication Details

Journal
iMetaOmics.
Published
2026-09-15
DOI
https://doi.org/10.1002/imo2.70140
Primary Topic
Genomics and Phylogenetic Studies
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Sequence‐identical microbial genomes with disparate taxonomy: A large‐scale redundancy analysis in NCBI genome database

Gaurav Sharma, Niharika Saraf
iMetaOmics.
Genomics and Phylogenetic Studies
article

Sequence‐identical microbial genomes with disparate taxonomy: A large‐scale redundancy analysis in NCBI genome database

Gaurav Sharma, Niharika Saraf
article en

Abstract

Public microbial genome repositories (such as NCBI) contain thousands of sequence-identical genome assemblies assigned to different taxonomic lineages, potentially inflating microbial diversity and introducing bias into comparative genomics, metagenomics, and pathogen surveillance. By systematically screening 729,559 NCBI prokaryotic genomes using assembly metadata filtering followed by MD5 hash-based sequence validation, we identified 3334 sets of completely identical genome assemblies with taxonomic inconsistencies. We propose a simple redundancy-aware genome submission workflow that combines assembly-level screening with cryptographic sequence identity checks to flag identical submissions, harmonize metadata, and strengthen the accuracy, transparency, and reliability of public genomic databases.

iMetaOmics.
Indian Institute of Technology Hyderabad (IN)
Science and Engineering Research Board
Openalex Percentile: Top 18%
Genomics and Phylogenetic Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.