Study on the joint application of multi-database to improve the ability of high-throughput sequencing of forensic diatoms

Objective: to apply the method of single or combined use of multiple databases to improve the accuracy of diatom classification based on high-throughput sequencing and to strengthen the accuracy of diagnosis and inference of drowning location of corpses in forensic practice. Methods: high throughput sequencing was carried out on the water samples of three major freshwater rivers in Hainan Province, China. Silva ribosomal RNA database, American National Biotechnology Information Center database, Diat.barcode diatom database and Chinese national GenBank database were selected to identify diatom species alone or in combination with the results of high-throughput sequencing. Combined with Wekits software, MEGA11 software and Hebert10 principle, the species and genus of diatoms were classified by sequence alignment, phylogenetic tree construction and interspecific distance calculation. The advantages and disadvantages of using database alone or in combination in the identification of diatom species were evaluated. Results: NCBI has an advantage in identifying diatom species, but the accuracy and completeness of information are not as good as Diat.barcode; although the number of CNGB genus-level identification is not as good as that of Silva, it is similar to the number of species-level sequences identified by Silva, and CNGB contains some sequences that can not be retrieved at diatom species level in other databases. When using the method of database combination, the number of species-level identification sequences obtained by all four databases is close to that of Silva+NCBI+CNGB combination, but the Silva+NCBI+CNGB combination does not use Diat.barcode database, which can reduce the requirement of bioinformatics knowledge. The more accurate identification results of diatom species obtained by database combination are helpful to distinguish different sampling sites, which provides a basis for the application of diatom species differences in forensic medicine to infer the site of drowning. Conclusion: different database combinations can improve the accuracy of diatom species identification based on high-throughput sequencing results, although it depends on the integrity of the database to some extent, but more species level sequences help to improve the ability to distinguish sampling points. In this study, we recommend using Silva+NCBI+CNGB database combination to identify diatom species with high-throughput sequencing results. The high accuracy of diatom classification will contribute to the diagnosis of drowning and the forensic application of inferring suspected drowning sites.

Authors

Publication Details

Journal
China National GeneBank DataBase
Published
2026-08-27
DOI
https://doi.org/10.26036/cnp0006175
Primary Topic
Injury Epidemiology and Prevention
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Study on the joint application of multi-database to improve the ability of high-throughput sequencing of forensic diatoms

覃小诗(Qin Xiaoshi)
China National GeneBank DataBase
Injury Epidemiology and Prevention
article

Study on the joint application of multi-database to improve the ability of high-throughput sequencing of forensic diatoms

覃小诗(Qin Xiaoshi)
article en

Abstract

Objective: to apply the method of single or combined use of multiple databases to improve the accuracy of diatom classification based on high-throughput sequencing and to strengthen the accuracy of diagnosis and inference of drowning location of corpses in forensic practice. Methods: high throughput sequencing was carried out on the water samples of three major freshwater rivers in Hainan Province, China. Silva ribosomal RNA database, American National Biotechnology Information Center database, Diat.barcode diatom database and Chinese national GenBank database were selected to identify diatom species alone or in combination with the results of high-throughput sequencing. Combined with Wekits software, MEGA11 software and Hebert10 principle, the species and genus of diatoms were classified by sequence alignment, phylogenetic tree construction and interspecific distance calculation. The advantages and disadvantages of using database alone or in combination in the identification of diatom species were evaluated. Results: NCBI has an advantage in identifying diatom species, but the accuracy and completeness of information are not as good as Diat.barcode; although the number of CNGB genus-level identification is not as good as that of Silva, it is similar to the number of species-level sequences identified by Silva, and CNGB contains some sequences that can not be retrieved at diatom species level in other databases. When using the method of database combination, the number of species-level identification sequences obtained by all four databases is close to that of Silva+NCBI+CNGB combination, but the Silva+NCBI+CNGB combination does not use Diat.barcode database, which can reduce the requirement of bioinformatics knowledge. The more accurate identification results of diatom species obtained by database combination are helpful to distinguish different sampling sites, which provides a basis for the application of diatom species differences in forensic medicine to infer the site of drowning. Conclusion: different database combinations can improve the accuracy of diatom species identification based on high-throughput sequencing results, although it depends on the integrity of the database to some extent, but more species level sequences help to improve the ability to distinguish sampling points. In this study, we recommend using Silva+NCBI+CNGB database combination to identify diatom species with high-throughput sequencing results. The high accuracy of diatom classification will contribute to the diagnosis of drowning and the forensic application of inferring suspected drowning sites.

China National GeneBank DataBase
Openalex Percentile: Top 8%
Injury Epidemiology and Prevention
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.