Direct microhaplotype genotyping for GT-seq (Genotyping-in-Thousands by Sequencing) using a diploid abundance model

GT-seq (Genotyping-in-Thousands by Sequencing) is widely used for high-throughput amplicon genotyping, but most analytical pipelines focus on single SNPs or rely on alignment-based variant calling. Here we present a direct microhaplotype genotyping framework that leverages the high read depth and low error rates typical of paired-end Illumina and Element sequencing. The pipeline first identifies primer-bounded reads and resolves paired-end sequences into quality-aware consensus amplicon sequences. Within each sample and locus, unique sequences are ranked by read abundance and the top one or two sequences are retained as directly observed haplotypes. These alleles are aggregated across samples to construct a catalog of observed haplotypes for each locus. In a second pass, reads are assigned to catalog haplotypes by exact sequence matching to produce diploid genotypes. Finally, catalog haplotype sequences are compared to identify phased SNP and collapsed indel variation. Optionally, catalog haplotypes may be aligned to a reference genome to project observed variants onto genomic coordinates and generate standards-compliant VCF output. This framework enables robust, microhaplotype genotyping directly from high-depth amplicon sequencing data. Comparison with an independent BWA/BCFtools alignment-based workflow demonstrated 99.67% genotype concordance across 102,520 genotype comparisons spanning 1,085 SNPs in 96 individuals. Genotype concordance remained above 99.4% even at the minimum supported sequencing depth of 10 reads per locus, demonstrating robust performance across a broad range of sequencing depths.

Authors

Institutions

Publication Details

Journal
PLoS Computational Biology
Published
2026-09-18
DOI
https://doi.org/10.1371/journal.pcbi.1014808
Primary Topic
Genomic variations and chromosomal abnormalities
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Direct microhaplotype genotyping for GT-seq (Genotyping-in-Thousands by Sequencing) using a diploid abundance model

Amanda J. Finger, Shannon Rose Kieran Blair, Amanda R. Campbell, Nathan Campbell
PLoS Computational Biology
Genomic variations and chromosomal abnormalities
article

Direct microhaplotype genotyping for GT-seq (Genotyping-in-Thousands by Sequencing) using a diploid abundance model

Amanda J. Finger, Shannon Rose Kieran Blair, Amanda R. Campbell, Nathan Campbell
article en

Abstract

GT-seq (Genotyping-in-Thousands by Sequencing) is widely used for high-throughput amplicon genotyping, but most analytical pipelines focus on single SNPs or rely on alignment-based variant calling. Here we present a direct microhaplotype genotyping framework that leverages the high read depth and low error rates typical of paired-end Illumina and Element sequencing. The pipeline first identifies primer-bounded reads and resolves paired-end sequences into quality-aware consensus amplicon sequences. Within each sample and locus, unique sequences are ranked by read abundance and the top one or two sequences are retained as directly observed haplotypes. These alleles are aggregated across samples to construct a catalog of observed haplotypes for each locus. In a second pass, reads are assigned to catalog haplotypes by exact sequence matching to produce diploid genotypes. Finally, catalog haplotype sequences are compared to identify phased SNP and collapsed indel variation. Optionally, catalog haplotypes may be aligned to a reference genome to project observed variants onto genomic coordinates and generate standards-compliant VCF output. This framework enables robust, microhaplotype genotyping directly from high-depth amplicon sequencing data. Comparison with an independent BWA/BCFtools alignment-based workflow demonstrated 99.67% genotype concordance across 102,520 genotype comparisons spanning 1,085 SNPs in 96 individuals. Genotype concordance remained above 99.4% even at the minimum supported sequencing depth of 10 reads per locus, demonstrating robust performance across a broad range of sequencing depths.

PLoS Computational BiologyVol. 22(9)
GfK (United States) (US), University of California, Davis (US)
U.S. Fish and Wildlife Service, Bureau of Reclamation
Openalex Percentile: Top 11%
Genomic variations and chromosomal abnormalities
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.