Genomic Selection in Rice Using Targeted Sequencing: Model Evaluation and Training Population Optimization for Breeding Applications

Background/Objectives: Genomic selection (GS) can accelerate crop improvement, but its application in breeding programs is often constrained by the cost of high-density genotyping. This study evaluated genomic prediction using 4647 targeted sequencing-derived SNPs in 829 rice accessions and investigated the effects of prediction model selection, population structure, training population optimization, and GWAS-assisted marker selection on predictive performance. Methods: Seventeen GS models were evaluated for eight agronomic traits using repeated five-fold cross-validation, cross-population prediction, STPGA-based training population optimization, and GWAS-derived marker subsets. Resaults: Predictive performance varied substantially among traits and models. For grain length (GL), EnsembleGS achieved the highest predictive performance among the 17 models (r = 0.838), while RKHS showed the highest performance for grain number per panicle (r = 0.698), RFR for heading date (r = 0.854), and EnsembleGS for length-to-width ratio (r = 0.844). Cross-population prediction generally resulted in lower predictive performance than within-population cross-validation, with substantial variation among transfer directions. STPGA retained substantial predictive performance with reduced training population sizes; for GL, the highest performance reached r = 0.800 with a 400-individual STPGA-selected training population. GWAS-guided marker prioritization also retained substantial predictive information after marker reduction. For GL, the highest performance reached r = 0.789 and 0.802 using the GWAS-Top500 and GWAS-Top1000 subsets, respectively, compared with 0.751 and 0.775 for the corresponding Random500 and Random1000 subsets. For length-to-width ratio, the corresponding values were r = 0.764 and 0.800 for GWAS-Top500 and GWAS-Top1000, compared with 0.727 and 0.783 for the random subsets. Conclusions: Integrating targeted sequencing with trait- and model-aware prediction, optimized training population design, and GWAS-guided marker prioritization can retain substantial genomic prediction performance while reducing marker and training population requirements, providing a potentially cost-effective framework for genomic selection in rice breeding.

Authors

Institutions

Publication Details

Journal
Genes
Published
2026-10-08
DOI
https://doi.org/10.3390/genes17101240
Primary Topic
Genetic and phenotypic traits in livestock
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Genomic Selection in Rice Using Targeted Sequencing: Model Evaluation and Training Population Optimization for Breeding Applications

Long Pan, Taijiao Hu, Shaohua Yang, Zaijie Chen et al.
Genes
Genetic and phenotypic traits in livestock
article

Genomic Selection in Rice Using Targeted Sequencing: Model Evaluation and Training Population Optimization for Breeding Applications

Long Pan, Taijiao Hu, Shaohua Yang, Zaijie Chen, Mingji Wu, Yana Song, Yan Lin, Huaqing Liu, Min Lin, Yue Wang
article en

Abstract

Background/Objectives: Genomic selection (GS) can accelerate crop improvement, but its application in breeding programs is often constrained by the cost of high-density genotyping. This study evaluated genomic prediction using 4647 targeted sequencing-derived SNPs in 829 rice accessions and investigated the effects of prediction model selection, population structure, training population optimization, and GWAS-assisted marker selection on predictive performance. Methods: Seventeen GS models were evaluated for eight agronomic traits using repeated five-fold cross-validation, cross-population prediction, STPGA-based training population optimization, and GWAS-derived marker subsets. Resaults: Predictive performance varied substantially among traits and models. For grain length (GL), EnsembleGS achieved the highest predictive performance among the 17 models (r = 0.838), while RKHS showed the highest performance for grain number per panicle (r = 0.698), RFR for heading date (r = 0.854), and EnsembleGS for length-to-width ratio (r = 0.844). Cross-population prediction generally resulted in lower predictive performance than within-population cross-validation, with substantial variation among transfer directions. STPGA retained substantial predictive performance with reduced training population sizes; for GL, the highest performance reached r = 0.800 with a 400-individual STPGA-selected training population. GWAS-guided marker prioritization also retained substantial predictive information after marker reduction. For GL, the highest performance reached r = 0.789 and 0.802 using the GWAS-Top500 and GWAS-Top1000 subsets, respectively, compared with 0.751 and 0.775 for the corresponding Random500 and Random1000 subsets. For length-to-width ratio, the corresponding values were r = 0.764 and 0.800 for GWAS-Top500 and GWAS-Top1000, compared with 0.727 and 0.783 for the random subsets. Conclusions: Integrating targeted sequencing with trait- and model-aware prediction, optimized training population design, and GWAS-guided marker prioritization can retain substantial genomic prediction performance while reducing marker and training population requirements, providing a potentially cost-effective framework for genomic selection in rice breeding.

GenesVol. 17(10)
Fujian Academy of Agricultural Sciences (CN)
Openalex Percentile: Top 14%
Genetic and phenotypic traits in livestock
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.