Improving multi-trait genomic prediction using synthetic traits from hyperspectral data based on co-heritability
KEY MESSAGE: Synthetic traits, wavelength ratios selected by co-heritability, raised multi-trait genomic predictive ability for leaf nitrogen and specific leaf area in sorghum by up to 17% over single-trait models. Genomic prediction (GP) is an essential tool in plant breeding as it can accelerate cultivar development by predicting the performance of unphenotyped lines. When using single-trait GP models, the precision of prediction is constrained by the heritability of the target trait. Multi-trait genomic prediction can be used to improve accuracy but requires identifying secondary traits with high heritability and genetic correlation with target traits. This study assessed the efficiency of multi-trait genomic prediction models powered by secondary traits derived from high-throughput phenotyping data when predicting leaf nitrogen content (N) and specific leaf area (SLA) in diverse sorghum accessions. We hypothesized that wavelength ratios from hyperspectral data could serve as synthetic secondary traits. Therefore, we developed models for direct measures of N and SLA, plus partial least squares regression (PLSR) predictions of them (Leaf N-PLSR, SLA-PLSR), totaling four target traits. Three synthetic traits (S1, S2, S3), each a ratio of two wavelengths within the hyperspectral data, were identified based on high co-heritability with target traits. Single-trait (genomic best linear unbiased prediction (GBLUP) served as the baseline model, followed by multi-trait GBLUP models combining synthetic and target traits. Model performance was assessed using fivefold cross-validation under single-trait, CV1, and CV2 schemes. Our approach improves multi-trait genomic prediction of target traits by using synthetic traits with no intrinsic biological meaning selected through co-heritability estimation. This demonstrates that there's more useful information in the spectra than is typically utilized, and this information can be leveraged to improve multi-trait prediction of a target trait.
Authors
- Andrew D. B. Leakey (ORCID: https://orcid.org/0000-0001-6251-024X)
- Samuel B. Fernandes (ORCID: https://orcid.org/0000-0001-8269-535X)
- Ashmita Upadhyay (ORCID: https://orcid.org/0000-0003-3933-8239)
- Alexander E. Lipka (ORCID: https://orcid.org/0000-0003-1571-8528)
- Rachel Paul (ORCID: https://orcid.org/0000-0002-1421-0637)
- Mohammed El-Kebir (ORCID: https://orcid.org/0000-0002-1468-2407)
- Stefan Ivanovic
- Ruhana Azam
- Sanmi Koyejo
- John N. Ferguson
- Meilu Yuan
Institutions
- University of Illinois Urbana-Champaign (US)
- University of Arkansas System (US)
- Carnegie Department of Plant Biology (US)
- University of Arkansas at Fayetteville (US)
- Stanford University (US)
Publication Details
- Journal
- Theoretical and Applied Genetics
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1007/s00122-026-05334-2
- Primary Topic
- Genetic Mapping and Diversity in Plants and Animals
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Institute of Food and Agriculture
- Advanced Research Projects Agency
- Biological and Environmental Research