The effect of train-test data splitting ratio on yield prediction from multispectral UAV imagery: effective reduction of training ratio with algorithm selection for single- and multi-date models

Abstract Context Optimizing the ratio of data required for training spectral grain yield (GY) prediction models is critical for efficient sampling of reference data, but has been rarely addressed so far. Given the high costs for collecting ground truth data, standard data splitting ratios which use the majority of the data for training and the rest for testing the models should be minimized, but in turn may affect prediction accuracy. Methods Therefore, this study evaluated 11 different train-test data splitting ratios (TSR), ranging from using 5% to 95% of the data for training, and the remaining portion as test sets. Models with six different machine learning algorithms were compared in winter wheat breeding trials conducted with each several thousand plots in two locations in Germany over a period of four years. The input data consisted of the (unmanned aerial vehicle) UAV-based NDRE index data from individual measurements dates as well as multi-date combinations. Results The results indicate that GY prediction remained relatively stable when decreasing TSR to about 0.30. Conversely, multi-date models tended to benefit more from higher TSR than single-date models. Support vector machine and random forest algorithms demonstrated relative advantage both for higher TSR and multi-date models, whereas partial least squares and ridge regression were the best algorithms for lowest TSR-values. Furthermore, analysis via repeated data splitting revealed minimum R² variability at a TSR of 0.30, but substantially increased variance at higher TSR values. Conclusions It is concluded that decreasing TSR while considering algorithm selection can potentially reduce costs without compromising prediction accuracy, therefore making spectral phenotyping methods more accessible and ready-to-use, whereas the absolute number of data points requires additional examination.

Authors

Institutions

Publication Details

Journal
Precision Agriculture
Published
2026-10-05
DOI
https://doi.org/10.1007/s11119-026-10464-0
Primary Topic
Remote Sensing in Agriculture
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

The effect of train-test data splitting ratio on yield prediction from multispectral UAV imagery: effective reduction of training ratio with algorithm selection for single- and multi-date models

Lukas Prey, Patrick Noack
Precision Agriculture
Remote Sensing in Agriculture
article

The effect of train-test data splitting ratio on yield prediction from multispectral UAV imagery: effective reduction of training ratio with algorithm selection for single- and multi-date models

Lukas Prey, Patrick Noack
article en

Abstract

Abstract Context Optimizing the ratio of data required for training spectral grain yield (GY) prediction models is critical for efficient sampling of reference data, but has been rarely addressed so far. Given the high costs for collecting ground truth data, standard data splitting ratios which use the majority of the data for training and the rest for testing the models should be minimized, but in turn may affect prediction accuracy. Methods Therefore, this study evaluated 11 different train-test data splitting ratios (TSR), ranging from using 5% to 95% of the data for training, and the remaining portion as test sets. Models with six different machine learning algorithms were compared in winter wheat breeding trials conducted with each several thousand plots in two locations in Germany over a period of four years. The input data consisted of the (unmanned aerial vehicle) UAV-based NDRE index data from individual measurements dates as well as multi-date combinations. Results The results indicate that GY prediction remained relatively stable when decreasing TSR to about 0.30. Conversely, multi-date models tended to benefit more from higher TSR than single-date models. Support vector machine and random forest algorithms demonstrated relative advantage both for higher TSR and multi-date models, whereas partial least squares and ridge regression were the best algorithms for lowest TSR-values. Furthermore, analysis via repeated data splitting revealed minimum R² variability at a TSR of 0.30, but substantially increased variance at higher TSR values. Conclusions It is concluded that decreasing TSR while considering algorithm selection can potentially reduce costs without compromising prediction accuracy, therefore making spectral phenotyping methods more accessible and ready-to-use, whereas the absolute number of data points requires additional examination.

Precision AgricultureVol. 27(6)
Weihenstephan-Triesdorf University of Applied Sciences (DE)
Openalex Percentile: Top 14%
Remote Sensing in Agriculture
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.