Benchmarking ensemble and deep learning models for drought-responsive gene prioritisation in rice using multi-omics data
Identifying drought-responsive genes is critical for climate-resilient crop breeding, yet computational gene–trait association often relies on single-omics or generalized genomic datasets. Optimal algorithm selection for high-dimensional, small multi-omics datasets also remains a persistent challenge. We present a machine learning framework that extends the QTG-Finder feature set by integrating transcriptomic and proteomic features to classify drought-responsive genes in rice. Using MBKBase annotations, we evaluated the ability of these models to prioritise 336 drought-responsive genes among 2,416 abiotic stress-associated genes, of which the remaining 2,080 constituted an operational background set. We tested four algorithms: Random Forest, XGBoost, one-dimensional convolutional neural network, and multilayer perceptron, while addressing severe class imbalance and assessing model stability. Tree-based ensemble methods achieved higher performance than the deep learning models in this dataset, which showed early overfitting. The fine-tuned Random Forest achieved a recall of 0.67 and an AUC-ROC of 0.87 under nested cross-validation, conditional on a feature matrix assembled before gene-level data splitting. Across repeated evaluations of a small internal holdout of 10 genes, mean accuracy and recall were 0.90 and 0.63; this set is too small to support a generalizability claim and is reported only as a consistency check. The newly added transcriptomic and proteomic features improved recall from 0.47 to 0.67 compared with genomics alone. Feature importance and ablation analyses indicated that transcriptomics-derived features, particularly measures of expression distribution, magnitude, and direction, contributed most to predictive performance. Functional enrichment analysis supported the biological relevance of the prioritised genes. These results demonstrate the value of multi-omics integration for gene prioritisation in small tabular datasets while highlighting current limitations of deep learning in this setting.
Authors
- Dharini Pathmanathan (ORCID: https://orcid.org/0000-0002-3279-4031)
- Rabiatul‐Adawiah Zainal‐Abidin (ORCID: https://orcid.org/0000-0002-3348-5636)
- Arpah Abu (ORCID: https://orcid.org/0000-0001-7072-0618)
- Kumanan N. Govaichelvan (ORCID: https://orcid.org/0000-0002-3456-3472)
Institutions
- Centre of Excellence in Mathematics (TH)
- Institute of Mathematical Sciences (IN)
- Malaysian Agricultural Research and Development Institute (MY)
Publication Details
- Journal
- PLoS ONE
- Published
- 2026-10-09
- DOI
- https://doi.org/10.1371/journal.pone.0360052
- Primary Topic
- Bioinformatics and Genomic Networks
- Type
- article
- Field-Weighted Citation Impact
- 0.00