Benchmarking ensemble and deep learning models for drought-responsive gene prioritisation in rice using multi-omics data

Identifying drought-responsive genes is critical for climate-resilient crop breeding, yet computational gene–trait association often relies on single-omics or generalized genomic datasets. Optimal algorithm selection for high-dimensional, small multi-omics datasets also remains a persistent challenge. We present a machine learning framework that extends the QTG-Finder feature set by integrating transcriptomic and proteomic features to classify drought-responsive genes in rice. Using MBKBase annotations, we evaluated the ability of these models to prioritise 336 drought-responsive genes among 2,416 abiotic stress-associated genes, of which the remaining 2,080 constituted an operational background set. We tested four algorithms: Random Forest, XGBoost, one-dimensional convolutional neural network, and multilayer perceptron, while addressing severe class imbalance and assessing model stability. Tree-based ensemble methods achieved higher performance than the deep learning models in this dataset, which showed early overfitting. The fine-tuned Random Forest achieved a recall of 0.67 and an AUC-ROC of 0.87 under nested cross-validation, conditional on a feature matrix assembled before gene-level data splitting. Across repeated evaluations of a small internal holdout of 10 genes, mean accuracy and recall were 0.90 and 0.63; this set is too small to support a generalizability claim and is reported only as a consistency check. The newly added transcriptomic and proteomic features improved recall from 0.47 to 0.67 compared with genomics alone. Feature importance and ablation analyses indicated that transcriptomics-derived features, particularly measures of expression distribution, magnitude, and direction, contributed most to predictive performance. Functional enrichment analysis supported the biological relevance of the prioritised genes. These results demonstrate the value of multi-omics integration for gene prioritisation in small tabular datasets while highlighting current limitations of deep learning in this setting.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-10-09
DOI
https://doi.org/10.1371/journal.pone.0360052
Primary Topic
Bioinformatics and Genomic Networks
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Benchmarking ensemble and deep learning models for drought-responsive gene prioritisation in rice using multi-omics data

Dharini Pathmanathan, Rabiatul‐Adawiah Zainal‐Abidin, Arpah Abu, Kumanan N. Govaichelvan
PLoS ONE
Bioinformatics and Genomic Networks
article

Benchmarking ensemble and deep learning models for drought-responsive gene prioritisation in rice using multi-omics data

Dharini Pathmanathan, Rabiatul‐Adawiah Zainal‐Abidin, Arpah Abu, Kumanan N. Govaichelvan
article en

Abstract

Identifying drought-responsive genes is critical for climate-resilient crop breeding, yet computational gene–trait association often relies on single-omics or generalized genomic datasets. Optimal algorithm selection for high-dimensional, small multi-omics datasets also remains a persistent challenge. We present a machine learning framework that extends the QTG-Finder feature set by integrating transcriptomic and proteomic features to classify drought-responsive genes in rice. Using MBKBase annotations, we evaluated the ability of these models to prioritise 336 drought-responsive genes among 2,416 abiotic stress-associated genes, of which the remaining 2,080 constituted an operational background set. We tested four algorithms: Random Forest, XGBoost, one-dimensional convolutional neural network, and multilayer perceptron, while addressing severe class imbalance and assessing model stability. Tree-based ensemble methods achieved higher performance than the deep learning models in this dataset, which showed early overfitting. The fine-tuned Random Forest achieved a recall of 0.67 and an AUC-ROC of 0.87 under nested cross-validation, conditional on a feature matrix assembled before gene-level data splitting. Across repeated evaluations of a small internal holdout of 10 genes, mean accuracy and recall were 0.90 and 0.63; this set is too small to support a generalizability claim and is reported only as a consistency check. The newly added transcriptomic and proteomic features improved recall from 0.47 to 0.67 compared with genomics alone. Feature importance and ablation analyses indicated that transcriptomics-derived features, particularly measures of expression distribution, magnitude, and direction, contributed most to predictive performance. Functional enrichment analysis supported the biological relevance of the prioritised genes. These results demonstrate the value of multi-omics integration for gene prioritisation in small tabular datasets while highlighting current limitations of deep learning in this setting.

PLoS ONEVol. 21(10)
Centre of Excellence in Mathematics (TH), Institute of Mathematical Sciences (IN), Malaysian Agricultural Research and Development Institute (MY)
Openalex Percentile: Top 23%
Bioinformatics and Genomic Networks
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.