GDA-Pred: Generative AI-Driven Data Augmentation for Improved Prediction of IL-6 and IL-13-Inducing Peptides

Abstract The identification of interleukin-6 (IL-6) and interleukin-13 (IL-13) inducing peptides is crucial for accelerating drug discovery targeting cancer, immune disorders, and infectious diseases. However, experimental screening methods remain time-consuming and costly. To address these limitations, various machine learning and deep learning models have been developed, yet their performance is still constrained by the limited availability of experimentally validated data. In this study, we propose a generative AI-driven data augmentation (GDA) framework and predictor, GDA-Pred, to improve the prediction performance of state-of-the-art (SOTA) classifiers for identifying IL-6 and IL-13 inducing peptides. GDA expands the training dataset by generating novel peptide sequences using three types of generative AI models: generative adversarial networks (GANs), diffusion models (DMs), and variational autoencoders (VAEs). The GDA framework is defined by four key parameters: the type of generative model, the sequence identity cutoff, the probability threshold (PT) for selecting generated peptides, and the augmentation ratio (AR) between generated and real peptides. Since optimizing these parameters is challenging with small datasets, we adopt a case study-oriented proof-of-concept approach using a moderately sized dataset of anti-inflammatory peptides (AIPs) to derive interpretable optimal settings. The performance of the optimized GDA was evaluated using stratified 5-fold cross-validation with cluster-based partitioning and a hold-out test on benchmark datasets. The optimized GDA was then applied to SOTA classifiers, collectively termed GDA-Pred, to identify IL-6 and IL-13 inducing peptides, both of which are limited by small dataset sizes. GDA-Pred substantially improved the prediction performance for both cytokine-inducing peptide tasks, demonstrating the feasibility of GDA-Pred as a robust and generalizable framework.

Authors

Institutions

Publication Details

Journal
International Journal of Molecular Sciences
Published
2026-08-27
DOI
https://doi.org/10.3390/ijms27177688
Primary Topic
vaccines and immunoinformatics approaches
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

GDA-Pred: Generative AI-Driven Data Augmentation for Improved Prediction of IL-6 and IL-13-Inducing Peptides

Hiroyuki Kurata, Md. Harun-Or-Roshid, Kazuhiro Maeda
International Journal of Molecular Sciences
vaccines and immunoinformatics approaches
article

GDA-Pred: Generative AI-Driven Data Augmentation for Improved Prediction of IL-6 and IL-13-Inducing Peptides

Hiroyuki Kurata, Md. Harun-Or-Roshid, Kazuhiro Maeda
article en

Abstract

Abstract The identification of interleukin-6 (IL-6) and interleukin-13 (IL-13) inducing peptides is crucial for accelerating drug discovery targeting cancer, immune disorders, and infectious diseases. However, experimental screening methods remain time-consuming and costly. To address these limitations, various machine learning and deep learning models have been developed, yet their performance is still constrained by the limited availability of experimentally validated data. In this study, we propose a generative AI-driven data augmentation (GDA) framework and predictor, GDA-Pred, to improve the prediction performance of state-of-the-art (SOTA) classifiers for identifying IL-6 and IL-13 inducing peptides. GDA expands the training dataset by generating novel peptide sequences using three types of generative AI models: generative adversarial networks (GANs), diffusion models (DMs), and variational autoencoders (VAEs). The GDA framework is defined by four key parameters: the type of generative model, the sequence identity cutoff, the probability threshold (PT) for selecting generated peptides, and the augmentation ratio (AR) between generated and real peptides. Since optimizing these parameters is challenging with small datasets, we adopt a case study-oriented proof-of-concept approach using a moderately sized dataset of anti-inflammatory peptides (AIPs) to derive interpretable optimal settings. The performance of the optimized GDA was evaluated using stratified 5-fold cross-validation with cluster-based partitioning and a hold-out test on benchmark datasets. The optimized GDA was then applied to SOTA classifiers, collectively termed GDA-Pred, to identify IL-6 and IL-13 inducing peptides, both of which are limited by small dataset sizes. GDA-Pred substantially improved the prediction performance for both cytokine-inducing peptide tasks, demonstrating the feasibility of GDA-Pred as a robust and generalizable framework.

International Journal of Molecular SciencesVol. 27(17)
Kyushu Institute of Technology (JP)
Japan Society for the Promotion of Science
Openalex Percentile: Top 99%
vaccines and immunoinformatics approaches
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.