TRACOS: realistic synthetic data generation via chaotic trajectory modeling in high-dimensional imbalanced datasets

Background Imbalanced data distribution is a pervasive challenge in supervised learning, frequently causing models to favor the majority class. Conventional oversampling techniques, such as Synthetic Minority Over-sampling Technique (SMOTE) and its variants, often struggle to preserve the structural complexity of minority class instances, leading to overlapping classes and reduced generalization. Methods To address these limitations, this article introduces TRAjectorybased Chaotic Oversampling Strategy (TRACOS), a novel hybrid oversampling approach. The method first constructs smooth trajectories using adaptive spline interpolation to accurately capture the nonlinear manifold of the minority class. Subsequently, it applies controlled randomness via logistic chaotic maps to generate realistic synthetic samples along these trajectories, enhancing diversity while strictly maintaining structural fidelity. Results TRACOS was rigorously evaluated on 35 multidimensional and highly imbalanced datasets from the Knowledge Extraction based on Evolutionary Learning (KEEL) repository using four distinct classifiers (K-Nearest Neighbors, Decision Tree, Naive Bayes, and Random Forest). The proposed method was benchmarked against seven widely used oversampling techniques, including SMOTE, Borderline-SMOTE variants, SMOTE with Density estimation (SMOTE_D), SMOTE with Rough Set Theory (SMOTE_RSB), and Adaptive Semi-Unsupervised Weighted Oversampling (A_SUWO), alongside the original imbalanced datasets. Based on Area Under the Curve (AUC) and Geometric Mean (G-Mean) metrics under a 5-fold cross-validation setting, results demonstrate that TRACOS consistently achieves superior classification performance. Notably, under the Random Forest model, TRACOS achieved an average G-Mean of 0.8243 and ranked first in 22 out of 35 datasets, exhibiting the highest average scores and the lowest performance variance across all experiments. Discussion The dual mechanism of TRACOS capturing intrinsic data geometry through spline trajectories while safely promoting diversity via chaotic dynamics proves highly effective. It offers a robust, adaptable, and statistically superior solution for real-world classification applications characterized by severe class imbalance.

Authors

Publication Details

Journal
PeerJ Computer Science
Published
2026-10-05
DOI
https://doi.org/10.7717/peerj-cs.4109
Primary Topic
Imbalanced Data Classification Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

TRACOS: realistic synthetic data generation via chaotic trajectory modeling in high-dimensional imbalanced datasets

Hüseyin Eldem
PeerJ Computer Science
Imbalanced Data Classification Techniques
article

TRACOS: realistic synthetic data generation via chaotic trajectory modeling in high-dimensional imbalanced datasets

Hüseyin Eldem
article en

Abstract

Background Imbalanced data distribution is a pervasive challenge in supervised learning, frequently causing models to favor the majority class. Conventional oversampling techniques, such as Synthetic Minority Over-sampling Technique (SMOTE) and its variants, often struggle to preserve the structural complexity of minority class instances, leading to overlapping classes and reduced generalization. Methods To address these limitations, this article introduces TRAjectorybased Chaotic Oversampling Strategy (TRACOS), a novel hybrid oversampling approach. The method first constructs smooth trajectories using adaptive spline interpolation to accurately capture the nonlinear manifold of the minority class. Subsequently, it applies controlled randomness via logistic chaotic maps to generate realistic synthetic samples along these trajectories, enhancing diversity while strictly maintaining structural fidelity. Results TRACOS was rigorously evaluated on 35 multidimensional and highly imbalanced datasets from the Knowledge Extraction based on Evolutionary Learning (KEEL) repository using four distinct classifiers (K-Nearest Neighbors, Decision Tree, Naive Bayes, and Random Forest). The proposed method was benchmarked against seven widely used oversampling techniques, including SMOTE, Borderline-SMOTE variants, SMOTE with Density estimation (SMOTE_D), SMOTE with Rough Set Theory (SMOTE_RSB), and Adaptive Semi-Unsupervised Weighted Oversampling (A_SUWO), alongside the original imbalanced datasets. Based on Area Under the Curve (AUC) and Geometric Mean (G-Mean) metrics under a 5-fold cross-validation setting, results demonstrate that TRACOS consistently achieves superior classification performance. Notably, under the Random Forest model, TRACOS achieved an average G-Mean of 0.8243 and ranked first in 22 out of 35 datasets, exhibiting the highest average scores and the lowest performance variance across all experiments. Discussion The dual mechanism of TRACOS capturing intrinsic data geometry through spline trajectories while safely promoting diversity via chaotic dynamics proves highly effective. It offers a robust, adaptable, and statistically superior solution for real-world classification applications characterized by severe class imbalance.

PeerJ Computer ScienceVol. 12
Openalex Percentile: Top 11%
Imbalanced Data Classification Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

TRACOS: realistic synthetic data generation via chaotic trajectory modeling in high-dimensional imbalanced datasets — Hüseyin Eldem · PeerJ Computer Science (2026) | TGRS Research Map | TGRS