A two-step feature selection using permutation based feature importance and wrapper-based bio-inspired algorithms for classifying clinical dataset

Abstract The proposed framework aims to address the challenge of high-dimensional, noisy, irrelevant, and redundant features in clinical datasets by using a two-step feature selection approach. The first step involves a wrapper-based method for calculating the Permutation Based Feature Importance (PBFI) score using the Support Vector Machine (SVM) classifier. The second step employs three wrapper-based bio-inspired optimization algorithms, namely Harris Hawk Optimization (HHO), Binary Grasshopper Optimization (BGOA), and Whale Optimization (WOA), with the weighted F1-Score, measured by the SVM classifier, as the fitness function. The selected features from both steps are then combined using a union operation and used to train four classifiers: SVM, K-Nearest Neighbor (K-NN), Linear Discriminant Analysis (LDA), and Naive Bayes (NB). The proposed framework is also evaluated using a non-parametric hypothesis testing method, the Kruskal-Wallis Test, on eight medical datasets from the Machine Learning Repository (MLR) maintained by the University of California, Irvine (UCI). The proposed approach is compared with feature selection using the Particle Swarm Optimization algorithm (PSOA) and Lion's Algorithm, and is shown to perform better.

Authors

Institutions

Publication Details

Journal
Intelligent Data Analysis
Published
2026-09-03
DOI
https://doi.org/10.1177/1088467x261475592
Primary Topic
Metaheuristic Optimization Algorithms Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A two-step feature selection using permutation based feature importance and wrapper-based bio-inspired algorithms for classifying clinical dataset

Nancy Jane Yesudhas, Khanna Harichandran Nehemiah, Kannan Arputharaj, Kochukrishnan Sivendran Navin
Intelligent Data Analysis
Metaheuristic Optimization Algorithms Research
article

A two-step feature selection using permutation based feature importance and wrapper-based bio-inspired algorithms for classifying clinical dataset

Nancy Jane Yesudhas, Khanna Harichandran Nehemiah, Kannan Arputharaj, Kochukrishnan Sivendran Navin
article en

Abstract

Abstract The proposed framework aims to address the challenge of high-dimensional, noisy, irrelevant, and redundant features in clinical datasets by using a two-step feature selection approach. The first step involves a wrapper-based method for calculating the Permutation Based Feature Importance (PBFI) score using the Support Vector Machine (SVM) classifier. The second step employs three wrapper-based bio-inspired optimization algorithms, namely Harris Hawk Optimization (HHO), Binary Grasshopper Optimization (BGOA), and Whale Optimization (WOA), with the weighted F1-Score, measured by the SVM classifier, as the fitness function. The selected features from both steps are then combined using a union operation and used to train four classifiers: SVM, K-Nearest Neighbor (K-NN), Linear Discriminant Analysis (LDA), and Naive Bayes (NB). The proposed framework is also evaluated using a non-parametric hypothesis testing method, the Kruskal-Wallis Test, on eight medical datasets from the Machine Learning Repository (MLR) maintained by the University of California, Irvine (UCI). The proposed approach is compared with feature selection using the Particle Swarm Optimization algorithm (PSOA) and Lion's Algorithm, and is shown to perform better.

Intelligent Data Analysis
University of Madras (IN), Indian Institute of Technology Madras (IN), Anna University, Chennai (IN), Vellore Institute of Technology University (IN)
Reduced inequalities
Openalex Percentile: Top 8%
Metaheuristic Optimization Algorithms Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A two-step feature selection using permutation based feature importance and wrapper-based bio-inspired algorithms for classifying clinical dataset — Nancy Jane Yesudhas, Khanna Harichandran Nehemiah, et al. · Intelligent Data Analysis (2026) | TGRS Research Map | TGRS