Multi-Objective Minimal Feature Selection for Explainable-by-Design Classification: Equivalence-Aware Selection and Rashomon-Set Analysis Across Six Domains

High-stakes and regulated decision-support systems require predictive models that are simultaneously accurate and interpretable, yet accuracy is frequently obtained through opaque models whose behaviour is difficult to audit. This paper develops an explainable-by-design classification framework in which interpretability is treated as a constructive objective rather than a post hoc add-on. We aggregate thirteen complementary relevance metrics into a single, stable ranking and then cast the selection of a minimal relevant feature subset as a multi-objective optimization problem that trades predictive quality against structural complexity. The resulting accuracy–complexity Pareto front is approximated with the NSGA-II algorithm, and an equivalence-aware rule selects, among statistically indistinguishable models, the most parsimonious and hence most transparent one. We further (i) characterize the feature-subset ε-Rashomon set of each problem, (ii) quantify the stability of the aggregated ranking under resampling and contrast it with the individual metrics, and (iii) benchmark the induced subsets against established selectors (mRMR, Boruta, RFE, LASSO) and against opaque full-feature baselines, including gradient boosting. The framework is validated with a leakage-free nested cross-validation—in which the ranking, the search and the selection are recomputed inside each training fold—across six heterogeneous domains (healthcare, industrial production, climate, socio-economics, and education). Empirically, reducing the feature space to about three features preserves predictive accuracy: under paired tests corrected for fold dependence (Nadeau–Bengio) and for multiple comparisons (Holm), the minimal model is statistically indistinguishable from the full-feature model on all six datasets and is never significantly worse: reducing to a handful of features is essentially free. The main payoff is therefore not an accuracy gain but the characterization it enables—a simple ranked prefix performs comparably to the equivalence-aware selection (a direct signature of the Rashomon effect), and every problem admits a large set of near-equivalent minimal subsets sharing a compact stable core, which we map explicitly. The aggregated ranking is also more stable than the average individual metric in four of six datasets. These results support an explainability-by-design paradigm in which the accuracy cost of transparency—incurred only on the hardest multiclass tasks—is small, explicit, and justified when auditability and human oversight are required.

Authors

Institutions

Publication Details

Journal
Machine Learning and Knowledge Extraction
Published
2026-10-08
DOI
https://doi.org/10.3390/make8100319
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Multi-Objective Minimal Feature Selection for Explainable-by-Design Classification: Equivalence-Aware Selection and Rashomon-Set Analysis Across Six Domains

Yair Rivera Julio, Roberto Porto Solano, José Manuel Molina, Antonio Berlanga
Machine Learning and Knowledge Extraction
Explainable Artificial Intelligence (XAI)
article

Multi-Objective Minimal Feature Selection for Explainable-by-Design Classification: Equivalence-Aware Selection and Rashomon-Set Analysis Across Six Domains

Yair Rivera Julio, Roberto Porto Solano, José Manuel Molina, Antonio Berlanga
article en

Abstract

High-stakes and regulated decision-support systems require predictive models that are simultaneously accurate and interpretable, yet accuracy is frequently obtained through opaque models whose behaviour is difficult to audit. This paper develops an explainable-by-design classification framework in which interpretability is treated as a constructive objective rather than a post hoc add-on. We aggregate thirteen complementary relevance metrics into a single, stable ranking and then cast the selection of a minimal relevant feature subset as a multi-objective optimization problem that trades predictive quality against structural complexity. The resulting accuracy–complexity Pareto front is approximated with the NSGA-II algorithm, and an equivalence-aware rule selects, among statistically indistinguishable models, the most parsimonious and hence most transparent one. We further (i) characterize the feature-subset ε-Rashomon set of each problem, (ii) quantify the stability of the aggregated ranking under resampling and contrast it with the individual metrics, and (iii) benchmark the induced subsets against established selectors (mRMR, Boruta, RFE, LASSO) and against opaque full-feature baselines, including gradient boosting. The framework is validated with a leakage-free nested cross-validation—in which the ranking, the search and the selection are recomputed inside each training fold—across six heterogeneous domains (healthcare, industrial production, climate, socio-economics, and education). Empirically, reducing the feature space to about three features preserves predictive accuracy: under paired tests corrected for fold dependence (Nadeau–Bengio) and for multiple comparisons (Holm), the minimal model is statistically indistinguishable from the full-feature model on all six datasets and is never significantly worse: reducing to a handful of features is essentially free. The main payoff is therefore not an accuracy gain but the characterization it enables—a simple ranked prefix performs comparably to the equivalence-aware selection (a direct signature of the Rashomon effect), and every problem admits a large set of near-equivalent minimal subsets sharing a compact stable core, which we map explicitly. The aggregated ranking is also more stable than the average individual metric in four of six datasets. These results support an explainability-by-design paradigm in which the accuracy cost of transparency—incurred only on the hardest multiclass tasks—is small, explicit, and justified when auditability and human oversight are required.

Machine Learning and Knowledge ExtractionVol. 8(10)
Corporación Universitaria Americana (CO), Universidad Carlos III de Madrid (ES)
Openalex Percentile: Top 13%
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.