Robust optimization reformulations for support vector machines under label and feature uncertainty

Classification models are widely used in machine learning to assign labels to observations based on their features. However, in many real-world applications, training data are affected by uncertainty. Labels may be incorrect due to annotation errors, subjective judgments, or recording mistakes, while feature values may be perturbed due to measurement errors, sensor noise, or preprocessing. These sources of uncertainty can reduce the reliability and generalization performance of classification models. This thesis develops robust optimization formulations for support vector machines (SVMs) under label and feature uncertainty. First, a compact reformulation is proposed for label-uncertain SVMs with cardinality-bounded label perturbations. Compared with an existing mixed-integer formulation, the proposed model reduces the number of binary variables from 2n to n, while preserving the same robustness guarantees. Computational experiments on synthetic and real-world datasets show that this compact formulation significantly reduces solver runtime, especially for larger datasets. Second, this thesis develops a column-and-constraint generation framework for structured label uncertainty. This approach allows additional linear constraints to be imposed on admissible label perturbations, making it possible to model group restrictions, application-specific rules, and monotonicity requirements. Numerical experiments demonstrate that the proposed framework can handle more flexible uncertainty structures while remaining computationally tractable. Third, feature uncertainty is studied by introducing a global uncertainty set that constrains the total perturbation budget across the full feature matrix. Unlike row-wise uncertainty sets, which perturb each observation independently, the proposed global model allows perturbations to be distributed across the dataset. A column-and-constraint generation algorithm is developed to solve the resulting two-stage robust optimization problem. Computational results show that the proposed feature-uncertain formulation is more conservative and computationally demanding, but provides stronger robustness under feature perturbations and improved out-of-sample performance in several experiments. Overall, this thesis contributes efficient and flexible robust optimization methods for SVM classification under uncertain training data. The proposed formulations improve computational performance for label uncertainty and provide broader modeling flexibility for structured label and feature uncertainty.

Authors

Publication Details

Journal
Open Collections
Published
2026-09-04
DOI
https://doi.org/10.14288/1.0455975
Primary Topic
Machine Learning and Data Classification
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Robust optimization reformulations for support vector machines under label and feature uncertainty

Mohammadsadra Nejati
Open Collections
Machine Learning and Data Classification
article

Robust optimization reformulations for support vector machines under label and feature uncertainty

Mohammadsadra Nejati
article en

Abstract

Classification models are widely used in machine learning to assign labels to observations based on their features. However, in many real-world applications, training data are affected by uncertainty. Labels may be incorrect due to annotation errors, subjective judgments, or recording mistakes, while feature values may be perturbed due to measurement errors, sensor noise, or preprocessing. These sources of uncertainty can reduce the reliability and generalization performance of classification models. This thesis develops robust optimization formulations for support vector machines (SVMs) under label and feature uncertainty. First, a compact reformulation is proposed for label-uncertain SVMs with cardinality-bounded label perturbations. Compared with an existing mixed-integer formulation, the proposed model reduces the number of binary variables from 2n to n, while preserving the same robustness guarantees. Computational experiments on synthetic and real-world datasets show that this compact formulation significantly reduces solver runtime, especially for larger datasets. Second, this thesis develops a column-and-constraint generation framework for structured label uncertainty. This approach allows additional linear constraints to be imposed on admissible label perturbations, making it possible to model group restrictions, application-specific rules, and monotonicity requirements. Numerical experiments demonstrate that the proposed framework can handle more flexible uncertainty structures while remaining computationally tractable. Third, feature uncertainty is studied by introducing a global uncertainty set that constrains the total perturbation budget across the full feature matrix. Unlike row-wise uncertainty sets, which perturb each observation independently, the proposed global model allows perturbations to be distributed across the dataset. A column-and-constraint generation algorithm is developed to solve the resulting two-stage robust optimization problem. Computational results show that the proposed feature-uncertain formulation is more conservative and computationally demanding, but provides stronger robustness under feature perturbations and improved out-of-sample performance in several experiments. Overall, this thesis contributes efficient and flexible robust optimization methods for SVM classification under uncertain training data. The proposed formulations improve computational performance for label uncertainty and provide broader modeling flexibility for structured label and feature uncertainty.

Open Collections
Openalex Percentile: Top 8%
Machine Learning and Data Classification
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.