Robust optimization reformulations for support vector machines under label and feature uncertainty
Classification models are widely used in machine learning to assign labels to observations based on their features. However, in many real-world applications, training data are affected by uncertainty. Labels may be incorrect due to annotation errors, subjective judgments, or recording mistakes, while feature values may be perturbed due to measurement errors, sensor noise, or preprocessing. These sources of uncertainty can reduce the reliability and generalization performance of classification models. This thesis develops robust optimization formulations for support vector machines (SVMs) under label and feature uncertainty. First, a compact reformulation is proposed for label-uncertain SVMs with cardinality-bounded label perturbations. Compared with an existing mixed-integer formulation, the proposed model reduces the number of binary variables from 2n to n, while preserving the same robustness guarantees. Computational experiments on synthetic and real-world datasets show that this compact formulation significantly reduces solver runtime, especially for larger datasets. Second, this thesis develops a column-and-constraint generation framework for structured label uncertainty. This approach allows additional linear constraints to be imposed on admissible label perturbations, making it possible to model group restrictions, application-specific rules, and monotonicity requirements. Numerical experiments demonstrate that the proposed framework can handle more flexible uncertainty structures while remaining computationally tractable. Third, feature uncertainty is studied by introducing a global uncertainty set that constrains the total perturbation budget across the full feature matrix. Unlike row-wise uncertainty sets, which perturb each observation independently, the proposed global model allows perturbations to be distributed across the dataset. A column-and-constraint generation algorithm is developed to solve the resulting two-stage robust optimization problem. Computational results show that the proposed feature-uncertain formulation is more conservative and computationally demanding, but provides stronger robustness under feature perturbations and improved out-of-sample performance in several experiments. Overall, this thesis contributes efficient and flexible robust optimization methods for SVM classification under uncertain training data. The proposed formulations improve computational performance for label uncertainty and provide broader modeling flexibility for structured label and feature uncertainty.
Authors
- Mohammadsadra Nejati
Publication Details
- Journal
- Open Collections
- Published
- 2026-09-04
- DOI
- https://doi.org/10.14288/1.0455975
- Primary Topic
- Machine Learning and Data Classification
- Type
- article
- Field-Weighted Citation Impact
- 0.00