Scalable multicollinearity recovery via mixed-integer optimization
In linear regression, multicollinearity and noise undermine coefficient estimation and predictive accuracy. This paper proposes Scalable Multicollinearity Recovery (SMR), a multicollinearity-detection framework that extends an existing mixed-integer quadratic optimization approach but differs in its formulation, scalability, and theoretical grounding. We propose a correlation-based screen and a parameter-free eigenvector screen, recast the minimum-support program with a Special Ordered Set constraint and a closed-form verification step, and introduce an irreducibility test and a residual-guided fast-path completion. We further establish its identifiability, stability, selection consistency, finite termination, and computational complexity. Together, these enhancements reduce false positives and allow detection to scale from 1000 to 10,000 predictors: at 10,000 predictors, the SMR procedure maintains detection accuracy at 100% while cutting the false-positive rate from 33% to 9% and runtime from about 6000 s to under 200 s. On six real-world datasets, one from OpenML and five from the UCI Machine Learning Repository, SMR recovers more genuine, irreducible multicollinear relationships than the original method and uncovers exact dependencies, namely perfect multicollinearity, as well as overlapping relationships that the original method either misses or reports in reducible form. Overall, SMR offers an accurate, scalable, and theoretically grounded approach to multicollinearity detection.
Authors
- Chih-Hua Hsu (ORCID: https://orcid.org/0000-0002-8175-436X)
- Ting-Yu Liao (ORCID: https://orcid.org/0000-0001-9143-1129)
Institutions
- Chung Yuan Christian University (TW)
Publication Details
- Journal
- Journal of the Chinese Institute of Engineers
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1080/02533839.2026.2727635
- Primary Topic
- Machine Learning and Data Classification
- Type
- article
- Field-Weighted Citation Impact
- 0.00