Optimization-Driven Cluster-Weighted Clusterwise Linear Regression

Clusterwise Linear Regression (CLR) integrates clustering and regression to uncover complex relationships within heterogeneous datasets by simultaneously partitioning data and fitting cluster-specific models. This dual capability offers a significant advantage over traditional regression methods, which assume homogeneity across the dataset. However, accurate data approximation often requires many linear functions, which can lead to overfitting. Existing CLR methods are not always efficient for prediction and may fail to partition the explanatory variable space distinctly. To address these limitations, we introduce a novel Cluster-Weighted CLR (CW-CLR) framework based on a nonsmooth, nonconvex optimization formulation that simultaneously solves the clustering and regression tasks. The framework assigns weights to explanatory variables under a sum-to-one constraint and incorporates a scaling parameter to balance clustering and regression errors, yielding a weighted prediction rule that improves robustness, mitigates overfitting, and better exploits cluster structure. To alleviate the challenges posed by nonconvexity, we propose an efficient incremental approach that begins with a single linear model and progressively constructs clusters, aided by a specialized initialization scheme designed to enhance optimization performance. We also propose an improved prediction rule for Optimization-Driven CW-CLR. Numerical experiments on 12 real-world datasets with varying sizes and structures show that CW-CLR achieves competitive or superior predictive performance compared with a broad range of state-of-the-art methods. In particular, CW-CLR attains the best RMSE and MAE on several benchmark datasets and delivers strong goodness-of-fit, with leading \(R^{2}\) and correlation values on multiple datasets. A non-parametric statistical analysis based on the Friedman test and post-hoc analysis confirms that CW-CLR significantly outperforms several competitors and is statistically comparable to the alternatives. Overall, the results demonstrate that CW-CLR provides a robust, accurate, and practically efficient alternative for regression on heterogeneous data.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Knowledge Discovery from Data
Published
2026-09-30
DOI
https://doi.org/10.1145/3850158
Primary Topic
Stochastic Gradient Optimization Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Optimization-Driven Cluster-Weighted Clusterwise Linear Regression

Mali Abdollahian, Adil Bagirov, Nargiz Sultanova, Sona Taheri et al.
ACM Transactions on Knowledge Discovery from Data
Stochastic Gradient Optimization Techniques
article

Optimization-Driven Cluster-Weighted Clusterwise Linear Regression

Mali Abdollahian, Adil Bagirov, Nargiz Sultanova, Sona Taheri, Nelusha Anne Perera
article en

Abstract

Clusterwise Linear Regression (CLR) integrates clustering and regression to uncover complex relationships within heterogeneous datasets by simultaneously partitioning data and fitting cluster-specific models. This dual capability offers a significant advantage over traditional regression methods, which assume homogeneity across the dataset. However, accurate data approximation often requires many linear functions, which can lead to overfitting. Existing CLR methods are not always efficient for prediction and may fail to partition the explanatory variable space distinctly. To address these limitations, we introduce a novel Cluster-Weighted CLR (CW-CLR) framework based on a nonsmooth, nonconvex optimization formulation that simultaneously solves the clustering and regression tasks. The framework assigns weights to explanatory variables under a sum-to-one constraint and incorporates a scaling parameter to balance clustering and regression errors, yielding a weighted prediction rule that improves robustness, mitigates overfitting, and better exploits cluster structure. To alleviate the challenges posed by nonconvexity, we propose an efficient incremental approach that begins with a single linear model and progressively constructs clusters, aided by a specialized initialization scheme designed to enhance optimization performance. We also propose an improved prediction rule for Optimization-Driven CW-CLR. Numerical experiments on 12 real-world datasets with varying sizes and structures show that CW-CLR achieves competitive or superior predictive performance compared with a broad range of state-of-the-art methods. In particular, CW-CLR attains the best RMSE and MAE on several benchmark datasets and delivers strong goodness-of-fit, with leading \(R^{2}\) and correlation values on multiple datasets. A non-parametric statistical analysis based on the Friedman test and post-hoc analysis confirms that CW-CLR significantly outperforms several competitors and is statistically comparable to the alternatives. Overall, the results demonstrate that CW-CLR provides a robust, accurate, and practically efficient alternative for regression on heterogeneous data.

ACM Transactions on Knowledge Discovery from Data
RMIT University (AU)
Openalex Percentile: Top 9%
Stochastic Gradient Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.