Discriminant analysis-guided multi-split segmentation for expectation maximization initialization in time-series data

In this paper, a discriminant analysis–guided multi-split segmentation approach referred to as K-between class segmentation (K-BCS) is presented, which is developed as an effective initialization strategy for the expectation–maximization (EM) algorithm in time-series data. The proposed approach is aimed at improving parameter estimation and convergence stability in the EM algorithm. Although EM is widely used for mixture modeling, its performance is highly sensitive to its initialization estimates, which often results in slow convergence, convergence to suboptimal local maxima, and unstable likelihood behavior. To address these limitations, K-BCS-EM leverages a proposed between-class measure to identify intrinsic data structures and generate informed initial parameter estimates. Unlike random or centroid-based methods, the proposed approach is fully data-driven and captures separation in both location and dispersion, thereby providing more reliable starting points for the EM algorithm. The proposed approach was evaluated via extensive Monte Carlo experiments across both synthetic and real-world datasets. The results obtained show that the K-BCS-EM approach consistently outperformed conventional initialization strategies, such as random sampling, random partitioning, jittered sampling, and k-means-based methods particularly in terms of their likelihood convergence, mean and variance RMSE performance. The K-BCS-EM demonstrates strong robustness in challenging scenarios such as in heavy-tailed distributions and highly overlapping mixtures while maintaining scalability as the number of components increases. Although the proposed approach introduces a modest computational overhead due to its preprocessing stage, the significant improvements in estimation accuracy, convergence stability, and robustness highlight K-BCS-EM as an effective initialization method for EM-based mixture modeling in time series data.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-09
DOI
https://doi.org/10.1038/s41598-026-70509-2
Primary Topic
Bayesian Methods and Mixture Models
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Discriminant analysis-guided multi-split segmentation for expectation maximization initialization in time-series data

Ipeleng Labius Machele, Adeiza James Onumanyi, Adnan M. Abu‐Mahfouz, Anish Kurien
Scientific Reports
Bayesian Methods and Mixture Models
article

Discriminant analysis-guided multi-split segmentation for expectation maximization initialization in time-series data

Ipeleng Labius Machele, Adeiza James Onumanyi, Adnan M. Abu‐Mahfouz, Anish Kurien
article en

Abstract

In this paper, a discriminant analysis–guided multi-split segmentation approach referred to as K-between class segmentation (K-BCS) is presented, which is developed as an effective initialization strategy for the expectation–maximization (EM) algorithm in time-series data. The proposed approach is aimed at improving parameter estimation and convergence stability in the EM algorithm. Although EM is widely used for mixture modeling, its performance is highly sensitive to its initialization estimates, which often results in slow convergence, convergence to suboptimal local maxima, and unstable likelihood behavior. To address these limitations, K-BCS-EM leverages a proposed between-class measure to identify intrinsic data structures and generate informed initial parameter estimates. Unlike random or centroid-based methods, the proposed approach is fully data-driven and captures separation in both location and dispersion, thereby providing more reliable starting points for the EM algorithm. The proposed approach was evaluated via extensive Monte Carlo experiments across both synthetic and real-world datasets. The results obtained show that the K-BCS-EM approach consistently outperformed conventional initialization strategies, such as random sampling, random partitioning, jittered sampling, and k-means-based methods particularly in terms of their likelihood convergence, mean and variance RMSE performance. The K-BCS-EM demonstrates strong robustness in challenging scenarios such as in heavy-tailed distributions and highly overlapping mixtures while maintaining scalability as the number of components increases. Although the proposed approach introduces a modest computational overhead due to its preprocessing stage, the significant improvements in estimation accuracy, convergence stability, and robustness highlight K-BCS-EM as an effective initialization method for EM-based mixture modeling in time series data.

Scientific Reports
Tshwane University of Technology (ZA), Council for Scientific and Industrial Research (ZA)
Reduced inequalities
Openalex Percentile: Top 8%
Bayesian Methods and Mixture Models
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.