A Big Data Analytics Framework for Corporate Financial Profiling: Unsupervised Machine Learning Applied to Ecuador’s Financial and Insurance Sector
The amount of financial information available through regulatory big data repositories, whose scale and update frequency are consistent with the defining characteristics of big data, represents an opportunity for implementing big data analytics and business intelligence in order to support corporate decision-making, although it presents challenges due to the correlated nature of high-dimensional financial indicators. This study proposes a big data analytics model that combines dimensionality reduction with unsupervised machine learning for the classification of financial profiles of firms belonging to Ecuador’s Financial and Insurance sector (sector K). A quantitative and non-experimental approach is employed with a dataset composed of a large panel of 7875 firm-year observations (2019–2024) from 2279 unique firms, extracted from the Superintendencia de Compañías, Valores y Seguros repository and including nine financial indicators related to size, profitability, liquidity, leverage, and firm age, with extreme values addressed through winsorization at the 1st and 99th percentiles rather than case deletion. PCA reduces the dimensionality of the data, while K-means clustering, validated with the elbow method, the Silhouette index, alternative clustering algorithms (PAM, hierarchical (Ward) clustering, and a Gaussian Mixture Model) and alternative values of K, segments 1894 firms that operate in 2024. PCA retains four principal components of the dataset—firm size, financial performance, operating structure, and capital structure and organizational maturity—that explain 76.04% of the total variance. Two interpretable corporate financial profiles with limited statistical separation (average Silhouette coefficient of 0.257) were identified by K-means clustering: a larger, older, less-leveraged profile (Cluster 1) and a smaller, younger, more highly leveraged profile with higher observed returns (Cluster 2). Beyond its regional contribution to the Ecuadorian financial and insurance sector, the study illustrates a fully reproducible big data analytics pipeline—including transparency and robustness diagnostics—that is applicable to other economic sectors and regulatory repositories that are continuously updated.
Authors
- María Estefanía Sánchez Pacheco (ORCID: https://orcid.org/0000-0002-2469-9018)
- Carlos Gabriel Parrales Chóez (ORCID: https://orcid.org/0000-0001-7006-4614)
- Fernando José Zambrano Farías (ORCID: https://orcid.org/0000-0001-6384-3353)
- Robin Xavier Martínez Mayorga (ORCID: https://orcid.org/0000-0002-8193-6251)
Institutions
- University of Guayaquil (EC)
- Universidad Internacional del Ecuador (EC)
Publication Details
- Journal
- Big Data and Cognitive Computing
- Published
- 2026-09-28
- DOI
- https://doi.org/10.3390/bdcc10100332
- Primary Topic
- Financial Distress and Bankruptcy Prediction
- Type
- article
- Field-Weighted Citation Impact
- 0.00