A Big Data Analytics Framework for Corporate Financial Profiling: Unsupervised Machine Learning Applied to Ecuador’s Financial and Insurance Sector

The amount of financial information available through regulatory big data repositories, whose scale and update frequency are consistent with the defining characteristics of big data, represents an opportunity for implementing big data analytics and business intelligence in order to support corporate decision-making, although it presents challenges due to the correlated nature of high-dimensional financial indicators. This study proposes a big data analytics model that combines dimensionality reduction with unsupervised machine learning for the classification of financial profiles of firms belonging to Ecuador’s Financial and Insurance sector (sector K). A quantitative and non-experimental approach is employed with a dataset composed of a large panel of 7875 firm-year observations (2019–2024) from 2279 unique firms, extracted from the Superintendencia de Compañías, Valores y Seguros repository and including nine financial indicators related to size, profitability, liquidity, leverage, and firm age, with extreme values addressed through winsorization at the 1st and 99th percentiles rather than case deletion. PCA reduces the dimensionality of the data, while K-means clustering, validated with the elbow method, the Silhouette index, alternative clustering algorithms (PAM, hierarchical (Ward) clustering, and a Gaussian Mixture Model) and alternative values of K, segments 1894 firms that operate in 2024. PCA retains four principal components of the dataset—firm size, financial performance, operating structure, and capital structure and organizational maturity—that explain 76.04% of the total variance. Two interpretable corporate financial profiles with limited statistical separation (average Silhouette coefficient of 0.257) were identified by K-means clustering: a larger, older, less-leveraged profile (Cluster 1) and a smaller, younger, more highly leveraged profile with higher observed returns (Cluster 2). Beyond its regional contribution to the Ecuadorian financial and insurance sector, the study illustrates a fully reproducible big data analytics pipeline—including transparency and robustness diagnostics—that is applicable to other economic sectors and regulatory repositories that are continuously updated.

Authors

Institutions

Publication Details

Journal
Big Data and Cognitive Computing
Published
2026-09-28
DOI
https://doi.org/10.3390/bdcc10100332
Primary Topic
Financial Distress and Bankruptcy Prediction
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Big Data Analytics Framework for Corporate Financial Profiling: Unsupervised Machine Learning Applied to Ecuador’s Financial and Insurance Sector

María Estefanía Sánchez Pacheco, Carlos Gabriel Parrales Chóez, Fernando José Zambrano Farías, Robin Xavier Martínez Mayorga
Big Data and Cognitive Computing
Financial Distress and Bankruptcy Prediction
article

A Big Data Analytics Framework for Corporate Financial Profiling: Unsupervised Machine Learning Applied to Ecuador’s Financial and Insurance Sector

María Estefanía Sánchez Pacheco, Carlos Gabriel Parrales Chóez, Fernando José Zambrano Farías, Robin Xavier Martínez Mayorga
article en

Abstract

The amount of financial information available through regulatory big data repositories, whose scale and update frequency are consistent with the defining characteristics of big data, represents an opportunity for implementing big data analytics and business intelligence in order to support corporate decision-making, although it presents challenges due to the correlated nature of high-dimensional financial indicators. This study proposes a big data analytics model that combines dimensionality reduction with unsupervised machine learning for the classification of financial profiles of firms belonging to Ecuador’s Financial and Insurance sector (sector K). A quantitative and non-experimental approach is employed with a dataset composed of a large panel of 7875 firm-year observations (2019–2024) from 2279 unique firms, extracted from the Superintendencia de Compañías, Valores y Seguros repository and including nine financial indicators related to size, profitability, liquidity, leverage, and firm age, with extreme values addressed through winsorization at the 1st and 99th percentiles rather than case deletion. PCA reduces the dimensionality of the data, while K-means clustering, validated with the elbow method, the Silhouette index, alternative clustering algorithms (PAM, hierarchical (Ward) clustering, and a Gaussian Mixture Model) and alternative values of K, segments 1894 firms that operate in 2024. PCA retains four principal components of the dataset—firm size, financial performance, operating structure, and capital structure and organizational maturity—that explain 76.04% of the total variance. Two interpretable corporate financial profiles with limited statistical separation (average Silhouette coefficient of 0.257) were identified by K-means clustering: a larger, older, less-leveraged profile (Cluster 1) and a smaller, younger, more highly leveraged profile with higher observed returns (Cluster 2). Beyond its regional contribution to the Ecuadorian financial and insurance sector, the study illustrates a fully reproducible big data analytics pipeline—including transparency and robustness diagnostics—that is applicable to other economic sectors and regulatory repositories that are continuously updated.

Big Data and Cognitive ComputingVol. 10(10)
University of Guayaquil (EC), Universidad Internacional del Ecuador (EC)
Openalex Percentile: Top 4%
Financial Distress and Bankruptcy Prediction
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.