Feature selection in cluster analysis for functional data with application to the labour market in Poland
The aim of this article is to apply methods of cluster analysis to functional data in order to determine if there are homogeneous groups among the provinces of Poland in terms of 20 features related to the labour market. In some existing studies all available features are used for cluster analysis, even though some of them may be insignificant for the task. Irrelevant features have a negative impact on the complexity and performance of clustering algorithms. In this study, before clusters of provinces were identified, features significantly differentiating provinces were selected. Our new feature selection criterion involves the use of coefficients of functional distance correlation and functional Hilbert-Schmidt correlation between the features characterizing provinces and their labels. The resulting clusters were nearly identical regardless of whether all 20 features were used or just the 9 significant ones, which indicates that irrelevant features do not contribute much to cluster identification. The methodology proposed in this article is flexible and can be successfully used in other research areas including, inter alia, poverty, socio-economic development or disability.
Authors
- Waldemar Wołyński (ORCID: https://orcid.org/0000-0002-0777-9163)
- Marcin Szymkowiak (ORCID: https://orcid.org/0000-0003-3432-4364)
- Mirosław Krzyśko (ORCID: https://orcid.org/0000-0001-8075-4432)
Institutions
- Poznań University of Economics and Business (PL)
- Uniwersytet Kaliski im. Prezydenta Stanisława Wojciechowskiego (PL)
- Adam Mickiewicz University in Poznań (PL)
Publication Details
- Journal
- Metrika
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1007/s00184-026-01050-5
- Primary Topic
- Advanced Clustering Algorithms Research
- Type
- article
- Field-Weighted Citation Impact
- 0.00