Kernel $$K$$-Means Clustering of Distributional Data
Abstract We consider the problem of clustering a sample of probability distributions from a random distribution on $$\mathbb R^d$$ R d . Our proposed partitioning method makes use of a symmetric, positive-definite kernel $$k$$ k and its associated reproducing kernel Hilbert space $$\mathscr {H}$$ H . By mapping each distribution to its corresponding kernel mean embedding in $$\mathscr {H}$$ H , we obtain a sample in this space where we carry out the $$K$$ K -means clustering procedure, which provides an unsupervised classification of the original sample. The procedure is simple and computationally feasible even for dimension $$d>1$$ d > 1 . The simulation studies provide insight into the choice of the kernel and its tuning parameter. The performance of the proposed clustering procedure is illustrated on a collection of synthetic aperture radar images and on meteorological data.
Authors
- Amparo Baı́llo (ORCID: https://orcid.org/0000-0003-1644-3992)
- José R. Berrendero (ORCID: https://orcid.org/0000-0003-0728-7748)
- Martín Sánchez-Signorini
Institutions
- Institute of Mathematical Sciences (ES)
- Universidad Carlos III de Madrid (ES)
- Universidad Autónoma de Madrid (ES)
Publication Details
- Journal
- Journal of Classification
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1007/s00357-026-09560-7
- Primary Topic
- Advanced Clustering Algorithms Research
- Type
- article
- Field-Weighted Citation Impact
- 0.00