Poisson subspace clustering: focusing on the essentials in count data

Abstract Count data represented as a matrix of non-negative integer values, such as contingency tables, are prevalent across diverse domains. When clustering such data sets, specific methods are required, as generic algorithms often fail to consider their unique distributional properties, leading to unreliable outputs. An effective strategy is to use well-established statistical models such as the Poisson and negative binomial distributions. We present 3CPO, a clustering algorithm based on statistically solid modeling of count data. In addition to the cluster labels, it identifies a subset of relevant columns, enhancing the interpretability of the results. We propose a simple iterative algorithm that maximizes the posterior probability to find good clustering solutions and discuss its properties. Extensive experiments demonstrate its ability to define high-quality clusters within associated subspaces for various data domains, ranging from gene expressions and texts to economics. Our findings suggest that 3CPO is a robust solution for clustering count data in a statistically sound and interpretable manner. Our code is available at https://github.com/collinleiber/3CPO .

Authors

Institutions

Publication Details

Journal
Data Mining and Knowledge Discovery
Published
2026-09-21
DOI
https://doi.org/10.1007/s10618-026-01230-x
Primary Topic
Advanced Clustering Algorithms Research
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Poisson subspace clustering: focusing on the essentials in count data

Heikki Mannila, Collin Leiber, Kai Puolamäki
Data Mining and Knowledge Discovery
Advanced Clustering Algorithms Research
article

Poisson subspace clustering: focusing on the essentials in count data

Heikki Mannila, Collin Leiber, Kai Puolamäki
article en

Abstract

Abstract Count data represented as a matrix of non-negative integer values, such as contingency tables, are prevalent across diverse domains. When clustering such data sets, specific methods are required, as generic algorithms often fail to consider their unique distributional properties, leading to unreliable outputs. An effective strategy is to use well-established statistical models such as the Poisson and negative binomial distributions. We present 3CPO, a clustering algorithm based on statistically solid modeling of count data. In addition to the cluster labels, it identifies a subset of relevant columns, enhancing the interpretability of the results. We propose a simple iterative algorithm that maximizes the posterior probability to find good clustering solutions and discuss its properties. Extensive experiments demonstrate its ability to define high-quality clusters within associated subspaces for various data domains, ranging from gene expressions and texts to economics. Our findings suggest that 3CPO is a robust solution for clustering count data in a statistically sound and interpretable manner. Our code is available at https://github.com/collinleiber/3CPO .

Data Mining and Knowledge DiscoveryVol. 40(6)
University of Helsinki (FI), Aalto University (FI)
Teknologiateollisuuden 100-Vuotisjuhlasäätiö
Openalex Percentile: Top 19%
Advanced Clustering Algorithms Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Poisson subspace clustering: focusing on the essentials in count data — Heikki Mannila, Collin Leiber, et al. · Data Mining and Knowledge Discovery (2026) | TGRS Research Map | TGRS