Minimizer Density Revisited: Models and Multiminimizers

High-throughput sequence analysis commonly relies on k-mers and sampling schemes to improve scalability and data locality. Among these, local schemes are widely used and are typically evaluated by their density, the expected fraction of selected positions. Minimizers are the most common example, but recent near-tight lower bounds suggest diminishing room for improvement under the classical notion of density. Here, we revisit density and broaden its scope. First, we establish a direct link between density and the distance between consecutive selected positions. Under the sole assumption that these distances are identically distributed, we show that density is exactly the inverse of their expected distance, without assumptions on the selection mechanism. This provides a new way to analyze schemes beyond classical local models. Second, we introduce multiminimizers, a meta-scheme combining N minimizer schemes and selecting the candidate whose extends farthest. Multiminimizers are not local schemes, but under independent random component orders and locally distinct m-mers, their expected density converges exponentially to the optimum. Experiments with random minimizers and open-closed mod-minimizers confirm a controllable computation-density trade-off. Third, we introduce deduplicated density, measuring the fraction of distinct minimizers needed to cover all k-mers in a sequence set. We show that multiminimizers also improve this metric, prove that global optimization is NP-complete, and propose an effective local heuristic. Finally, we provide an efficient SIMD-accelerated Rust implementation and demonstrate reduced memory usage on core sequence-analysis tasks.

Authors

Institutions

Publication Details

Journal
Journal of Computational Biology
Published
2026-10-08
DOI
https://doi.org/10.1177/15578666261495123
Primary Topic
Algorithms and Data Compression
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Minimizer Density Revisited: Models and Multiminimizers

Antoine Limasset, Lucas Robidou, Camille Marchet, Florian Ingels et al.
Journal of Computational Biology
Algorithms and Data Compression
article

Minimizer Density Revisited: Models and Multiminimizers

Antoine Limasset, Lucas Robidou, Camille Marchet, Florian Ingels, Igor Martayan
article en

Abstract

High-throughput sequence analysis commonly relies on k-mers and sampling schemes to improve scalability and data locality. Among these, local schemes are widely used and are typically evaluated by their density, the expected fraction of selected positions. Minimizers are the most common example, but recent near-tight lower bounds suggest diminishing room for improvement under the classical notion of density. Here, we revisit density and broaden its scope. First, we establish a direct link between density and the distance between consecutive selected positions. Under the sole assumption that these distances are identically distributed, we show that density is exactly the inverse of their expected distance, without assumptions on the selection mechanism. This provides a new way to analyze schemes beyond classical local models. Second, we introduce multiminimizers, a meta-scheme combining N minimizer schemes and selecting the candidate whose extends farthest. Multiminimizers are not local schemes, but under independent random component orders and locally distinct m-mers, their expected density converges exponentially to the optimum. Experiments with random minimizers and open-closed mod-minimizers confirm a controllable computation-density trade-off. Third, we introduce deduplicated density, measuring the fraction of distinct minimizers needed to cover all k-mers in a sequence set. We show that multiminimizers also improve this metric, prove that global optimization is NP-complete, and propose an effective local heuristic. Finally, we provide an efficient SIMD-accelerated Rust implementation and demonstrate reduced memory usage on core sequence-analysis tasks.

Journal of Computational Biology
Centre National de la Recherche Scientifique (FR), Université de Lille (FR), Commissariat à l'Énergie Atomique et aux Énergies Alternatives (FR), Université Paris-Saclay (FR), Institut de Biologie Intégrative de la Cellule (FR), Centre de Recherche en Informatique, Signal et Automatique de Lille (FR), École Centrale de Lille (FR)
Openalex Percentile: Top 12%
Algorithms and Data Compression
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.