Spark: Phylogenetic Analysis Using Series-Parallel Resistor-Derived Features from K-Mer and Substring Positional Accumulation Sum

The transformation of genomic sequences into k-mer-based feature vectors offers an efficient and convenient means for interpreting genomic signals and analyzing their similarity. However, genomic attribute signals derived directly from k-mers may contain some random background. Existing studies have confirmed that the original k-mer signals can be purified by assigning weights, entropy-based quantification, specific patterns, or simulating the original signals. Here, we propose Spark, a novel feature extraction algorithm that treats each k-mer as a resistor: the positional accumulation sum of a k-mer represents its resistance, so the sequential arrangement of k-mers mimics resistors in series, while splitting a k-mer into two shorter substrings and combining them mimics resistors in parallel to simulate the k-mer signal from its background. To compare features derived from different perspectives, we evaluate feature vectors extracted from the original signal, the simulated signal, and the ratio of the two on six genomic datasets. Theoretical analysis shows that this ratio ranges within (0, 2), and experiments reveal that on average 0.898 of original signals fall in (0, 1), with the proportion of such ratios exceeding 0.965 in two-thirds of the datasets. We therefore apply an odds transformation to ratios greater than 1 and then nonlinearly normalize all ratios with the Sigmoid function. The resulting ratio-based features improve the quantification of genomic differences: in a comparison against the best-performing alignment-free method, Spark achieves the smallest RF distance on most datasets and is within 2 of the best on the remaining ones. Spark thus offers a novel and effective approach to extracting features from the positional information of k-mers for genomic sequence vectorization, with potential applicability to other genomic analyses.

Authors

Institutions

Publication Details

Journal
Mathematics
Published
2026-09-14
DOI
https://doi.org/10.3390/math14183327
Primary Topic
Fractal and DNA sequence analysis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Spark: Phylogenetic Analysis Using Series-Parallel Resistor-Derived Features from K-Mer and Substring Positional Accumulation Sum

Runbin Tang, Jingjing Zhang, Jianwen Huang, Zhifeng Xiao
Mathematics
Fractal and DNA sequence analysis
article

Spark: Phylogenetic Analysis Using Series-Parallel Resistor-Derived Features from K-Mer and Substring Positional Accumulation Sum

Runbin Tang, Jingjing Zhang, Jianwen Huang, Zhifeng Xiao
article en

Abstract

The transformation of genomic sequences into k-mer-based feature vectors offers an efficient and convenient means for interpreting genomic signals and analyzing their similarity. However, genomic attribute signals derived directly from k-mers may contain some random background. Existing studies have confirmed that the original k-mer signals can be purified by assigning weights, entropy-based quantification, specific patterns, or simulating the original signals. Here, we propose Spark, a novel feature extraction algorithm that treats each k-mer as a resistor: the positional accumulation sum of a k-mer represents its resistance, so the sequential arrangement of k-mers mimics resistors in series, while splitting a k-mer into two shorter substrings and combining them mimics resistors in parallel to simulate the k-mer signal from its background. To compare features derived from different perspectives, we evaluate feature vectors extracted from the original signal, the simulated signal, and the ratio of the two on six genomic datasets. Theoretical analysis shows that this ratio ranges within (0, 2), and experiments reveal that on average 0.898 of original signals fall in (0, 1), with the proportion of such ratios exceeding 0.965 in two-thirds of the datasets. We therefore apply an odds transformation to ratios greater than 1 and then nonlinearly normalize all ratios with the Sigmoid function. The resulting ratio-based features improve the quantification of genomic differences: in a comparison against the best-performing alignment-free method, Spark achieves the smallest RF distance on most datasets and is within 2 of the best on the remaining ones. Spark thus offers a novel and effective approach to extracting features from the positional information of k-mers for genomic sequence vectorization, with potential applicability to other genomic analyses.

MathematicsVol. 14(18)
Chongqing Normal University (CN)
Openalex Percentile: Top 18%
Fractal and DNA sequence analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.