Hyb-Adam-UA: additivity-aware refinement of minimax-initialized mtDNA distance matrices

Mitochondrial DNA (mtDNA) distance matrices are standard inputs for distance-based phylogenetic inference. Missing entries can affect both topology reconstruction and branch-length estimation, whereas generic matrix-completion methods do not explicitly promote the tree-metric structure relevant to phylogenetic interpretation. We propose Hyb-Adam-UA (hybrid Adam, ultrametrically initialized and additivity-aware), a two-stage completion method that initializes missing entries by minimax-path distances on the observed graph and then refines only those entries using a four-point additivity objective with a triangle-inequality guard, while preserving all observed distances. We evaluated Hyb-Adam-UA on two 15 × 15 mtDNA benchmarks: a closely related Cercopithecidae dataset and a taxonomically heterogeneous primate dataset. Complete reference matrices were constructed from MAFFT multiple-sequence alignments using pairwise-deletion p -distances. Symmetric missingness masks were applied at 30%, 50%, 65%, and 85% missingness, with 30 replicates per level. Hyb-Adam-UA and its Stage 1-only ablation were compared with MW ⋆ -proj, NJ ⋆ -proj, LRMC, KNN-impute, and MDS-SMACOF using hidden-entry error, topology, patristic-distance, and branch-length criteria. For the heterogeneous dataset, the Stage 2 refinement significantly reduced hidden-entry RMSE relative to Stage 1 at 30%, 50%, and 65% missingness and relative to MW ⋆ -proj at 30%, 65%, and 85%. It also improved several branch-length results. For the Cercopithecidae dataset, however, the refinement provided no consistent advantage and was inferior to Stage 1 in some settings. Improvements in hidden-entry reconstruction produced only limited and inconsistent improvements in tree topology. A five-replicate synthetic 30 × 30 benchmark further demonstrated the effectiveness of Hyb-Adam-UA beyond the empirical 15 × 15 setting: among methods successful in all five replicates, it achieved the lowest mean hidden-entry RMSE at three of the four missingness levels and the lowest mean MAE at all four. Hyb-Adam-UA provides an additivity-aware framework for completing partially observed phylogenetic distance matrices without imposing a strict molecular-clock assumption. Its benefit is dataset-dependent: the method can improve hidden-distance reconstruction and branch-length estimation for taxonomically heterogeneous data, but the minimax-path initialization is itself a strong completion method. Lower matrix-reconstruction error does not necessarily produce a more accurate phylogenetic topology; matrix-level, branch-length, and topology criteria should therefore be evaluated separately.

Authors

Institutions

Publication Details

Journal
BMC Bioinformatics
Published
2026-09-16
DOI
https://doi.org/10.1186/s12859-026-06629-3
Primary Topic
DNA and Biological Computing
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Hyb-Adam-UA: additivity-aware refinement of minimax-initialized mtDNA distance matrices

Boris Melnikov, Ye Zhang, Dmitrii Chaikovskii, Yuehong Zhao et al.
BMC Bioinformatics
DNA and Biological Computing
article

Hyb-Adam-UA: additivity-aware refinement of minimax-initialized mtDNA distance matrices

Boris Melnikov, Ye Zhang, Dmitrii Chaikovskii, Yuehong Zhao, Weilai Qu
article en

Abstract

Mitochondrial DNA (mtDNA) distance matrices are standard inputs for distance-based phylogenetic inference. Missing entries can affect both topology reconstruction and branch-length estimation, whereas generic matrix-completion methods do not explicitly promote the tree-metric structure relevant to phylogenetic interpretation. We propose Hyb-Adam-UA (hybrid Adam, ultrametrically initialized and additivity-aware), a two-stage completion method that initializes missing entries by minimax-path distances on the observed graph and then refines only those entries using a four-point additivity objective with a triangle-inequality guard, while preserving all observed distances. We evaluated Hyb-Adam-UA on two 15 × 15 mtDNA benchmarks: a closely related Cercopithecidae dataset and a taxonomically heterogeneous primate dataset. Complete reference matrices were constructed from MAFFT multiple-sequence alignments using pairwise-deletion p -distances. Symmetric missingness masks were applied at 30%, 50%, 65%, and 85% missingness, with 30 replicates per level. Hyb-Adam-UA and its Stage 1-only ablation were compared with MW ⋆ -proj, NJ ⋆ -proj, LRMC, KNN-impute, and MDS-SMACOF using hidden-entry error, topology, patristic-distance, and branch-length criteria. For the heterogeneous dataset, the Stage 2 refinement significantly reduced hidden-entry RMSE relative to Stage 1 at 30%, 50%, and 65% missingness and relative to MW ⋆ -proj at 30%, 65%, and 85%. It also improved several branch-length results. For the Cercopithecidae dataset, however, the refinement provided no consistent advantage and was inferior to Stage 1 in some settings. Improvements in hidden-entry reconstruction produced only limited and inconsistent improvements in tree topology. A five-replicate synthetic 30 × 30 benchmark further demonstrated the effectiveness of Hyb-Adam-UA beyond the empirical 15 × 15 setting: among methods successful in all five replicates, it achieved the lowest mean hidden-entry RMSE at three of the four missingness levels and the lowest mean MAE at all four. Hyb-Adam-UA provides an additivity-aware framework for completing partially observed phylogenetic distance matrices without imposing a strict molecular-clock assumption. Its benefit is dataset-dependent: the method can improve hidden-distance reconstruction and branch-length estimation for taxonomically heterogeneous data, but the minimax-path initialization is itself a strong completion method. Lower matrix-reconstruction error does not necessarily produce a more accurate phylogenetic topology; matrix-level, branch-length, and topology criteria should therefore be evaluated separately.

BMC Bioinformatics
Beijing Institute of Technology (CN), Shenzhen University (CN), University Town of Shenzhen (CN), Tsinghua–Berkeley Shenzhen Institute (CN), Shenzhen Technology University (CN)
National Natural Science Foundation of China, National Key Research and Development Program of China, Shenzhen Science and Technology Innovation Program
Openalex Percentile: Top 18%
DNA and Biological Computing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.