From known chemical space to unannotated metabolites: a cluster-guided retention-time driven framework for biologically informed annotation

{"INTRODUCTION:":[0],"Untargeted":[1],"metabolomics":[2],"often":[3],"results":[4],"in":[5,233],"a":[6,44,131,225],"significant":[7],"portion":[8],"of":[9,75,99,142],"unannotated":[10,32,117],"metabolites,":[11],"or":[12],"\\"metabolic":[13],"dark":[14,230],"matter,\\"":[15],"which":[16,143],"hinders":[17],"biological":[18,108],"interpretation.":[19],"OBJECTIVES:":[20],"A":[21],"two-step":[22],"analytical":[23],"approach":[24],"was":[25],"developed":[26],"to":[27,106,158,171],"systematically":[28],"prioritize":[29],"and":[30,87,96,101,110,126,163,184,206,223],"interpret":[31],"metabolites":[33,57,79,232],"using":[34,123],"plasma":[35],"LC-MS/MS":[36],"data":[37],"from":[38,121],"pregnant":[39],"women":[40],"with":[41],"obesity":[42],"as":[43],"biologically":[45,71],"relevant":[46,72],"test":[47],"dataset.":[48,77],"METHODS:":[49],"The":[50],"first":[51],"step":[52],"involved":[53],"clustering":[54],"1,021":[55],"known":[56],"into":[58],"ten":[59],"structurally":[60,160,216],"coherent":[61],"groups":[62],"based":[63],"on":[64],"the":[65,70,76,155,197],"Tanimoto":[66,151],"similarity,":[67],"thus":[68],"defining":[69],"chemical":[73],"space":[74],"These":[78],"were":[80,119,147],"further":[81,167],"characterized":[82],"by":[83,201],"Absorption,":[84],"Distribution,":[85],"Metabolism,":[86],"Excretion":[88],"(ADME)":[89],"profiling,":[90],"protein":[91],"target":[92],"prediction,":[93],"molecular":[94,124,127,199],"docking":[95],"Kyoto":[97],"Encyclopedia":[98],"Genes":[100],"Genomes":[102],"pathway":[103],"mapping":[104],"analysis,":[105],"establish":[107],"plausibility":[109],"functional":[111],"perspective.":[112],"Candidate":[113],"structures":[114,146],"for":[115,228],"1,836":[116],"features":[118],"retrieved":[120],"PubChem":[122],"formula":[125,200],"weight":[128],"matching":[129],"within":[130],"±0.5":[132],"Da":[133],"tolerance.":[134],"RESULTS:":[135],"This":[136,212],"search":[137],"yielded":[138],"569,115":[139],"candidate":[140,156,175],"structures,":[141],"368,197":[144],"unique":[145],"retained":[148],"after":[149],"curation.":[150],"coefficient":[152],"filtering":[153],"reduced":[154],"pool":[157],"19,868":[159],"plausible":[161],"candidates,":[162,218],"retention":[164,209],"time-based":[165],"prioritization":[166,191],"refined":[168],"this":[169],"set":[170],"418":[172],"high":[173],"confidence":[174],"annotations,":[176],"including":[177],"83":[178],"database-supported":[179],"candidates":[180],"identified":[181],"through":[182],"HMDB":[183],"LIPID":[185],"MAPS":[186],"structure":[187],"database":[188],"cross-referencing.":[189],"RT-based":[190],"effectively":[192],"distinguished":[193],"positional":[194],"isomers":[195],"sharing":[196],"same":[198],"incorporating":[202],"agreement":[203],"between":[204],"predicted":[205],"experimentally":[207],"observed":[208],"times.":[210],"CONCLUSION:":[211],"improved":[213],"discrimination":[214],"among":[215],"similar":[217],"expanded":[219],"metabolite":[220],"annotation":[221],"confidence,":[222],"provided":[224],"scalable":[226],"framework":[227],"prioritizing":[229],"matter":[231],"untargeted":[234],"metabolomics.":[235]}

Authors

Institutions

Publication Details

Journal
Metabolomics
Published
2026-08-26
DOI
https://doi.org/10.1007/s11306-026-02520-7
Primary Topic
Computational Drug Discovery Methods
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

From known chemical space to unannotated metabolites: a cluster-guided retention-time driven framework for biologically informed annotation

Renny S. Lan, Henry A. Paz, Sree V. Chintapalli, Hailemariam Abrha Assress et al.
Metabolomics
Computational Drug Discovery Methods
article

From known chemical space to unannotated metabolites: a cluster-guided retention-time driven framework for biologically informed annotation

Renny S. Lan, Henry A. Paz, Sree V. Chintapalli, Hailemariam Abrha Assress, Brian D. Piccolo, Elisabet Børsheim, Adepu Kiran Kumar, Colin D. Kay, Ahmad Mani‐Varnosfaderani, Dipendra Bhandari, Keith Henderson
article en

Abstract

INTRODUCTION: Untargeted metabolomics often results in a significant portion of unannotated metabolites, or "metabolic dark matter," which hinders biological interpretation. OBJECTIVES: A two-step analytical approach was developed to systematically prioritize and interpret unannotated metabolites using plasma LC-MS/MS data from pregnant women with obesity as a biologically relevant test dataset. METHODS: The first step involved clustering 1,021 known metabolites into ten structurally coherent groups based on the Tanimoto similarity, thus defining the biologically relevant chemical space of the dataset. These metabolites were further characterized by Absorption, Distribution, Metabolism, and Excretion (ADME) profiling, protein target prediction, molecular docking and Kyoto Encyclopedia of Genes and Genomes pathway mapping analysis, to establish biological plausibility and functional perspective. Candidate structures for 1,836 unannotated features were retrieved from PubChem using molecular formula and molecular weight matching within a ±0.5 Da tolerance. RESULTS: This search yielded 569,115 candidate structures, of which 368,197 unique structures were retained after curation. Tanimoto coefficient filtering reduced the candidate pool to 19,868 structurally plausible candidates, and retention time-based prioritization further refined this set to 418 high confidence candidate annotations, including 83 database-supported candidates identified through HMDB and LIPID MAPS structure database cross-referencing. RT-based prioritization effectively distinguished positional isomers sharing the same molecular formula by incorporating agreement between predicted and experimentally observed retention times. CONCLUSION: This improved discrimination among structurally similar candidates, expanded metabolite annotation confidence, and provided a scalable framework for prioritizing dark matter metabolites in untargeted metabolomics.

MetabolomicsVol. 22(5)
Arkansas Children's Nutrition Center (US), University of Arkansas for Medical Sciences (US)
U.S. Department of Agriculture, Agricultural Research Service
Zero hunger
Openalex Percentile: Top 9%
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.