From known chemical space to unannotated metabolites: a cluster-guided retention-time driven framework for biologically informed annotation.

Dipendra Bhandari, Henry A Paz, Keith Henderson, Kiran Kumar Adepu, Ahmad Mani-Varnosfaderani, Hailemariam Abrha Assress, Brian D Piccolo, Renny S Lan, Elisabet Børsheim, Colin D Kay, Sree V Chintapalli

Journal: Metabolomics : Official journal of the Metabolomic Society 2026;22(5):

PMID: 42649459

Abstract

INTRODUCTION

Untargeted metabolomics often results in a significant portion of unannotated metabolites, or "metabolic dark matter," which hinders biological interpretation.

OBJECTIVES

A two-step analytical approach was developed to systematically prioritize and interpret unannotated metabolites using plasma LC-MS/MS data from pregnant women with obesity as a biologically relevant test dataset.

METHODS

The first step involved clustering 1,021 known metabolites into ten structurally coherent groups based on the Tanimoto similarity, thus defining the biologically relevant chemical space of the dataset. These metabolites were further characterized by Absorption, Distribution, Metabolism, and Excretion (ADME) profiling, protein target prediction, molecular docking and Kyoto Encyclopedia of Genes and Genomes pathway mapping analysis, to establish biological plausibility and functional perspective. Candidate structures for 1,836 unannotated features were retrieved from PubChem using molecular formula and molecular weight matching within a ±0.5 Da tolerance.

RESULTS

This search yielded 569,115 candidate structures, of which 368,197 unique structures were retained after curation. Tanimoto coefficient filtering reduced the candidate pool to 19,868 structurally plausible candidates, and retention time-based prioritization further refined this set to 418 high confidence candidate annotations, including 83 database-supported candidates identified through HMDB and LIPID MAPS structure database cross-referencing. RT-based prioritization effectively distinguished positional isomers sharing the same molecular formula by incorporating agreement between predicted and experimentally observed retention times.

CONCLUSION

This improved discrimination among structurally similar candidates, expanded metabolite annotation confidence, and provided a scalable framework for prioritizing dark matter metabolites in untargeted metabolomics.

© 2026. The Author(s).

Address: Arkansas Children's Nutrition Center and Arkansas Children's Research Institute, 15 Children's Way, Little Rock, AR, 72202, USA.; Department of Pediatrics, University of Arkansas for Medical Sciences, Little Rock, AR, USA.; Arkansas Children's Nutrition Center and Arkansas Children's Research Institute, 15 Children's Way, Little Rock, AR, 72202, USA.; Department of Pediatrics, University of Arkansas for Medical Sciences, Little Rock, AR, USA.; Arkansas Children's Nutrition Center and Arkansas Children's Research Institute, 15 Children's Way, Little Rock, AR, 72202, USA. [email protected].; Department of Pediatrics, University of Arkansas for Medical Sciences, Little Rock, AR, USA. [email protected].
Bant logo

© Copyright 2026, Nutrition Evidence

NED wishes to thank the following organisations for their support:

We use cookies to improve your experience and analyze site traffic with Google Analytics. By continuing to use our site, you agree to our use of cookies. Learn more.