MSI: A Mahalanobis‐Based Molecular Similarity Index for High‐Dimensional Embeddings

Quantifying molecular similarity is crucial for drug discovery and for exploring chemical space. A similarity assessment always combines two independent ingredients: a molecular representation and a similarity coefficient. The most common pairing, binary substructure fingerprints scored with the Tanimoto coefficient, depends strongly on fingerprint bit density and, because it compares unweighted sets of substructure identifiers, is blind to the multiplicity of repeated fragments and frequently returns ranking ties that obscure meaningful chemical relationships. Here, we introduce the Mahalanobis Similarity Index (MSI), which pairs continuous Mol2Vec embeddings with a covariance‐aware Mahalanobis distance ( d M ) and an associated Mahalanobis angle ( θ M ) to give a statistically grounded assessment that is invariant under invertible linear reparametrization of the descriptor space. We evaluated MSI on five chemically distinct reference compounds: aspirin, a salicylate nonsteroidal anti‐inflammatory drug (NSAID); aniline, an industrial aromatic amine; curcumin, a polyphenolic natural product; ibuprofen, a propionic‐acid NSAID; and digitoxin, a cardiac glycoside. Relative to the Tanimoto coefficient computed on ECFP4 fingerprints, MSI improves the analysis in three specific respects: it promotes chemically reasonable analogs that the fingerprint deprioritises; it resolves ranking ties, recovering between 7 and 10 distinct scores among the 10 nearest neighbors where Tanimoto recovers only 3–7; and its geometry varies systematically with HOMO–LUMO energy gaps in the QM9 dataset, indicating that the embedding tracks electronic structure even though it was trained on structural context alone. Polar plots and three‐dimensional similarity maps reveal anisotropy within the embedding space and define practical applicability domains for high‐similarity retrieval. MSI retains discriminatory power in the regime where the Tanimoto coefficient saturates near zero, and sparse peripheral regions suggest scaffold‐hopping opportunities. The dual radial–angular description supports hypothesis‐driven reasoning about how structural modifications shift electronic properties. MSI is computationally efficient, chemically interpretable, and offers a practical way to navigate high‐dimensional chemical space. Beyond drug discovery, it is applicable to materials science, toxicology, and chemical biology, wherever continuous molecular embeddings are used for property‐driven screening.

Authors

Institutions

Publication Details

Journal
Molecular Informatics
Published
2026-08-27
DOI
https://doi.org/10.1002/minf.70051
Primary Topic
Computational Drug Discovery Methods
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

MSI: A Mahalanobis‐Based Molecular Similarity Index for High‐Dimensional Embeddings

Felipe Aparicio, Leon Alday-Toledo, Roberto Bernal-Jaquez, Emiliano Montoya et al.
Molecular Informatics
Computational Drug Discovery Methods
article

MSI: A Mahalanobis‐Based Molecular Similarity Index for High‐Dimensional Embeddings

Felipe Aparicio, Leon Alday-Toledo, Roberto Bernal-Jaquez, Emiliano Montoya, Elliot Ridout‐Buhl
article en

Abstract

Quantifying molecular similarity is crucial for drug discovery and for exploring chemical space. A similarity assessment always combines two independent ingredients: a molecular representation and a similarity coefficient. The most common pairing, binary substructure fingerprints scored with the Tanimoto coefficient, depends strongly on fingerprint bit density and, because it compares unweighted sets of substructure identifiers, is blind to the multiplicity of repeated fragments and frequently returns ranking ties that obscure meaningful chemical relationships. Here, we introduce the Mahalanobis Similarity Index (MSI), which pairs continuous Mol2Vec embeddings with a covariance‐aware Mahalanobis distance ( d M ) and an associated Mahalanobis angle ( θ M ) to give a statistically grounded assessment that is invariant under invertible linear reparametrization of the descriptor space. We evaluated MSI on five chemically distinct reference compounds: aspirin, a salicylate nonsteroidal anti‐inflammatory drug (NSAID); aniline, an industrial aromatic amine; curcumin, a polyphenolic natural product; ibuprofen, a propionic‐acid NSAID; and digitoxin, a cardiac glycoside. Relative to the Tanimoto coefficient computed on ECFP4 fingerprints, MSI improves the analysis in three specific respects: it promotes chemically reasonable analogs that the fingerprint deprioritises; it resolves ranking ties, recovering between 7 and 10 distinct scores among the 10 nearest neighbors where Tanimoto recovers only 3–7; and its geometry varies systematically with HOMO–LUMO energy gaps in the QM9 dataset, indicating that the embedding tracks electronic structure even though it was trained on structural context alone. Polar plots and three‐dimensional similarity maps reveal anisotropy within the embedding space and define practical applicability domains for high‐similarity retrieval. MSI retains discriminatory power in the regime where the Tanimoto coefficient saturates near zero, and sparse peripheral regions suggest scaffold‐hopping opportunities. The dual radial–angular description supports hypothesis‐driven reasoning about how structural modifications shift electronic properties. MSI is computationally efficient, chemically interpretable, and offers a practical way to navigate high‐dimensional chemical space. Beyond drug discovery, it is applicable to materials science, toxicology, and chemical biology, wherever continuous molecular embeddings are used for property‐driven screening.

Molecular InformaticsVol. 45(9)
Universidad Autónoma de la Ciudad de México (MX), Universidad Autónoma Metropolitana (MX)
Universidad Autónoma Metropolitana
Reduced inequalities
Openalex Percentile: Top 8%
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.