A theoretical and empirical evaluation of product connectivity descriptors for improving antiviral QSAR and QSPR prediction

Degree-based topological indices are attractive 2D molecular descriptors because they are interpretable and inexpensive, yet classical additive edge aggregations can be dominated by highly connected substructures. To address this imbalance, we introduce product-connectivity indices, which reweight edge contributions by the multiplicative normalization \\({w}_{PC}(u,v)=1/\\sqrt{{\\delta }_{u}{\\delta }_{v}}\\) to attenuate hub effects while emphasizing endpoint compatibility. We formulate a 14-member product-connectivity family, establish relations and degree-extreme bounds against the corresponding unweighted indices, and benchmark PCI under leakage-controlled feature-block ablations. The empirical study uses 10,558 curated ChEMBL antiviral records (10,486 unique compounds), 12 QSPR endpoints, antiviral p E C 50 QSAR, and a fixed five-regressor suite evaluated by hold-out testing for QSPR and 5-fold cross-validation for QSAR. At an equal 14-feature budget, PCI improved QSPR over classical indices (best R 2 : 0.8662 vs. 0.8589; RMSE: 9.4834 vs. 9.8811; MAE: 5.9517 vs. 6.1045) and gave a modest QSAR advantage ( R 2 : 0.4887 vs. 0.4839). Rich descriptor sets reached R 2 ≈ 0.71, whereas adding PCI to RDKit-rich representations was essentially neutral, indicating that the practical benefit is concentrated in compressed descriptor settings. The novelty is a general product-normalization that can be applied to an existing degree kernel rather than defining one additional fixed index. This construction yields paired weighted/unweighted descriptor families with explicit scale bounds and is evaluated under matched feature budgets. Its practical advantage is a compact, interpretable 2D representation that adds measurable signal without requiring conformer generation or quantum-chemical calculations. The observed gains are dataset-dependent and should not be interpreted as universal superiority over richer descriptor spaces. PCI provides a low-cost, theory-grounded descriptor family with a measurable advantage in constrained feature settings and should be viewed as complementary to, rather than a replacement for, richer molecular representations.

Authors

Institutions

Publication Details

Journal
BMC Bioinformatics
Published
2026-09-21
DOI
https://doi.org/10.1186/s12859-026-06654-2
Primary Topic
Computational Drug Discovery Methods
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A theoretical and empirical evaluation of product connectivity descriptors for improving antiviral QSAR and QSPR prediction

Yusuf Zeren, Zaied Alhaj, Azzam Altairi, Mohammed Alsharafi
BMC Bioinformatics
Computational Drug Discovery Methods
article

A theoretical and empirical evaluation of product connectivity descriptors for improving antiviral QSAR and QSPR prediction

Yusuf Zeren, Zaied Alhaj, Azzam Altairi, Mohammed Alsharafi
article en

Abstract

Degree-based topological indices are attractive 2D molecular descriptors because they are interpretable and inexpensive, yet classical additive edge aggregations can be dominated by highly connected substructures. To address this imbalance, we introduce product-connectivity indices, which reweight edge contributions by the multiplicative normalization \({w}_{PC}(u,v)=1/\sqrt{{\delta }_{u}{\delta }_{v}}\) to attenuate hub effects while emphasizing endpoint compatibility. We formulate a 14-member product-connectivity family, establish relations and degree-extreme bounds against the corresponding unweighted indices, and benchmark PCI under leakage-controlled feature-block ablations. The empirical study uses 10,558 curated ChEMBL antiviral records (10,486 unique compounds), 12 QSPR endpoints, antiviral p E C 50 QSAR, and a fixed five-regressor suite evaluated by hold-out testing for QSPR and 5-fold cross-validation for QSAR. At an equal 14-feature budget, PCI improved QSPR over classical indices (best R 2 : 0.8662 vs. 0.8589; RMSE: 9.4834 vs. 9.8811; MAE: 5.9517 vs. 6.1045) and gave a modest QSAR advantage ( R 2 : 0.4887 vs. 0.4839). Rich descriptor sets reached R 2 ≈ 0.71, whereas adding PCI to RDKit-rich representations was essentially neutral, indicating that the practical benefit is concentrated in compressed descriptor settings. The novelty is a general product-normalization that can be applied to an existing degree kernel rather than defining one additional fixed index. This construction yields paired weighted/unweighted descriptor families with explicit scale bounds and is evaluated under matched feature budgets. Its practical advantage is a compact, interpretable 2D representation that adds measurable signal without requiring conformer generation or quantum-chemical calculations. The observed gains are dataset-dependent and should not be interpreted as universal superiority over richer descriptor spaces. PCI provides a low-cost, theory-grounded descriptor family with a measurable advantage in constrained feature settings and should be viewed as complementary to, rather than a replacement for, richer molecular representations.

BMC Bioinformatics
University of Science and Technology (YE), Sana'a University (YE), Yıldız Technical University (TR), Istanbul University-Cerrahpaşa (TR), Istanbul Commerce University (TR), University of Aden (YE)
Openalex Percentile: Top 9%
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.