A theoretical and empirical evaluation of product connectivity descriptors for improving antiviral QSAR and QSPR prediction
Degree-based topological indices are attractive 2D molecular descriptors because they are interpretable and inexpensive, yet classical additive edge aggregations can be dominated by highly connected substructures. To address this imbalance, we introduce product-connectivity indices, which reweight edge contributions by the multiplicative normalization \\({w}_{PC}(u,v)=1/\\sqrt{{\\delta }_{u}{\\delta }_{v}}\\) to attenuate hub effects while emphasizing endpoint compatibility. We formulate a 14-member product-connectivity family, establish relations and degree-extreme bounds against the corresponding unweighted indices, and benchmark PCI under leakage-controlled feature-block ablations. The empirical study uses 10,558 curated ChEMBL antiviral records (10,486 unique compounds), 12 QSPR endpoints, antiviral p E C 50 QSAR, and a fixed five-regressor suite evaluated by hold-out testing for QSPR and 5-fold cross-validation for QSAR. At an equal 14-feature budget, PCI improved QSPR over classical indices (best R 2 : 0.8662 vs. 0.8589; RMSE: 9.4834 vs. 9.8811; MAE: 5.9517 vs. 6.1045) and gave a modest QSAR advantage ( R 2 : 0.4887 vs. 0.4839). Rich descriptor sets reached R 2 ≈ 0.71, whereas adding PCI to RDKit-rich representations was essentially neutral, indicating that the practical benefit is concentrated in compressed descriptor settings. The novelty is a general product-normalization that can be applied to an existing degree kernel rather than defining one additional fixed index. This construction yields paired weighted/unweighted descriptor families with explicit scale bounds and is evaluated under matched feature budgets. Its practical advantage is a compact, interpretable 2D representation that adds measurable signal without requiring conformer generation or quantum-chemical calculations. The observed gains are dataset-dependent and should not be interpreted as universal superiority over richer descriptor spaces. PCI provides a low-cost, theory-grounded descriptor family with a measurable advantage in constrained feature settings and should be viewed as complementary to, rather than a replacement for, richer molecular representations.
Authors
- Yusuf Zeren (ORCID: https://orcid.org/0000-0001-8346-2208)
- Zaied Alhaj (ORCID: https://orcid.org/0009-0007-8820-5009)
- Azzam Altairi
- Mohammed Alsharafi
Institutions
- University of Science and Technology (YE)
- Sana'a University (YE)
- Yıldız Technical University (TR)
- Istanbul University-Cerrahpaşa (TR)
- Istanbul Commerce University (TR)
- University of Aden (YE)
Publication Details
- Journal
- BMC Bioinformatics
- Published
- 2026-09-21
- DOI
- https://doi.org/10.1186/s12859-026-06654-2
- Primary Topic
- Computational Drug Discovery Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00