SpectralMol: Spectral Graph Bisection Achieves Universal Molecular Fragment Decomposition for Drug-Like Molecule Generation
Abstract Generating valid, diverse, and drug-like molecules remains a central challenge in computational drug discovery. A key bottleneck for fragment-based generative models is decomposition coverage: existing junction-tree methods fail on bridged bicyclic and spiro ring systems common in drug-like scaffolds. Our primary contribution is Fiedler-BFS (breadth-first search) molecular decomposition, a spectral bisection algorithm achieving 100% decomposition success on all 250,000 ZINC-250K molecules, including bridged and spiro systems where junction-tree methods fail. A decomposition ablation (5,000 molecules, 5 seeds) confirms that spectral structure drives the improvement: 91.4% vs 7.1% for random bisection. Compared to BRICS (break retrosynthetically interesting chemical substructures; a rule-based retrosynthesis method), Fiedler-BFS produces fewer fragments per molecule (2.31 vs 5.13), better suited for hierarchical generative modeling. As a secondary contribution, we instantiate the decomposition in a Tree-VAE (variational autoencoder) with classifier-free guidance over molecular properties, Inverse Multi-Quadratic MMD (maximum mean discrepancy) regularization, and gradient-based latent-space steering. Under a matched protocol in which the identical Tree-VAE is retrained while only the cut-bond set varies, Fiedler-BFS yields shorter, more learnable fragment trees than BRICS and junction-tree decomposition (validation reconstruction loss of 0.012 vs 0.048 vs 0.083). On ZINC-250K, the model achieves 100% chemical validity, 99.87% novelty, a mean QED (quantitative estimate of drug-likeness) of 0.603 (matching the ZINC-250K training mean), and an SA (synthetic accessibility score) of 2.70, confirming compatibility with a hierarchical generative pipeline for drug-like molecule generation. Property conditioning reaches only a 12.2% QED hit rate and, under robust metrics, does not outperform an unconditioned baseline; the v2 vocabulary naturally generates drug-like molecules with QED ≈ 0.60, so stronger property guidance remains an open challenge.
Authors
- Zhizhe Lin (ORCID: https://orcid.org/0000-0002-0088-7241)
- Weihua Bai (ORCID: https://orcid.org/0000-0001-8333-7415)
- Teng Zhou (ORCID: https://orcid.org/0000-0003-1920-8891)
- Zhifeng Hao (ORCID: https://orcid.org/0000-0002-9731-1504)
- Keqin Li
- Gang Li
Institutions
- State University of New York (US)
- Zhaoqing University (CN)
- Hainan University (CN)
- Shantou University (CN)
- Department of Commerce (AU)
Publication Details
- Journal
- Journal of Chemical Information and Modeling
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1021/acs.jcim.6c01818
- Primary Topic
- Computational Drug Discovery Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00