MSGRL: A Motif-Driven Self-Supervised Graph Representation Learning Framework for Interpretable Molecular Property Prediction

Molecular property prediction is a fundamental task in drug discovery and chemical biology, where effective molecular representations are essential for accurate prediction. Learning transferable motif-level representations remains challenging because explicit motif annotations are scarce and existing representations may be altered during downstream supervised optimization. In this study, we propose MSGRL, a motif-driven self-supervised graph representation learning framework for interpretable molecular property prediction. MSGRL represents each molecule through a hierarchical graph structure, consisting of a motif-based graph for inter-motif organization and motif-specific atom-based graphs for intra-motif atomic structure. Its central design is to decouple label-agnostic intra-motif structural learning from label-dependent inter-motif property learning. An MPNN-GRU encoder is pretrained on motif-specific atom-based graphs using a variational motif graph autoencoder (VMGAE), which reconstructs the internal bond topology of motifs in a self-supervised manner. After pretraining, the intra-motif encoder is kept frozen, while the downstream module pools atom-level feature matrices into motif vectors, propagates information over the motif-based graph, and applies attention-based pooling for molecular property prediction. This design keeps the pretrained intra-motif representations fixed while allowing the inter-motif network and prediction head to adapt to individual downstream tasks. Experiments on eight MoleculeNet benchmark datasets show that MSGRL achieves the highest ROC-AUC scores on all five classification datasets and the lowest RMSE on Lipophilicity, while its performance on ESOL and FreeSolv is more mixed. Ablation studies further support the contributions of encoder freezing, motif-based graph construction, and attention-based pooling. Motif-level attribution analyses provide qualitative and dataset-level evidence regarding the substructures emphasized by the model. These results demonstrate the effectiveness of the proposed hierarchical representation strategy, particularly for the evaluated molecular classification tasks.

Authors

Institutions

Publication Details

Journal
Molecules
Published
2026-08-27
DOI
https://doi.org/10.3390/molecules31173008
Primary Topic
Computational Drug Discovery Methods
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

MSGRL: A Motif-Driven Self-Supervised Graph Representation Learning Framework for Interpretable Molecular Property Prediction

Guosheng Zhu, You Wu, Wen Wang, Haitao Fu et al.
Molecules
Computational Drug Discovery Methods
article

MSGRL: A Motif-Driven Self-Supervised Graph Representation Learning Framework for Interpretable Molecular Property Prediction

Guosheng Zhu, You Wu, Wen Wang, Haitao Fu, Cheng Zeng, Xiaoyun Qi, Qiyu Tang, Yuxin Jiang
article en

Abstract

Molecular property prediction is a fundamental task in drug discovery and chemical biology, where effective molecular representations are essential for accurate prediction. Learning transferable motif-level representations remains challenging because explicit motif annotations are scarce and existing representations may be altered during downstream supervised optimization. In this study, we propose MSGRL, a motif-driven self-supervised graph representation learning framework for interpretable molecular property prediction. MSGRL represents each molecule through a hierarchical graph structure, consisting of a motif-based graph for inter-motif organization and motif-specific atom-based graphs for intra-motif atomic structure. Its central design is to decouple label-agnostic intra-motif structural learning from label-dependent inter-motif property learning. An MPNN-GRU encoder is pretrained on motif-specific atom-based graphs using a variational motif graph autoencoder (VMGAE), which reconstructs the internal bond topology of motifs in a self-supervised manner. After pretraining, the intra-motif encoder is kept frozen, while the downstream module pools atom-level feature matrices into motif vectors, propagates information over the motif-based graph, and applies attention-based pooling for molecular property prediction. This design keeps the pretrained intra-motif representations fixed while allowing the inter-motif network and prediction head to adapt to individual downstream tasks. Experiments on eight MoleculeNet benchmark datasets show that MSGRL achieves the highest ROC-AUC scores on all five classification datasets and the lowest RMSE on Lipophilicity, while its performance on ESOL and FreeSolv is more mixed. Ablation studies further support the contributions of encoder freezing, motif-based graph construction, and attention-based pooling. Motif-level attribution analyses provide qualitative and dataset-level evidence regarding the substructures emphasized by the model. These results demonstrate the effectiveness of the proposed hierarchical representation strategy, particularly for the evaluated molecular classification tasks.

MoleculesVol. 31(17)
Roslin Institute (GB), Huazhong Agricultural University (CN), Hubei University (CN), University of Edinburgh (GB)
Openalex Percentile: Top 8%
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.