Cracking ERα Y537S Resistance: Explainable Machine Learning-Guided Discovery and Molecular Dynamics Validation of Stable Candidate Ligands

Background/Objectives: Endocrine resistance due to activating mutations in the estrogen receptor alpha (ERα), especially Y537S, is still a big challenge in the treatment of hormone receptor-positive breast cancer, and there is a need to develop new small-molecule inhibitors that can overcome ligand-independent receptor activation. We used an integrated machine learning and molecular dynamics (MD)-based virtual screening (VS) approach to identify potential computationally prioritized hits of ERα Y537S in this study. Methods: The dataset (~3353 compounds) was characterized using molecular fingerprints and graphically represented by PCA and t-SNE to show structurally different clusters linked to potency (pIC50). To guarantee robust downstream model training, the diagnosis of structural and influential outliers was completed through the application of rigorous data quality control methods, including the Williams plot, Mahalanobis distance, and Cook’s distance diagnostics. This was followed by multiple benchmarks of machine learning classifiers (random forest, SVM, KNN, and CNN) for activity classification, yielding the highest AUC of 0.95 for the random forest classifier. Interpretable structure–activity insights were gained by the identification of key substructures of the fingerprint by SHAP and feature importance analysis that influence the predicted potency. Three lead candidates (Hit-1, Hit-2, and Hit-3) were screened and tested with a 200 ns all-atom MD simulation in comparison with a reference control to evaluate the binding stability. Comprehensive trajectory analysis, such as PCA, FEL, RDF, salt bridges, and DCCM, was completed following the structure-based identification of the most stable ligands. Results: Three candidates were identified as Hit-1, Hit-2, and Hit-3. Hit-3 was found to be bound in the most stable binding pose and had the most similar conformational behavior to the control, whereas Hit-2 showed transient pose instability associated with increased anti-correlated domain-level motions and an alternate high-salt-bridge conformational state. As no computational metric can be an absolute measure of experimental affinity, Hit-1 has the most favorable calculated end-point binding-energy estimate, while Hit-3 showed conformational stability in MD simulations. Conclusions: The computationally prioritized candidates can be subjected to further experimental testing, as the conducted docking, MD simulations, and end-point free energy analyses cannot establish overall compound binding, cell-based activity, pharmacological mechanism, or therapeutic efficacy.

Authors

Institutions

Publication Details

Journal
Pharmaceuticals
Published
2026-09-16
DOI
https://doi.org/10.3390/ph19091471
Primary Topic
Computational Drug Discovery Methods
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Cracking ERα Y537S Resistance: Explainable Machine Learning-Guided Discovery and Molecular Dynamics Validation of Stable Candidate Ligands

Abdulmohsen M. Alruwetei
Pharmaceuticals
Computational Drug Discovery Methods
article

Cracking ERα Y537S Resistance: Explainable Machine Learning-Guided Discovery and Molecular Dynamics Validation of Stable Candidate Ligands

Abdulmohsen M. Alruwetei
article en

Abstract

Background/Objectives: Endocrine resistance due to activating mutations in the estrogen receptor alpha (ERα), especially Y537S, is still a big challenge in the treatment of hormone receptor-positive breast cancer, and there is a need to develop new small-molecule inhibitors that can overcome ligand-independent receptor activation. We used an integrated machine learning and molecular dynamics (MD)-based virtual screening (VS) approach to identify potential computationally prioritized hits of ERα Y537S in this study. Methods: The dataset (~3353 compounds) was characterized using molecular fingerprints and graphically represented by PCA and t-SNE to show structurally different clusters linked to potency (pIC50). To guarantee robust downstream model training, the diagnosis of structural and influential outliers was completed through the application of rigorous data quality control methods, including the Williams plot, Mahalanobis distance, and Cook’s distance diagnostics. This was followed by multiple benchmarks of machine learning classifiers (random forest, SVM, KNN, and CNN) for activity classification, yielding the highest AUC of 0.95 for the random forest classifier. Interpretable structure–activity insights were gained by the identification of key substructures of the fingerprint by SHAP and feature importance analysis that influence the predicted potency. Three lead candidates (Hit-1, Hit-2, and Hit-3) were screened and tested with a 200 ns all-atom MD simulation in comparison with a reference control to evaluate the binding stability. Comprehensive trajectory analysis, such as PCA, FEL, RDF, salt bridges, and DCCM, was completed following the structure-based identification of the most stable ligands. Results: Three candidates were identified as Hit-1, Hit-2, and Hit-3. Hit-3 was found to be bound in the most stable binding pose and had the most similar conformational behavior to the control, whereas Hit-2 showed transient pose instability associated with increased anti-correlated domain-level motions and an alternate high-salt-bridge conformational state. As no computational metric can be an absolute measure of experimental affinity, Hit-1 has the most favorable calculated end-point binding-energy estimate, while Hit-3 showed conformational stability in MD simulations. Conclusions: The computationally prioritized candidates can be subjected to further experimental testing, as the conducted docking, MD simulations, and end-point free energy analyses cannot establish overall compound binding, cell-based activity, pharmacological mechanism, or therapeutic efficacy.

PharmaceuticalsVol. 19(9)
Qassim University (SA), Buraydah Colleges (SA)
Openalex Percentile: Top 8%
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.