T-ANALYZE as a Workflow for Matched Molecule Pair Analysis on Peptides
Abstract Matched molecular pairs (MMP) is an approach to finding local changes in structure that give rise to consistent changes in activity. Assuming molecules A and B are in a matched pair, one can imagine a “transformation”, i.e., adding, removing, or exchanging a small set of atoms that changes A to B. Typical machine-learning algorithms are meant to handle individual molecules A, and model the activity difference between A and the average activity, and are agnostic to the presence of similar compounds in the training set. In contrast, MMP looks for differences between similar compounds A and B and is agnostic to the absolute value of the activities of A and B. Machine learning models tend to be poor at predicting activity cliffs, while MMP is designed to identify them. Previously, we (Sheridan et al., J. Chem. Inf. Model. 2006, 46, 180–192) suggested a workflow T-ANALYZE that organizes and displays sets of closely related drug-sized compounds such that any consistent MMP-based structure-activity in a data set can be easily seen. Nowadays, peptides are under consideration as candidate drugs and the MMP approach should be extended to them. Peptides are small enough that they can be treated as connection tables, i.e., atoms connected with bonds. Alternatively, they can be written as sequences, i.e., a one-dimensional list of monomer names. In this paper, we adapt the T-ANALYZE workflow to peptides represented both as connection tables and as sequences. We demonstrate this on the PAMPA activity of the CycPeptMPDB dataset.
Authors
- Robert P. Sheridan (ORCID: https://orcid.org/0000-0002-6549-1635)
Institutions
- Merck & Co., Inc., Rahway, NJ, USA (United States) (US)
Publication Details
- Journal
- Journal of Chemical Information and Modeling
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1021/acs.jcim.6c01414
- Primary Topic
- Computational Drug Discovery Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00