Chemical Similarity and Predicted Susceptibility in ATLAS: Structural–Functional Alignment and Divergences
This working paper (WP05 - with FAIR) documents chemical similarity and predicted susceptibility profile correspondence within the ATLAS framework. Its purpose is to describe the observed pairwise patterns and their sensitivity to analytical choices, rather than to establish biological equivalence or shared molecular mechanisms. The analysis covers 97 AACF entries, 4,656 unordered entry pairs, and predicted profiles aligned across 1,999 cell-line records. Chemical similarity is calculated using corresponding-role Morgan fingerprints and Tanimoto scores, with precursor similarities averaged using equal weights. Predicted-profile similarity is assessed using Pearson and Spearman correlations. Absolute and centred profile differences complement these measures, distinguishing co-variation from agreement in prediction levels. Registry-based quartile thresholds define four relative chemical–profile classes: high/high (HH), high/low (HL), low/high (LH), and low/low (LL). Agreement across 16 specification combinations is reported as descriptive consistency. Entry-label permutations simultaneously rearrange the rows and columns of profile similarity matrices, preserving their internal pairwise dependence. The principal observations are: No global monotonic association met the stated multiplicity-adjusted threshold under the evaluated specifications. Across-grid-consistent HH pairs showed enrichment under unrestricted all-pair permutations in both the primary and extreme-value sensitivity datasets. This enrichment did not meet the adjusted threshold under broader structural-identity exclusions or registry-prefix-restricted permutations. The primary dataset contained 158 across-grid-consistent LH pairs, without significant class enrichment. Relatively high profile correlation could coexist with differences in mean prediction levels. A targeted audit identified one prediction with substantial influence on profile dispersion, individual correlations, and class assignments. Its omission was examined separately, preserving the primary data. WP05 provides a documented comparison of structural representations, predicted-profile correlations, absolute differences, and structural identity. All observations are conditional on the current registry, selected representations, pair eligibility, and randomization models. The profiles are computational predictions, not experimentally measured responses; the analysis does not constitute external biological validation.
Authors
- Vasil Tsanov (ORCID: https://orcid.org/0000-0002-5695-1601)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23121928
- Primary Topic
- Computational Drug Discovery Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00