Computational Drug Effects Modelling
These files represent R scripts used to perform computation modelling of drug effects across human transcriptomic profiles previously described by Jarr et al., 2025, in Circulation scientific journal. More detailed information can be found in the public Github repository. Method summary This R script implements the SHP-1i in-silico treatment predictive model using a Random Forest (RF) multi-class classification approach (see Figure 1). To develop this model, we first identified 121 differentially expressed genes (DEGs, FDR < 0.1) in human macrophages following in-vitro SHP-1i treatment. During model construction, these genes were then narrowed down to 119 through LASSO feature selection. Using the original, unaltered expression data for these 119 genes from 654 human plaque samples, we trained a random forest classification model to predict plaque subtypes. The data was partitioned into 1,000 randomly generated subsets, each split into 80% training and 20% testing sets. RF model performance was evaluated using multiclass area under the ROC curve (AUC), with an average AUC of 75% or higher considered indicative of a well performing model. To simulate in-silico treatment of SHP-1i, the expression changes of the same 119 DESeq2-derived, LASSO-selected DEGs were projected onto the 654 human plaque samples. This was achieved by multiplying the expression values by the antilogarithmic transformed fold changes, resulting in a “perturbed” dataset (with patients as rows and genes as columns). The trained predictive model was then applied to classify the perturbed dataset, and these classifications were used to quantify the effects of the in-silico SHP-1i treatment. After 1,000 iterations, an averaged output was generated, revealing predicted changes in plaque classification after SHP-1i in-silico treatment. The shifts in plaque subtypes following this projection were assessed for statistical significance using McNemar’s test, which evaluates paired data to determine whether the observed shifts (e.g., from more vulnerable to less vulnerable plaques) are meaningful. To further the interpret results, odds ratios (ORs) were calculated by dividing the number of plaques predicted to shift favorably (e.g., transitioning from more vulnerable to less vulnerable subtypes) by the number of plaques predicted to shift unfavorably (e.g., transitioning from less vulnerable to more vulnerable subtypes). Corresponding confidence intervals were calculated using standard formulas and p-values were derived from McNemar's Chi-squared test. Libraries used: progress 1.2.3, pROC 1.18.5, caret 6.0-94, glmnet 4.1-8, randomForest 4.7-1.1, EnsDb.Hsapiens.v86 2.99.0, biomaRT 2.60.1, ggalluvial 0.12.5, limma 3.60.4, ggplot2 3.5.1 dplyr 1.1.4
Authors
- Michal Mokry (ORCID: https://orcid.org/0000-0002-7719-6668)
- Kaylin C.A. Palm (ORCID: https://orcid.org/0000-0002-0816-8962)
Institutions
- University Medical Center Utrecht (NL)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23062149
- Primary Topic
- Computational Drug Discovery Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00