Assay-Level Modeling and External Validation of Antibody Developability Using Physicochemical Descriptors and Sequence Embeddings
Abstract Antibody developability is inherently multidimensional, involving thermal stability, aggregation and hydrophobicity, nonspecific interactions, and expression-related manufacturability. Instead of treating developability as a single composite label, we investigated assay-level modeling using GDPa1 as a benchmark data set. We developed a two-branch framework combining interpretable developability descriptors and data-driven calibration. One branch fits assay-specific descriptor models to derive end point-specific physicochemical weights, while the other integrates frozen pretrained VH/VL sequence embeddings with computed descriptors in a multitask regression architecture. Internal validation across eight GDPa1 assays showed assay-dependent predictability, with hydrophobicity- and aggregation-related readouts generally showing stronger signals than expression-related and some thermal stability end points. Descriptor coefficient analysis revealed distinct physicochemical patterns across assays, supporting the view that antibody developability cannot be represented by a single universal descriptor weighting scheme. External held-out validation on GDPa3 revealed limited cross-data set transferability. Across seven directly matched assays, mean Spearman correlations were 0.126 for the full descriptor model, 0.139 for the sequence-only descriptor model, and 0.125 for the hybrid sequence–descriptor model. HIC retention time was the most transferable end point, whereas SEC monomer percentage and several thermal stability end points showed weak external correlations. The hybrid model did not consistently outperform descriptor-only baselines, indicating that pretrained sequence embeddings provided limited and assay-dependent incremental value in this small-data setting. GDPa1 sequence-cluster analysis showed that internal validation estimates depended on clustering threshold and fold construction, whereas GDPa3 provided a more direct external assessment of cross-data set transferability. These results support a cautious assay-level framework for antibody developability modeling. Descriptor-based representations remain strong and interpretable baselines, while sequence embeddings require end point-specific and external validation. The proposed workflow can convert assay-level predictions into domain-level profiles for computational triage, but prospective candidate prioritization requires binding assessment and experimental validation.
Authors
- Hei Wun Kan (ORCID: https://orcid.org/0000-0003-4566-5746)
- Haiping Zhang (ORCID: https://orcid.org/0000-0003-2133-1768)
- Mark Akinola Ige
- Zhiyuan Zheng
- John Z. H. Zhang
- Xiaochun Wan
- Shasha Jiang
- Junxin Li
Institutions
- University of Lagos (NG)
- Southern University of Science and Technology (CN)
- Shenzhen Institutes of Advanced Technology (CN)
- Shenzhen Technology University (CN)
- University of Chinese Academy of Sciences (CN)
Publication Details
- Journal
- Journal of Chemical Information and Modeling
- Published
- 2026-09-14
- DOI
- https://doi.org/10.1021/acs.jcim.6c02385
- Primary Topic
- Protein purification and stability
- Type
- article
- Field-Weighted Citation Impact
- 0.00