Assay-Level Modeling and External Validation of Antibody Developability Using Physicochemical Descriptors and Sequence Embeddings

Abstract Antibody developability is inherently multidimensional, involving thermal stability, aggregation and hydrophobicity, nonspecific interactions, and expression-related manufacturability. Instead of treating developability as a single composite label, we investigated assay-level modeling using GDPa1 as a benchmark data set. We developed a two-branch framework combining interpretable developability descriptors and data-driven calibration. One branch fits assay-specific descriptor models to derive end point-specific physicochemical weights, while the other integrates frozen pretrained VH/VL sequence embeddings with computed descriptors in a multitask regression architecture. Internal validation across eight GDPa1 assays showed assay-dependent predictability, with hydrophobicity- and aggregation-related readouts generally showing stronger signals than expression-related and some thermal stability end points. Descriptor coefficient analysis revealed distinct physicochemical patterns across assays, supporting the view that antibody developability cannot be represented by a single universal descriptor weighting scheme. External held-out validation on GDPa3 revealed limited cross-data set transferability. Across seven directly matched assays, mean Spearman correlations were 0.126 for the full descriptor model, 0.139 for the sequence-only descriptor model, and 0.125 for the hybrid sequence–descriptor model. HIC retention time was the most transferable end point, whereas SEC monomer percentage and several thermal stability end points showed weak external correlations. The hybrid model did not consistently outperform descriptor-only baselines, indicating that pretrained sequence embeddings provided limited and assay-dependent incremental value in this small-data setting. GDPa1 sequence-cluster analysis showed that internal validation estimates depended on clustering threshold and fold construction, whereas GDPa3 provided a more direct external assessment of cross-data set transferability. These results support a cautious assay-level framework for antibody developability modeling. Descriptor-based representations remain strong and interpretable baselines, while sequence embeddings require end point-specific and external validation. The proposed workflow can convert assay-level predictions into domain-level profiles for computational triage, but prospective candidate prioritization requires binding assessment and experimental validation.

Authors

Institutions

Publication Details

Journal
Journal of Chemical Information and Modeling
Published
2026-09-14
DOI
https://doi.org/10.1021/acs.jcim.6c02385
Primary Topic
Protein purification and stability
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Assay-Level Modeling and External Validation of Antibody Developability Using Physicochemical Descriptors and Sequence Embeddings

Hei Wun Kan, Haiping Zhang, Mark Akinola Ige, Zhiyuan Zheng et al.
Journal of Chemical Information and Modeling
Protein purification and stability
article

Assay-Level Modeling and External Validation of Antibody Developability Using Physicochemical Descriptors and Sequence Embeddings

Hei Wun Kan, Haiping Zhang, Mark Akinola Ige, Zhiyuan Zheng, John Z. H. Zhang, Xiaochun Wan, Shasha Jiang, Junxin Li
article en

Abstract

Abstract Antibody developability is inherently multidimensional, involving thermal stability, aggregation and hydrophobicity, nonspecific interactions, and expression-related manufacturability. Instead of treating developability as a single composite label, we investigated assay-level modeling using GDPa1 as a benchmark data set. We developed a two-branch framework combining interpretable developability descriptors and data-driven calibration. One branch fits assay-specific descriptor models to derive end point-specific physicochemical weights, while the other integrates frozen pretrained VH/VL sequence embeddings with computed descriptors in a multitask regression architecture. Internal validation across eight GDPa1 assays showed assay-dependent predictability, with hydrophobicity- and aggregation-related readouts generally showing stronger signals than expression-related and some thermal stability end points. Descriptor coefficient analysis revealed distinct physicochemical patterns across assays, supporting the view that antibody developability cannot be represented by a single universal descriptor weighting scheme. External held-out validation on GDPa3 revealed limited cross-data set transferability. Across seven directly matched assays, mean Spearman correlations were 0.126 for the full descriptor model, 0.139 for the sequence-only descriptor model, and 0.125 for the hybrid sequence–descriptor model. HIC retention time was the most transferable end point, whereas SEC monomer percentage and several thermal stability end points showed weak external correlations. The hybrid model did not consistently outperform descriptor-only baselines, indicating that pretrained sequence embeddings provided limited and assay-dependent incremental value in this small-data setting. GDPa1 sequence-cluster analysis showed that internal validation estimates depended on clustering threshold and fold construction, whereas GDPa3 provided a more direct external assessment of cross-data set transferability. These results support a cautious assay-level framework for antibody developability modeling. Descriptor-based representations remain strong and interpretable baselines, while sequence embeddings require end point-specific and external validation. The proposed workflow can convert assay-level predictions into domain-level profiles for computational triage, but prospective candidate prioritization requires binding assessment and experimental validation.

Journal of Chemical Information and Modeling
University of Lagos (NG), Southern University of Science and Technology (CN), Shenzhen Institutes of Advanced Technology (CN), Shenzhen Technology University (CN), University of Chinese Academy of Sciences (CN)
Decent work and economic growth
Openalex Percentile: Top 18%
Protein purification and stability
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.