Predicting non-specific binding of VHHs using machine learning models with cluster-aware validation

Propensity for nonspecific binding—also known as polyreactivity—is a serious developability risk factor for biotherapeutic candidates. To minimize this risk, drug companies are increasingly relying on in silico tools utilizing machine learning methods, but developing these tools is challenging. For example, the available data often contains many closely related sequences originating from drug pipeline projects, which can introduce significant biases in the in silico models training and benchmarking, leading to poor generalizability on new data. We present here a workflow designed to diagnose and mitigate some of the problems associated with using pipeline data. The workflow is based on a custom cross-validation procedure that can evaluate model performance on unseen data in different contexts. As a demonstration of the workflow, we use it to train a model to predict variable heavy-chain only fragment antibodies (VHH) binding to baculovirus particles (BVP)—a widely used assay for nonspecific binding. Using descriptors based on computed protein structures, the workflow identifies several risk factors that correlate with higher polyreactivity levels.

Authors

Institutions

Publication Details

Journal
mAbs
Published
2026-09-16
DOI
https://doi.org/10.1080/19420862.2026.2732787
Primary Topic
Machine Learning in Bioinformatics
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Predicting non-specific binding of VHHs using machine learning models with cluster-aware validation

Rebecca Croasdale-Wood, Sharfa Farzandh, Valentin Stanev, Tony Pham et al.
mAbs
Machine Learning in Bioinformatics
article

Predicting non-specific binding of VHHs using machine learning models with cluster-aware validation

Rebecca Croasdale-Wood, Sharfa Farzandh, Valentin Stanev, Tony Pham, Maryam Pouryahya, Jenna Caldwell, Gilad Kaplan, Kuan-Lin Chen, Tom Diethe, Federico Devalle, Andrew Dippel, Mehdi Boroumand, Chacko Chakiath, Isabelle Sermadiras, Bismark Amofah, Rohan Jain, Jen DiChiara, Mark Hutchinson, Jay Hyun Jo
article en

Abstract

Propensity for nonspecific binding—also known as polyreactivity—is a serious developability risk factor for biotherapeutic candidates. To minimize this risk, drug companies are increasingly relying on in silico tools utilizing machine learning methods, but developing these tools is challenging. For example, the available data often contains many closely related sequences originating from drug pipeline projects, which can introduce significant biases in the in silico models training and benchmarking, leading to poor generalizability on new data. We present here a workflow designed to diagnose and mitigate some of the problems associated with using pipeline data. The workflow is based on a custom cross-validation procedure that can evaluate model performance on unseen data in different contexts. As a demonstration of the workflow, we use it to train a model to predict variable heavy-chain only fragment antibodies (VHH) binding to baculovirus particles (BVP)—a widely used assay for nonspecific binding. Using descriptors based on computed protein structures, the workflow identifies several risk factors that correlate with higher polyreactivity levels.

mAbsVol. 18(1)
AstraZeneca (Brazil) (BR)
AstraZeneca
Openalex Percentile: Top 18%
Machine Learning in Bioinformatics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.