Homology-independent function prediction as a drop-in layer for DNA synthesis screening: calibrated k-mer and protein language model baselines with a non-invasive plugin for the Common Mechanism
DNA synthesis screening — the last technical checkpoint before a digital sequence becomes physical DNA — relies almost entirely on homology search: a sequence is flagged only if it resembles a known hazard. Engineered, diverged, or de novo designed proteins with toxin-like function can therefore pass current screens undetected. Here we present a homology-independent function-prediction layer that scores whether a sequence encodes a toxin or virulence factor directly from the protein sequence, and a plugin architecture that adds this capability to the open-source Common Mechanism (commec) screening tool without modifying it. Using only public annotated data (UniProt Tox-Prot and keyword-curated virulence factors), a cluster-disjoint evaluation protocol in which no test sequence shares more than ~30% identity with any training sequence, and screening-realistic metrics (recall at fixed low false-positive rates), we benchmark a deliberately simple amino-acid k-mer logistic-regression baseline against a frozen ESM-2 protein language model embedding classifier on a 4,003-sequence homology-held-out test set. Both models generalize without homology (AUROC 0.912 and 0.933), but they differ decisively at screening operating points: at a false-positive rate of 0.1%, the ESM-2 model detects 3.3x more held-out hazards than the k-mer baseline (recall 0.086 vs 0.026). Under controlled sequence drift, the k-mer model collapses (recall 0.33 to 0.004 by 50% identity) while the ESM-2 model retains detectable signal down to 20% identity — far beyond the reach of homology search. In an in-situ evaluation against a full installation of the Common Mechanism, 40 held-out toxins reverse-translated into fresh, never-before-seen DNA were screened: commec flagged 35 of 40 (87.5%), while the four toxins it passed entirely were scored in the 0.69–0.94 range by the ESM-2 layer — elevated but below the frozen threshold — and were flagged at query level only through short-ORF aggregation artifacts, a failure mode we characterize openly. All code is released as a drop-in companion to commec. We argue that calibrated, homology-independent scoring is a practical near-term complement to homology screening, and we release a full pipeline for training, benchmarking, and deploying such models.
Authors
- Soham Bhole
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23159128
- Primary Topic
- Machine Learning in Bioinformatics
- Type
- preprint