Homology-independent function prediction as a drop-in layer for DNA synthesis screening: calibrated k-mer and protein language model baselines with a non-invasive plugin for the Common Mechanism

DNA synthesis screening — the last technical checkpoint before a digital sequence becomes physical DNA — relies almost entirely on homology search: a sequence is flagged only if it resembles a known hazard. Engineered, diverged, or de novo designed proteins with toxin-like function can therefore pass current screens undetected. Here we present a homology-independent function-prediction layer that scores whether a sequence encodes a toxin or virulence factor directly from the protein sequence, and a plugin architecture that adds this capability to the open-source Common Mechanism (commec) screening tool without modifying it. Using only public annotated data (UniProt Tox-Prot and keyword-curated virulence factors), a cluster-disjoint evaluation protocol in which no test sequence shares more than ~30% identity with any training sequence, and screening-realistic metrics (recall at fixed low false-positive rates), we benchmark a deliberately simple amino-acid k-mer logistic-regression baseline against a frozen ESM-2 protein language model embedding classifier on a 4,003-sequence homology-held-out test set. Both models generalize without homology (AUROC 0.912 and 0.933), but they differ decisively at screening operating points: at a false-positive rate of 0.1%, the ESM-2 model detects 3.3x more held-out hazards than the k-mer baseline (recall 0.086 vs 0.026). Under controlled sequence drift, the k-mer model collapses (recall 0.33 to 0.004 by 50% identity) while the ESM-2 model retains detectable signal down to 20% identity — far beyond the reach of homology search. In an in-situ evaluation against a full installation of the Common Mechanism, 40 held-out toxins reverse-translated into fresh, never-before-seen DNA were screened: commec flagged 35 of 40 (87.5%), while the four toxins it passed entirely were scored in the 0.69–0.94 range by the ESM-2 layer — elevated but below the frozen threshold — and were flagged at query level only through short-ORF aggregation artifacts, a failure mode we characterize openly. All code is released as a drop-in companion to commec. We argue that calibrated, homology-independent scoring is a practical near-term complement to homology screening, and we release a full pipeline for training, benchmarking, and deploying such models.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.22999874
Primary Topic
Machine Learning in Bioinformatics
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Homology-independent function prediction as a drop-in layer for DNA synthesis screening: calibrated k-mer and protein language model baselines with a non-invasive plugin for the Common Mechanism

Soham Bhole
Zenodo (CERN European Organization for Nuclear Research)
Machine Learning in Bioinformatics
preprint

Homology-independent function prediction as a drop-in layer for DNA synthesis screening: calibrated k-mer and protein language model baselines with a non-invasive plugin for the Common Mechanism

Soham Bhole
preprint en

Abstract

DNA synthesis screening — the last technical checkpoint before a digital sequence becomes physical DNA — relies almost entirely on homology search: a sequence is flagged only if it resembles a known hazard. Engineered, diverged, or de novo designed proteins with toxin-like function can therefore pass current screens undetected. Here we present a homology-independent function-prediction layer that scores whether a sequence encodes a toxin or virulence factor directly from the protein sequence, and a plugin architecture that adds this capability to the open-source Common Mechanism (commec) screening tool without modifying it. Using only public annotated data (UniProt Tox-Prot and keyword-curated virulence factors), a cluster-disjoint evaluation protocol in which no test sequence shares more than ~30% identity with any training sequence, and screening-realistic metrics (recall at fixed low false-positive rates), we benchmark a deliberately simple amino-acid k-mer logistic-regression baseline against a frozen ESM-2 protein language model embedding classifier on a 4,003-sequence homology-held-out test set. Both models generalize without homology (AUROC 0.912 and 0.933), but they differ decisively at screening operating points: at a false-positive rate of 0.1%, the ESM-2 model detects 3.3x more held-out hazards than the k-mer baseline (recall 0.086 vs 0.026). Under controlled sequence drift, the k-mer model collapses (recall 0.33 to 0.004 by 50% identity) while the ESM-2 model retains detectable signal down to 20% identity — far beyond the reach of homology search. In an in-situ evaluation against a full installation of the Common Mechanism, 40 held-out toxins reverse-translated into fresh, never-before-seen DNA were screened: commec flagged 35 of 40 (87.5%), while the four toxins it passed entirely were scored in the 0.69–0.94 range by the ESM-2 layer — elevated but below the frozen threshold — and were flagged at query level only through short-ORF aggregation artifacts, a failure mode we characterize openly. All code is released as a drop-in companion to commec. We argue that calibrated, homology-independent scoring is a practical near-term complement to homology screening, and we release a full pipeline for training, benchmarking, and deploying such models.

Zenodo (CERN European Organization for Nuclear Research)
Machine Learning in Bioinformatics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Homology-independent function prediction as a drop-in layer for DNA synthesis screening: calibrated k-mer and protein language model baselines with a non-invasive plugin for the Common Mechanism — Soham Bhole · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS