Multi-feature classification to improve colorimetric loop-mediated isothermal amplification fidelity

Loop-mediated isothermal amplification (LAMP) is a cost-effective and portable assay technique for performing nucleic acid-based diagnostics in the field whose adoption is hindered by design and reproducibility issues. This is due to a complex primer design process that fine-tunes parameters across 6–8 binding regions. The likelihood of assay success depends on satisfying thermodynamic and secondary structure constraints while maintaining target specificity and avoiding overlaps between multiple primers. Software such as the NEB ® LAMP Primer Design Tool, PREMIER Biosoft LAMP Designer, Primer3, PCR Signature Erosion Tool (PSET), and PrimerExplorer enable automation of this task for researchers. However, in our experience, these programs can sometimes yield inconsistent results in laboratory testing. Our approach trained multiple machine learning (ML) models on primer sets targeting various organisms from working assays and failing ones to determine significant features and improve predictions prior to ordering primer sets. A literature review produced an initial list of primer sets ( n = 116), which were then filtered based on reference template availability to discern their FIP/BIP components (F2/F1c and B1c/B2). Additional filtering removed assays with any shared primers to ensure sample independence during cross-validation. The final training set ( n = 101) included sequence and thermodynamic features derived from primers collected from the review ( n = 71) and those designed in-house with PSET ( n = 30). To remove dependence on arbitrary forward and backward labels, the minimum and maximum were computed for each feature across primer pairs. Failing assays were difficult to obtain from the publications, so we provided our own ( n = 22). Using WEKA Experimenter, models were created based on decision tree and Bayesian learning algorithms using an experimental scheme that performed a parameter grid search, seeded replicates, feature selection, and cross-validation while avoiding data-leakage and outputting logs for model comparison, feature analysis, and overfit assessment. Thermodynamic features associated with the inner loop primers (F1c/B1c) consistently appeared in the top ranks according to consensus between information gain, class-correlation, and model-based feature ranking. For classification, all non-baseline models were statistically equivalent. The highest-performing model, J48, achieved TP and TN rates of 0.92 (± 0.02) and 0.80 (± 0.08) while achieving Cohen’s kappa coefficient and F-score values of 0.70 (± 0.07) and 0.93 (± 0.01), respectively. This work highlights how a practical model was built from a small, imbalanced training set incorporating negative research results, of which more are needed to improve generalization and refine parameters critical to assay success.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-15
DOI
https://doi.org/10.1038/s41598-026-70510-9
Primary Topic
Biosensors and Analytical Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Multi-feature classification to improve colorimetric loop-mediated isothermal amplification fidelity

Shanmuga Sozhamannan, Nicholas Tolli, Bradley W. Abramson, Daniel Negrón et al.
Scientific Reports
Biosensors and Analytical Detection
article

Multi-feature classification to improve colorimetric loop-mediated isothermal amplification fidelity

Shanmuga Sozhamannan, Nicholas Tolli, Bradley W. Abramson, Daniel Negrón, Gabrielle Melton, Bryan D. Necciai, Katharine Jennings, Sveta Jagannathan, Kelsey Hauser
article en

Abstract

Loop-mediated isothermal amplification (LAMP) is a cost-effective and portable assay technique for performing nucleic acid-based diagnostics in the field whose adoption is hindered by design and reproducibility issues. This is due to a complex primer design process that fine-tunes parameters across 6–8 binding regions. The likelihood of assay success depends on satisfying thermodynamic and secondary structure constraints while maintaining target specificity and avoiding overlaps between multiple primers. Software such as the NEB ® LAMP Primer Design Tool, PREMIER Biosoft LAMP Designer, Primer3, PCR Signature Erosion Tool (PSET), and PrimerExplorer enable automation of this task for researchers. However, in our experience, these programs can sometimes yield inconsistent results in laboratory testing. Our approach trained multiple machine learning (ML) models on primer sets targeting various organisms from working assays and failing ones to determine significant features and improve predictions prior to ordering primer sets. A literature review produced an initial list of primer sets ( n = 116), which were then filtered based on reference template availability to discern their FIP/BIP components (F2/F1c and B1c/B2). Additional filtering removed assays with any shared primers to ensure sample independence during cross-validation. The final training set ( n = 101) included sequence and thermodynamic features derived from primers collected from the review ( n = 71) and those designed in-house with PSET ( n = 30). To remove dependence on arbitrary forward and backward labels, the minimum and maximum were computed for each feature across primer pairs. Failing assays were difficult to obtain from the publications, so we provided our own ( n = 22). Using WEKA Experimenter, models were created based on decision tree and Bayesian learning algorithms using an experimental scheme that performed a parameter grid search, seeded replicates, feature selection, and cross-validation while avoiding data-leakage and outputting logs for model comparison, feature analysis, and overfit assessment. Thermodynamic features associated with the inner loop primers (F1c/B1c) consistently appeared in the top ranks according to consensus between information gain, class-correlation, and model-based feature ranking. For classification, all non-baseline models were statistically equivalent. The highest-performing model, J48, achieved TP and TN rates of 0.92 (± 0.02) and 0.80 (± 0.08) while achieving Cohen’s kappa coefficient and F-score values of 0.70 (± 0.07) and 0.93 (± 0.01), respectively. This work highlights how a practical model was built from a small, imbalanced training set incorporating negative research results, of which more are needed to improve generalization and refine parameters critical to assay success.

Scientific Reports
Noblis (US), Stafford College (GB), Chemical Dynamics (United States) (US)
Openalex Percentile: Top 20%
Biosensors and Analytical Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.