Toward Reproducible Local Evaluation of Prompt Injection in Open-Weight LLMs: Design and Threat Model

Open-weight language models are increasingly deployed in local and edge environments, yet their robustnessagainst prompt injection and system-prompt leakage remains poorly quantified under reproducible conditions.Existing large-scale benchmarks largely target hosted, API-served models and can be difficult to run end-to-endon consumer hardware, leaving practitioners who deploy or fine-tune models locally with few reproduciblebaselines. We present the design of a fully local, deterministic evaluation pipeline intended to measure thesevulnerabilities across multiple open-weight models without relying on proprietary APIs. The design specifiesfour modular components (model loader, attack generator, harness, and scorer), a taxonomy of four attackcategories—direct injection, multi-turn progressive injection, system-prompt leakage probes, and simulatedindirect injection—a fixed threat model, and a statistically grounded protocol for a 160-instance attack suite(40 instances per category) intended to be run against six representative instruction-tuned open-weight modelsspanning approximately 3–9B parameters. The paper’s contribution is the design, threat model, and taxonomy;implementing the pipeline, running the evaluation, and releasing the resulting code and transcripts are plannedfuture work. We discuss the statistical precision such a suite would provide, the ethical safeguards intended forthe eventual evaluation, and how the design relates to prior benchmarks aimed primarily at hosted models.Keywords: prompt injection, system-prompt leakage, open-weight models, reproducible evaluation, LLM security,red-teaming, threat modeling

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-18
DOI
https://doi.org/10.5281/zenodo.22819018
Primary Topic
Security and Verification in Computing
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Toward Reproducible Local Evaluation of Prompt Injection in Open-Weight LLMs: Design and Threat Model

MD Tanjim Molla
Zenodo (CERN European Organization for Nuclear Research)
Security and Verification in Computing
preprint

Toward Reproducible Local Evaluation of Prompt Injection in Open-Weight LLMs: Design and Threat Model

MD Tanjim Molla
preprint en

Abstract

Open-weight language models are increasingly deployed in local and edge environments, yet their robustnessagainst prompt injection and system-prompt leakage remains poorly quantified under reproducible conditions.Existing large-scale benchmarks largely target hosted, API-served models and can be difficult to run end-to-endon consumer hardware, leaving practitioners who deploy or fine-tune models locally with few reproduciblebaselines. We present the design of a fully local, deterministic evaluation pipeline intended to measure thesevulnerabilities across multiple open-weight models without relying on proprietary APIs. The design specifiesfour modular components (model loader, attack generator, harness, and scorer), a taxonomy of four attackcategories—direct injection, multi-turn progressive injection, system-prompt leakage probes, and simulatedindirect injection—a fixed threat model, and a statistically grounded protocol for a 160-instance attack suite(40 instances per category) intended to be run against six representative instruction-tuned open-weight modelsspanning approximately 3–9B parameters. The paper’s contribution is the design, threat model, and taxonomy;implementing the pipeline, running the evaluation, and releasing the resulting code and transcripts are plannedfuture work. We discuss the statistical precision such a suite would provide, the ethical safeguards intended forthe eventual evaluation, and how the design relates to prior benchmarks aimed primarily at hosted models.Keywords: prompt injection, system-prompt leakage, open-weight models, reproducible evaluation, LLM security,red-teaming, threat modeling

Zenodo (CERN European Organization for Nuclear Research)
Daffodil International University (BD)
Security and Verification in Computing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.