Benchmarking Open-source Security-Tuned Language Models on Consumer Hardware
Security-tuned large language models (LLMs) are increasingly available as open-weight models for cybersecurity tasks. However, many published evaluations emphasize model capability benchmarks rather than the practical constraints faced by students, independent researchers, and small security teams using consumer hardware. This paper presents a reproducible framework for evaluating security-tuned models in resource-constrained environments and supplements the proposed framework with published benchmark results for selected models. The study focuses on models in the 7B–8B parameter range and considers cybersecurity knowledge, penetration-testing question answering, MITRE ATT&CK-oriented tasks, quantization, latency, throughput, and memory footprint. Published evidence is used only where an external source reports the corresponding measurement; no unexecuted local hardware benchmark is presented as an original experiment. The paper therefore distinguishes between secondary benchmark evidence and the proposed consumer-hardware experiment.
Authors
- Aditya Nikam
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-29
- DOI
- https://doi.org/10.5281/zenodo.23043048
- Primary Topic
- Information and Cyber Security
- Type
- article
- Field-Weighted Citation Impact
- 0.00