Benchmarking Open-source Security-Tuned Language Models on Consumer Hardware

Security-tuned large language models (LLMs) are increasingly available as open-weight models for cybersecurity tasks. However, many published evaluations emphasize model capability benchmarks rather than the practical constraints faced by students, independent researchers, and small security teams using consumer hardware. This paper presents a reproducible framework for evaluating security-tuned models in resource-constrained environments and supplements the proposed framework with published benchmark results for selected models. The study focuses on models in the 7B–8B parameter range and considers cybersecurity knowledge, penetration-testing question answering, MITRE ATT&CK-oriented tasks, quantization, latency, throughput, and memory footprint. Published evidence is used only where an external source reports the corresponding measurement; no unexecuted local hardware benchmark is presented as an original experiment. The paper therefore distinguishes between secondary benchmark evidence and the proposed consumer-hardware experiment.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-29
DOI
https://doi.org/10.5281/zenodo.23043048
Primary Topic
Information and Cyber Security
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Benchmarking Open-source Security-Tuned Language Models on Consumer Hardware

Aditya Nikam
Zenodo (CERN European Organization for Nuclear Research)
Information and Cyber Security
article

Benchmarking Open-source Security-Tuned Language Models on Consumer Hardware

Aditya Nikam
article en

Abstract

Security-tuned large language models (LLMs) are increasingly available as open-weight models for cybersecurity tasks. However, many published evaluations emphasize model capability benchmarks rather than the practical constraints faced by students, independent researchers, and small security teams using consumer hardware. This paper presents a reproducible framework for evaluating security-tuned models in resource-constrained environments and supplements the proposed framework with published benchmark results for selected models. The study focuses on models in the 7B–8B parameter range and considers cybersecurity knowledge, penetration-testing question answering, MITRE ATT&CK-oriented tasks, quantization, latency, throughput, and memory footprint. Published evidence is used only where an external source reports the corresponding measurement; no unexecuted local hardware benchmark is presented as an original experiment. The paper therefore distinguishes between secondary benchmark evidence and the proposed consumer-hardware experiment.

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 4%
Information and Cyber Security
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Benchmarking Open-source Security-Tuned Language Models on Consumer Hardware — Aditya Nikam · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS