Open or Frontier? A Cost- and Energy-Aware Benchmark of Large Language Models for Software Vulnerability Detection

Large language models (LLMs) are increasingly applied to software vulnerability detection, but evaluations report accuracy while ignoring inference cost and energy and under-represent open-weight models relative to proprietary systems. We benchmark eight LLMs, three frontier, and five open-weight on a stratified 1549-function subset of label-clean PrimeVul, treating cost and energy as first-class axes alongside detection quality. Cost is measured directly; energy is measured on-GPU across a concurrency sweep for three locally servable open models and FLOP-estimated with a sensitivity range for the API-served ones. Efficiency is the robust finding: open-weight models occupy the quality-efficiency Pareto frontier in every configuration tested, and no frontier model is Pareto-optimal; this is a 4-billion-parameter model matching the best frontier system’s quality at one sixty-sixth of the list price. On quality, at a matched output budget, the best open model significantly exceeds every frontier model (0.711 balanced accuracy against 0.606–0.653), although the strongest frontier system is level with the next two open models. Two findings bound the practical reading. A 125-million-parameter detector fine-tuned on PrimeVul outperforms all eight LLMs (0.765), so where in-distribution labels exist, a small task-specific model is the better instrument. It should be noted that nothing here is deployment-ready: at the natural 1:44 prevalence, precision is 2.3–8.2%.

Authors

Institutions

Publication Details

Journal
Computers
Published
2026-09-16
DOI
https://doi.org/10.3390/computers15090623
Primary Topic
Software Engineering Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Open or Frontier? A Cost- and Energy-Aware Benchmark of Large Language Models for Software Vulnerability Detection

Wolfgang Slany, Patrick Deininger
Computers
Software Engineering Research
article

Open or Frontier? A Cost- and Energy-Aware Benchmark of Large Language Models for Software Vulnerability Detection

Wolfgang Slany, Patrick Deininger
article en

Abstract

Large language models (LLMs) are increasingly applied to software vulnerability detection, but evaluations report accuracy while ignoring inference cost and energy and under-represent open-weight models relative to proprietary systems. We benchmark eight LLMs, three frontier, and five open-weight on a stratified 1549-function subset of label-clean PrimeVul, treating cost and energy as first-class axes alongside detection quality. Cost is measured directly; energy is measured on-GPU across a concurrency sweep for three locally servable open models and FLOP-estimated with a sensitivity range for the API-served ones. Efficiency is the robust finding: open-weight models occupy the quality-efficiency Pareto frontier in every configuration tested, and no frontier model is Pareto-optimal; this is a 4-billion-parameter model matching the best frontier system’s quality at one sixty-sixth of the list price. On quality, at a matched output budget, the best open model significantly exceeds every frontier model (0.711 balanced accuracy against 0.606–0.653), although the strongest frontier system is level with the next two open models. Two findings bound the practical reading. A 125-million-parameter detector fine-tuned on PrimeVul outperforms all eight LLMs (0.765), so where in-distribution labels exist, a small task-specific model is the better instrument. It should be noted that nothing here is deployment-ready: at the natural 1:44 prevalence, precision is 2.3–8.2%.

ComputersVol. 15(9)
FH JOANNEUM University of Applied Sciences (AT), Graz University of Technology (AT)
Openalex Percentile: Top 4%
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Open or Frontier? A Cost- and Energy-Aware Benchmark of Large Language Models for Software Vulnerability Detection — Wolfgang Slany, Patrick Deininger · Computers (2026) | TGRS Research Map | TGRS