Known-Bad Names, Unknown-Bad Uses: What Local Code Models Detect When They Review Cryptographic API Misuse

Small code language models are now easy to run on a developer's own laptop, and one thing people ask them to do is a quick security pass over code before it ships. I wanted to know how far that trust holds for one narrow but dangerous flaw family, cryptographic API misuse. I built CryptoBench, a set of thirty-six vulnerable snippets spread across nine misuse classes, plus eighteen secure counterparts, two per class, so a model earns nothing by calling everything unsafe, and I ran it against seven open code models from 0.5B to 14B parameters. Detection turned out to be strongly class dependent. Pooled over the three models that separate safe from unsafe code, MD5 or SHA-1 misuse was flagged 82% of the time against a 3% false-positive rate on its secure controls, and ECB mode 90% against 33%. The controls also cut the other way. Disabled TLS verification was flagged 97% of the time, yet correctly verified TLS code was flagged 87% of the time, and hardcoded keys (58%) were flagged less often than keys read from the environment (63%), so neither class shows real discrimination. The models miss misuse that hides inside ordinary use of a general-purpose API. The clearest case is a weak random number generator standing in for a cryptographic one, caught 38% of the time with no false alarms, and going from a 7B to a 14B model raised that from 6 to 11 of 20, a difference that is not statistically significant. Two of the highest raw scores in my set came from models that label almost everything vulnerable, which a matched secure control exposes at once. I release the benchmark and the full per-trial results.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23113860
Primary Topic
Advanced Malware Detection Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Known-Bad Names, Unknown-Bad Uses: What Local Code Models Detect When They Review Cryptographic API Misuse

Sunny Chokshi
Zenodo (CERN European Organization for Nuclear Research)
Advanced Malware Detection Techniques
preprint

Known-Bad Names, Unknown-Bad Uses: What Local Code Models Detect When They Review Cryptographic API Misuse

Sunny Chokshi
preprint en

Abstract

Small code language models are now easy to run on a developer's own laptop, and one thing people ask them to do is a quick security pass over code before it ships. I wanted to know how far that trust holds for one narrow but dangerous flaw family, cryptographic API misuse. I built CryptoBench, a set of thirty-six vulnerable snippets spread across nine misuse classes, plus eighteen secure counterparts, two per class, so a model earns nothing by calling everything unsafe, and I ran it against seven open code models from 0.5B to 14B parameters. Detection turned out to be strongly class dependent. Pooled over the three models that separate safe from unsafe code, MD5 or SHA-1 misuse was flagged 82% of the time against a 3% false-positive rate on its secure controls, and ECB mode 90% against 33%. The controls also cut the other way. Disabled TLS verification was flagged 97% of the time, yet correctly verified TLS code was flagged 87% of the time, and hardcoded keys (58%) were flagged less often than keys read from the environment (63%), so neither class shows real discrimination. The models miss misuse that hides inside ordinary use of a general-purpose API. The clearest case is a weak random number generator standing in for a cryptographic one, caught 38% of the time with no false alarms, and going from a 7B to a 14B model raised that from 6 to 11 of 20, a difference that is not statistically significant. Two of the highest raw scores in my set came from models that label almost everything vulnerable, which a matched secure control exposes at once. I release the benchmark and the full per-trial results.

Zenodo (CERN European Organization for Nuclear Research)
University of the Cumberlands (US)
Advanced Malware Detection Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.