VSAT: Estimating the Completeness of Multi-Lens LLM Security Audits of Web Applications via Capture-Recapture

{"VSAT":[0,59,98],"(Vulnerability":[1],"Saturation":[2],"Auditing)":[3],"makes":[4],"the":[5,25,37,56,94,149],"completeness":[6,75,155],"of":[7,55],"an":[8,139],"LLM":[9,21],"security":[10,165],"audit":[11,22],"a":[12,33,42,51,66,73,84,90,116,121,125,158],"measurable,":[13],"calibrated":[14],"quantity.":[15],"It":[16],"runs":[17],"several":[18],"deliberately":[19],"diverse":[20],"\\"lenses\\"":[23],"over":[24,120],"same":[26],"codebase":[27],"and":[28,49,82,103,128,131,157,177],"treats":[29],"each":[30],"lens":[31],"as":[32,89],"capture":[34],"occasion,":[35],"so":[36],"overlap":[38],"structure":[39],"yields":[40],"(i)":[41],"far":[43],"more":[44],"complete":[45],"union":[46],"vulnerability":[47],"list":[48],"(ii)":[50],"Chao2":[52],"richness":[53],"estimate":[54,135],"undiscovered":[57],"population.":[58],"combines":[60],"this":[61,147,181],"statistical":[62],"discovery":[63,118],"saturation":[64,159],"with":[65,124],"deterministic":[67],"OWASP":[68],"ASVS":[69],"structural":[70],"coverage":[71],"into":[72],"single":[74,122],"score":[76],"(Phi":[77],"=":[78],"C_struct":[79],"x":[80],"C_hat)":[81],"derives":[83],"saturation-based":[85],"stopping":[86,160],"rule.":[87],"Implemented":[88],"security-audit":[91],"skill":[92],"on":[93,107],"cc-rsg-web":[95],"agentic":[96],"platform,":[97],"attains":[99],"94.7-100%":[100],"category":[101],"recall":[102],"99.1%":[104],"code-verified":[105],"precision":[106],"three":[108],"documented":[109],"benchmark":[110],"applications":[111],"(NodeGoat,":[112],"django.nV,":[113],"DVWA),":[114],"delivers":[115],"2.4-3.4x":[117],"uplift":[119],"pass":[123],"per-finding":[126],"proof-of-concept":[127],"regression":[129],"test,":[130],"its":[132],"non-zero":[133],"residual":[134],"is":[136,148],"corroborated":[137],"by":[138],"independent":[140],"real-world":[141],"field":[142],"validation.":[143],"To":[144],"our":[145],"knowledge":[146],"first":[150],"method":[151],"to":[152,162],"bring":[153],"capture-recapture":[154],"estimation":[156],"rule":[161],"LLM-based":[163],"web-application":[164],"auditing.":[166],"Reproduction":[167],"artefacts":[168],"(metrics,":[169],"incidence":[170],"matrices,":[171],"analysis":[172],"scripts,":[173],"injection":[174],"answer":[175],"keys,":[176],"matcher)":[178],"added":[179],"in":[180],"version.":[182]}

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-21
DOI
https://doi.org/10.5281/zenodo.21402201
Primary Topic
Web Application Security Vulnerabilities
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

VSAT: Estimating the Completeness of Multi-Lens LLM Security Audits of Web Applications via Capture-Recapture

Daishiro Hirashima
Zenodo (CERN European Organization for Nuclear Research)
Web Application Security Vulnerabilities
preprint

VSAT: Estimating the Completeness of Multi-Lens LLM Security Audits of Web Applications via Capture-Recapture

Daishiro Hirashima
preprint en

Abstract

VSAT (Vulnerability Saturation Auditing) makes the completeness of an LLM security audit a measurable, calibrated quantity. It runs several deliberately diverse LLM audit "lenses" over the same codebase and treats each lens as a capture occasion, so the overlap structure yields (i) a far more complete union vulnerability list and (ii) a Chao2 richness estimate of the undiscovered population. VSAT combines this statistical discovery saturation with a deterministic OWASP ASVS structural coverage into a single completeness score (Phi = C_struct x C_hat) and derives a saturation-based stopping rule. Implemented as a security-audit skill on the cc-rsg-web agentic platform, VSAT attains 94.7-100% category recall and 99.1% code-verified precision on three documented benchmark applications (NodeGoat, django.nV, DVWA), delivers a 2.4-3.4x discovery uplift over a single pass with a per-finding proof-of-concept and regression test, and its non-zero residual estimate is corroborated by an independent real-world field validation. To our knowledge this is the first method to bring capture-recapture completeness estimation and a saturation stopping rule to LLM-based web-application security auditing. Reproduction artefacts (metrics, incidence matrices, analysis scripts, injection answer keys, and matcher) added in this version.

Zenodo (CERN European Organization for Nuclear Research)
Toyobo (Japan) (JP), Toyo Engineering (Japan) (JP)
Peace, Justice and strong institutions
Web Application Security Vulnerabilities
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.