VeriCrypt-Agent: Evidence-Grounded Multi-Agent Verification for Cryptographic Dependency Updates

Cryptographic dependency updates can install cleanly and pass visible regression tests while changing project-facing behavior or leaving an application inside a vulnerable version range. This study presents VeriCrypt-Agent, an evidence-grounded workflow for bounded-autonomy cryptographic maintenance. The system combines a frozen snapshot of release/advisory evidence, a public catalog of project-level executable probes, cross-version execution under old and candidate dependency versions, role-specialized LLM reasoning, and a deterministic decision gate. Automatic merging is allowed only when mandatory probes pass, policy constraints are satisfied, and the audit record is complete. On CryptoUpdate-Mini-30, ten unique package-version transitions are evaluated through 30 controlled executions. Visible tests accept all five non-mergeable transitions, while advisory-only screening accepts two of five. The deterministic All-Probes Gate performs best on this benchmark, with perfect transition-level decisions and no review cases, when all relevant probes are already known and inexpensive. The full workflow observes no unsafe auto-merges (0/5; exact 95% CI: 0.000–0.522) and routes 9/30 executions to review (0.30; exact 95% CI: 0.147–0.494). Removing grounding accepts three of five non-mergeable transitions, and a fixed probe budget accepts two of five. On 48 blinded static crypto-API audit files, including 32 unsafe cases, two-pass review detects 20/32 unsafe cases (0.625; exact 95% CI: 0.437–0.789) and reaches macro-F1 =0.556. These point estimates characterize evidence-backed release decisions only under the evaluated conditions; they do not establish general superiority over exhaustive deterministic testing or a deployable static vulnerability detector.

Authors

Institutions

Publication Details

Journal
Computers
Published
2026-09-21
DOI
https://doi.org/10.3390/computers15090638
Primary Topic
Advanced Malware Detection Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

VeriCrypt-Agent: Evidence-Grounded Multi-Agent Verification for Cryptographic Dependency Updates

А. С. Бородулин, В С Тынченко, Vladimir Nelyub, Andrei Gantimurov et al.
Computers
Advanced Malware Detection Techniques
article

VeriCrypt-Agent: Evidence-Grounded Multi-Agent Verification for Cryptographic Dependency Updates

А. С. Бородулин, В С Тынченко, Vladimir Nelyub, Andrei Gantimurov, Ivan Pavlovich Malashin, Dmitry Martysyuk
article en

Abstract

Cryptographic dependency updates can install cleanly and pass visible regression tests while changing project-facing behavior or leaving an application inside a vulnerable version range. This study presents VeriCrypt-Agent, an evidence-grounded workflow for bounded-autonomy cryptographic maintenance. The system combines a frozen snapshot of release/advisory evidence, a public catalog of project-level executable probes, cross-version execution under old and candidate dependency versions, role-specialized LLM reasoning, and a deterministic decision gate. Automatic merging is allowed only when mandatory probes pass, policy constraints are satisfied, and the audit record is complete. On CryptoUpdate-Mini-30, ten unique package-version transitions are evaluated through 30 controlled executions. Visible tests accept all five non-mergeable transitions, while advisory-only screening accepts two of five. The deterministic All-Probes Gate performs best on this benchmark, with perfect transition-level decisions and no review cases, when all relevant probes are already known and inexpensive. The full workflow observes no unsafe auto-merges (0/5; exact 95% CI: 0.000–0.522) and routes 9/30 executions to review (0.30; exact 95% CI: 0.147–0.494). Removing grounding accepts three of five non-mergeable transitions, and a fixed probe budget accepts two of five. On 48 blinded static crypto-API audit files, including 32 unsafe cases, two-pass review detects 20/32 unsafe cases (0.625; exact 95% CI: 0.437–0.789) and reaches macro-F1 =0.556. These point estimates characterize evidence-backed release decisions only under the evaluated conditions; they do not establish general superiority over exhaustive deterministic testing or a deployable static vulnerability detector.

ComputersVol. 15(9)
Bauman Moscow State Technical University (RU)
Openalex Percentile: Top 10%
Advanced Malware Detection Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.