Beyond Compilation: an Empirical Framework-aware Assessment of Security and Reliability in Spring Boot Applications Generated with Large Language Models
Today's large language models (LLMs) can now create full-scale back-end applications. As such, evaluation becomes more complex than simply checking if the application compiles. Within Spring Boot, the code may compile successfully, but still fail to start up, violate an expected security property, or utilize a configuration that would be unsafe for practical use. SpringLLM-Bench was developed to address this gap. The pilot benchmark evaluates full Spring Boot back-ends created by LLMs in five enterprise-level tasks: JWT-based authentication with role-based access control, transactional money transfers, concurrent inventory purchases, secure JPA-backed searches, and Kafka-based payment consumption. Each of three modern LLMs produced three individual applications for each of the five tasks under a one-shot, no-repair protocol. This resulted in 45 frozen applications. We assessed each of these applications in separate stages instead of reducing each sample to a single pass/fail result: Maven build, stable Spring Boot startup, independent functional testing, task-specific security/reliability testing, and source-level static analysis. Thirty-six of the 45 applications (80.0%) successfully compiled; 32 (71.1%) reached a stable startup state; and 31 (68.9%) passed the functional oracle. The performance of each task was not uniform. All nine concurrency applications were functionally correct; however, only one of the nine Kafka applications passed. We then performed a source-level Semgrep Community Edition 1.178.0 scan on the remaining 1,177 files (after excluding build artifacts and preserving raw outputs) using 203 rules. The scan identified three potential secrets. Manual data-flow analysis confirmed that all three potential secrets were hard-coded JWT signing keys used to generate and verify tokens. Two of the three affected applications were otherwise functionally correct.This translates into a Functionally Correct but Security-Defective (FCSD) rate of 2/45 (4.4%) or 2/31 (6.5%) for the subset of functionally correct applications, while there were no instances of Functionally Correct but Reliability-Defective (FCRD) applications.
Authors
- Toshinee Bhasin
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23066919
- Primary Topic
- Scientific Computing and Data Management
- Type
- preprint