Beyond Compilation: an Empirical Framework-aware Assessment of Security and Reliability in Spring Boot Applications Generated with Large Language Models

Today's large language models (LLMs) can now create full-scale back-end applications. As such, evaluation becomes more complex than simply checking if the application compiles. Within Spring Boot, the code may compile successfully, but still fail to start up, violate an expected security property, or utilize a configuration that would be unsafe for practical use. SpringLLM-Bench was developed to address this gap. The pilot benchmark evaluates full Spring Boot back-ends created by LLMs in five enterprise-level tasks: JWT-based authentication with role-based access control, transactional money transfers, concurrent inventory purchases, secure JPA-backed searches, and Kafka-based payment consumption. Each of three modern LLMs produced three individual applications for each of the five tasks under a one-shot, no-repair protocol. This resulted in 45 frozen applications. We assessed each of these applications in separate stages instead of reducing each sample to a single pass/fail result: Maven build, stable Spring Boot startup, independent functional testing, task-specific security/reliability testing, and source-level static analysis. Thirty-six of the 45 applications (80.0%) successfully compiled; 32 (71.1%) reached a stable startup state; and 31 (68.9%) passed the functional oracle. The performance of each task was not uniform. All nine concurrency applications were functionally correct; however, only one of the nine Kafka applications passed. We then performed a source-level Semgrep Community Edition 1.178.0 scan on the remaining 1,177 files (after excluding build artifacts and preserving raw outputs) using 203 rules. The scan identified three potential secrets. Manual data-flow analysis confirmed that all three potential secrets were hard-coded JWT signing keys used to generate and verify tokens. Two of the three affected applications were otherwise functionally correct.This translates into a Functionally Correct but Security-Defective (FCSD) rate of 2/45 (4.4%) or 2/31 (6.5%) for the subset of functionally correct applications, while there were no instances of Functionally Correct but Reliability-Defective (FCRD) applications.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.23066919
Primary Topic
Scientific Computing and Data Management
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Beyond Compilation: an Empirical Framework-aware Assessment of Security and Reliability in Spring Boot Applications Generated with Large Language Models

Toshinee Bhasin
Zenodo (CERN European Organization for Nuclear Research)
Scientific Computing and Data Management
preprint

Beyond Compilation: an Empirical Framework-aware Assessment of Security and Reliability in Spring Boot Applications Generated with Large Language Models

Toshinee Bhasin
preprint en

Abstract

Today's large language models (LLMs) can now create full-scale back-end applications. As such, evaluation becomes more complex than simply checking if the application compiles. Within Spring Boot, the code may compile successfully, but still fail to start up, violate an expected security property, or utilize a configuration that would be unsafe for practical use. SpringLLM-Bench was developed to address this gap. The pilot benchmark evaluates full Spring Boot back-ends created by LLMs in five enterprise-level tasks: JWT-based authentication with role-based access control, transactional money transfers, concurrent inventory purchases, secure JPA-backed searches, and Kafka-based payment consumption. Each of three modern LLMs produced three individual applications for each of the five tasks under a one-shot, no-repair protocol. This resulted in 45 frozen applications. We assessed each of these applications in separate stages instead of reducing each sample to a single pass/fail result: Maven build, stable Spring Boot startup, independent functional testing, task-specific security/reliability testing, and source-level static analysis. Thirty-six of the 45 applications (80.0%) successfully compiled; 32 (71.1%) reached a stable startup state; and 31 (68.9%) passed the functional oracle. The performance of each task was not uniform. All nine concurrency applications were functionally correct; however, only one of the nine Kafka applications passed. We then performed a source-level Semgrep Community Edition 1.178.0 scan on the remaining 1,177 files (after excluding build artifacts and preserving raw outputs) using 203 rules. The scan identified three potential secrets. Manual data-flow analysis confirmed that all three potential secrets were hard-coded JWT signing keys used to generate and verify tokens. Two of the three affected applications were otherwise functionally correct.This translates into a Functionally Correct but Security-Defective (FCSD) rate of 2/45 (4.4%) or 2/31 (6.5%) for the subset of functionally correct applications, while there were no instances of Functionally Correct but Reliability-Defective (FCRD) applications.

Zenodo (CERN European Organization for Nuclear Research)
Scientific Computing and Data Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Beyond Compilation: an Empirical Framework-aware Assessment of Security and Reliability in Spring Boot Applications Generated with Large Language Models — Toshinee Bhasin · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS