Benchmarking Generative AI Models for Cybersecurity Vulnerability Identification Using 50 Curated Vulnerability Cases: An Exploratory Single-Pass Comparative Study of ChatGPT, Gemini, and Claude
LLMs are being rapidly examined for cybersecurity applications, including identifying software vulnerabilities, conducting secure code reviews, and performing security analysis. Despite recent developments in Generative Artificial Intelligence (GenAI), there is a lack of empirical evidence on the relative performance of state-of-the-art commercial LLMs on discovering and categorizing software vulnerabilities. In this paper, we benchmarked the cybersecurity vulnerability-identification capabilities of ChatGPT, Gemini, and Claude on a dataset of fifty curated vulnerability cases from various Common Weakness Enumeration (CWE) categories. All the models were given the same prompts in a comparative benchmarking setup in a standardized testing environment. Responses were scored using a CWE-based scoring structure that includes exact matches, related matches based on established MITRE CWE links, and incorrect classifications. Descriptive statistics and the Friedman test were used to analyze performance. Descriptive analyses showed that ChatGPT obtained the highest weighted score (72/100), followed by Claude (66/100) and Gemini (64/100). ChatGPT also achieved the highest exact classification rate and the shortest observed response time under the experimental conditions. However, the inferential analysis indicated no statistically significant differences among the evaluated models (Friedman χ²(2) = 3.170, p = .205). These findings suggest broadly comparable vulnerability identification across the three LLMs under the exploratory, single-iteration benchmark conditions employed in the study. All evaluated models demonstrated meaningful capability for cybersecurity vulnerability identification. Although ChatGPT obtained the highest descriptive benchmark scores, inferential statistical analysis indicated no statistically significant performance differences among the evaluated models.
Authors
- Kristine T. Soberano
- Emanuel Dungon
- Mark Jan Irog
- Hans Luther Salva
- Rod Albert Aspera
- Rosa Mae Bascar
- Serafin Palmares
- Julse Lorenz Alvior
Institutions
- State University of Northern Negros (PH)
Publication Details
- Journal
- International Journal of Transformative Multidisciplinary Studies
- Published
- 2026-10-05
- DOI
- https://doi.org/10.66074/r9l7p4j8w
- Primary Topic
- Web Application Security Vulnerabilities
- Type
- article
- Field-Weighted Citation Impact
- 0.00