Benchmarking Generative AI Models for Cybersecurity Vulnerability Identification Using 50 Curated Vulnerability Cases: An Exploratory Single-Pass Comparative Study of ChatGPT, Gemini, and Claude

LLMs are being rapidly examined for cybersecurity applications, including identifying software vulnerabilities, conducting secure code reviews, and performing security analysis. Despite recent developments in Generative Artificial Intelligence (GenAI), there is a lack of empirical evidence on the relative performance of state-of-the-art commercial LLMs on discovering and categorizing software vulnerabilities. In this paper, we benchmarked the cybersecurity vulnerability-identification capabilities of ChatGPT, Gemini, and Claude on a dataset of fifty curated vulnerability cases from various Common Weakness Enumeration (CWE) categories. All the models were given the same prompts in a comparative benchmarking setup in a standardized testing environment. Responses were scored using a CWE-based scoring structure that includes exact matches, related matches based on established MITRE CWE links, and incorrect classifications. Descriptive statistics and the Friedman test were used to analyze performance. Descriptive analyses showed that ChatGPT obtained the highest weighted score (72/100), followed by Claude (66/100) and Gemini (64/100). ChatGPT also achieved the highest exact classification rate and the shortest observed response time under the experimental conditions. However, the inferential analysis indicated no statistically significant differences among the evaluated models (Friedman χ²(2) = 3.170, p = .205). These findings suggest broadly comparable vulnerability identification across the three LLMs under the exploratory, single-iteration benchmark conditions employed in the study. All evaluated models demonstrated meaningful capability for cybersecurity vulnerability identification. Although ChatGPT obtained the highest descriptive benchmark scores, inferential statistical analysis indicated no statistically significant performance differences among the evaluated models.

Authors

Institutions

Publication Details

Journal
International Journal of Transformative Multidisciplinary Studies
Published
2026-10-05
DOI
https://doi.org/10.66074/r9l7p4j8w
Primary Topic
Web Application Security Vulnerabilities
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Benchmarking Generative AI Models for Cybersecurity Vulnerability Identification Using 50 Curated Vulnerability Cases: An Exploratory Single-Pass Comparative Study of ChatGPT, Gemini, and Claude

Kristine T. Soberano, Emanuel Dungon, Mark Jan Irog, Hans Luther Salva et al.
International Journal of Transformative Multidisciplinary Studies
Web Application Security Vulnerabilities
article

Benchmarking Generative AI Models for Cybersecurity Vulnerability Identification Using 50 Curated Vulnerability Cases: An Exploratory Single-Pass Comparative Study of ChatGPT, Gemini, and Claude

Kristine T. Soberano, Emanuel Dungon, Mark Jan Irog, Hans Luther Salva, Rod Albert Aspera, Rosa Mae Bascar, Serafin Palmares, Julse Lorenz Alvior
article en

Abstract

LLMs are being rapidly examined for cybersecurity applications, including identifying software vulnerabilities, conducting secure code reviews, and performing security analysis. Despite recent developments in Generative Artificial Intelligence (GenAI), there is a lack of empirical evidence on the relative performance of state-of-the-art commercial LLMs on discovering and categorizing software vulnerabilities. In this paper, we benchmarked the cybersecurity vulnerability-identification capabilities of ChatGPT, Gemini, and Claude on a dataset of fifty curated vulnerability cases from various Common Weakness Enumeration (CWE) categories. All the models were given the same prompts in a comparative benchmarking setup in a standardized testing environment. Responses were scored using a CWE-based scoring structure that includes exact matches, related matches based on established MITRE CWE links, and incorrect classifications. Descriptive statistics and the Friedman test were used to analyze performance. Descriptive analyses showed that ChatGPT obtained the highest weighted score (72/100), followed by Claude (66/100) and Gemini (64/100). ChatGPT also achieved the highest exact classification rate and the shortest observed response time under the experimental conditions. However, the inferential analysis indicated no statistically significant differences among the evaluated models (Friedman χ²(2) = 3.170, p = .205). These findings suggest broadly comparable vulnerability identification across the three LLMs under the exploratory, single-iteration benchmark conditions employed in the study. All evaluated models demonstrated meaningful capability for cybersecurity vulnerability identification. Although ChatGPT obtained the highest descriptive benchmark scores, inferential statistical analysis indicated no statistically significant performance differences among the evaluated models.

International Journal of Transformative Multidisciplinary StudiesVol. 2(4)
State University of Northern Negros (PH)
Openalex Percentile: Top 5%
Web Application Security Vulnerabilities
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.