RiskyPy: Toward Automated Prompt Generation Tasks Calibrated to Model Proficiency for Joint Security and Functionality Evaluation of Generated Code

Evaluating approaches that improve the security of LLM code generation requires datasets adequate for the joint assessment of security and functionality while remaining adapted to the capabilities of the models being studied. However, existing benchmarks are costly to produce and are generally not calibrated to a chosen range of model capabilities. We introduce RiskyPy, an automated pipeline that generates security-oriented code-generation datasets calibrated to model proficiency using functionality and vulnerability canaries. We instantiate RiskyPy for Python using domains covering 18 CWEs from the 2025 CWE Top 25, producing 73 prompts. Experiments on Qwen3-4B and Seed-Coder-8B show that the generated tasks remain largely functionally accessible while frequently eliciting insecure implementations, leaving room for joint evaluation of security and functionality of security improvement approaches.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-01
DOI
https://doi.org/10.5281/zenodo.23088676
Primary Topic
Software Engineering Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

RiskyPy: Toward Automated Prompt Generation Tasks Calibrated to Model Proficiency for Joint Security and Functionality Evaluation of Generated Code

Naouel Moha, Jacques Klein, Florent Avellaneda, Tegawendé François Bissyandé et al.
Zenodo (CERN European Organization for Nuclear Research)
Software Engineering Research
article

RiskyPy: Toward Automated Prompt Generation Tasks Calibrated to Model Proficiency for Joint Security and Functionality Evaluation of Generated Code

Naouel Moha, Jacques Klein, Florent Avellaneda, Tegawendé François Bissyandé, Melissa Tessa, Daniele Lunghi, Zacharie Chenail-Larcher
article en

Abstract

Evaluating approaches that improve the security of LLM code generation requires datasets adequate for the joint assessment of security and functionality while remaining adapted to the capabilities of the models being studied. However, existing benchmarks are costly to produce and are generally not calibrated to a chosen range of model capabilities. We introduce RiskyPy, an automated pipeline that generates security-oriented code-generation datasets calibrated to model proficiency using functionality and vulnerability canaries. We instantiate RiskyPy for Python using domains covering 18 CWEs from the 2025 CWE Top 25, producing 73 prompts. Experiments on Qwen3-4B and Seed-Coder-8B show that the generated tasks remain largely functionally accessible while frequently eliciting insecure implementations, leaving room for joint evaluation of security and functionality of security improvement approaches.

Zenodo (CERN European Organization for Nuclear Research)
Université du Québec à Montréal (CA), University of Luxembourg (LU), École de Technologie Supérieure (CA)
Openalex Percentile: Top 6%
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

RiskyPy: Toward Automated Prompt Generation Tasks Calibrated to Model Proficiency for Joint Security and Functionality Evaluation of Generated Code — Naouel Moha, Jacques Klein, et al. · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS