RiskyPy: Toward Automated Prompt Generation Tasks Calibrated to Model Proficiency for Joint Security and Functionality Evaluation of Generated Code
Evaluating approaches that improve the security of LLM code generation requires datasets adequate for the joint assessment of security and functionality while remaining adapted to the capabilities of the models being studied. However, existing benchmarks are costly to produce and are generally not calibrated to a chosen range of model capabilities. We introduce RiskyPy, an automated pipeline that generates security-oriented code-generation datasets calibrated to model proficiency using functionality and vulnerability canaries. We instantiate RiskyPy for Python using domains covering 18 CWEs from the 2025 CWE Top 25, producing 73 prompts. Experiments on Qwen3-4B and Seed-Coder-8B show that the generated tasks remain largely functionally accessible while frequently eliciting insecure implementations, leaving room for joint evaluation of security and functionality of security improvement approaches.
Authors
- Naouel Moha (ORCID: https://orcid.org/0000-0001-9252-9937)
- Jacques Klein (ORCID: https://orcid.org/0000-0003-4052-475X)
- Florent Avellaneda (ORCID: https://orcid.org/0000-0003-1030-5388)
- Tegawendé François Bissyandé (ORCID: https://orcid.org/0000-0001-7270-9869)
- Melissa Tessa (ORCID: https://orcid.org/0009-0008-7525-5522)
- Daniele Lunghi (ORCID: https://orcid.org/0000-0003-1324-981X)
- Zacharie Chenail-Larcher (ORCID: https://orcid.org/0009-0002-7464-6659)
Institutions
- Université du Québec à Montréal (CA)
- University of Luxembourg (LU)
- École de Technologie Supérieure (CA)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-01
- DOI
- https://doi.org/10.5281/zenodo.23088676
- Primary Topic
- Software Engineering Research
- Type
- article
- Field-Weighted Citation Impact
- 0.00