LLM4VV: Building an LLMJ Training Dataset for OpenMP and OpenACC Compiler Test Evaluation
Compilers that implement directive-based parallel programming models require extensive test suites to validate their outputs. These test suites are complex and time-consuming to produce by hand, so being able to create them with LLMs could vastly improve the quality of compiler implementations by quickly increasing the number of test cases compiler developers can validate against. Prior work \cite{sollenberger2025arxiv} have explored such systems, with LLM4VV suggesting a Dual-Agent system comprised of a generative agent and a discriminative agent. This work aims to create the training dataset necessary to fine-tune an LLM into a quality discriminative agent as specified by LLM4VV by exploring alternative methods for providing relevant feature information from applicable specifications. The process resulting from our exploration led to the creation of high-quality synthetic data that adequately captures syntactical, semantic, and conceptual errors that may be present in generated parallel code. We then formatted the relevant information and synthetic data into both a training dataset and a task-specific benchmark that determines an LLM's aptitude as a discriminative agent. Our dataset enables the training of LLMs into discriminative agents and makes it possible to compare discriminative agents against each other, ensuring the best possible quality control for a Dual-Agent system.
Authors
- Sunita Chandrasekaran (ORCID: https://orcid.org/0000-0002-3560-9428)
- Zachariah Sollenberger (ORCID: https://orcid.org/0009-0008-6669-8778)
- Saieda Ali Zada (ORCID: https://orcid.org/0009-0000-5368-4512)
- Rahul Patel (ORCID: https://orcid.org/0009-0009-8739-7138)
Institutions
- University of Delaware (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-08
- DOI
- https://doi.org/10.5281/zenodo.23249421
- Primary Topic
- Software Testing and Debugging Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00