LLM4VV: Building an LLMJ Training Dataset for OpenMP and OpenACC Compiler Test Evaluation

Compilers that implement directive-based parallel programming models require extensive test suites to validate their outputs. These test suites are complex and time-consuming to produce by hand, so being able to create them with LLMs could vastly improve the quality of compiler implementations by quickly increasing the number of test cases compiler developers can validate against. Prior work \cite{sollenberger2025arxiv} have explored such systems, with LLM4VV suggesting a Dual-Agent system comprised of a generative agent and a discriminative agent. This work aims to create the training dataset necessary to fine-tune an LLM into a quality discriminative agent as specified by LLM4VV by exploring alternative methods for providing relevant feature information from applicable specifications. The process resulting from our exploration led to the creation of high-quality synthetic data that adequately captures syntactical, semantic, and conceptual errors that may be present in generated parallel code. We then formatted the relevant information and synthetic data into both a training dataset and a task-specific benchmark that determines an LLM's aptitude as a discriminative agent. Our dataset enables the training of LLMs into discriminative agents and makes it possible to compare discriminative agents against each other, ensuring the best possible quality control for a Dual-Agent system.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-08
DOI
https://doi.org/10.5281/zenodo.23249421
Primary Topic
Software Testing and Debugging Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

LLM4VV: Building an LLMJ Training Dataset for OpenMP and OpenACC Compiler Test Evaluation

Sunita Chandrasekaran, Zachariah Sollenberger, Saieda Ali Zada, Rahul Patel
Zenodo (CERN European Organization for Nuclear Research)
Software Testing and Debugging Techniques
article

LLM4VV: Building an LLMJ Training Dataset for OpenMP and OpenACC Compiler Test Evaluation

Sunita Chandrasekaran, Zachariah Sollenberger, Saieda Ali Zada, Rahul Patel
article en

Abstract

Compilers that implement directive-based parallel programming models require extensive test suites to validate their outputs. These test suites are complex and time-consuming to produce by hand, so being able to create them with LLMs could vastly improve the quality of compiler implementations by quickly increasing the number of test cases compiler developers can validate against. Prior work \cite{sollenberger2025arxiv} have explored such systems, with LLM4VV suggesting a Dual-Agent system comprised of a generative agent and a discriminative agent. This work aims to create the training dataset necessary to fine-tune an LLM into a quality discriminative agent as specified by LLM4VV by exploring alternative methods for providing relevant feature information from applicable specifications. The process resulting from our exploration led to the creation of high-quality synthetic data that adequately captures syntactical, semantic, and conceptual errors that may be present in generated parallel code. We then formatted the relevant information and synthetic data into both a training dataset and a task-specific benchmark that determines an LLM's aptitude as a discriminative agent. Our dataset enables the training of LLMs into discriminative agents and makes it possible to compare discriminative agents against each other, ensuring the best possible quality control for a Dual-Agent system.

Zenodo (CERN European Organization for Nuclear Research)
University of Delaware (US)
Openalex Percentile: Top 3%
Software Testing and Debugging Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.