Automated Robust Dynamic Generation of KGQA Benchmark Datasets

Knowledge Graph Question Answering (KGQA) systems transform natural-language questions into structured SPARQL queries to retrieve information from Knowledge Graphs (KGs). However, existing KGQA benchmarks are static, prone to obsolescence as KGs evolve, and increasingly unreliable for evaluating systems based on Large Language Models (LLMs) due to memorization effects (i.e., a model has reproduced the correct output from existing training data). This paper presents DynBench, a fully automated and robust framework for the dynamic generation of KGQA benchmark datasets. Unlike prior approaches, DynBench is enabled to automatically generate an arbitrary number of new benchmark datasets while preserving the natural-language surface forms and structural complexity of the original SPARQL queries in comparison to the given original KGQA dataset. Using entity and property substitutions within the same KG, DynBench produces semantically consistent questionquery pairs and automatically validates them without human intervention. We introduce an automatic validation mechanism that ensures semantic and syntactic integrity through backtransformation and metric-based validation. Two dynamic benchmarks were generated from the well-known datasets QALD-9-plus as well as LC-QuAD 2.0 and evaluated using both human experts and automated metrics to show DynBench’s datasetagnostic capabilities. Given our results, we can infer that the Levenshtein distance serves as the most reliable automatic validation measure, achieving a precision of up to 0.96 and demonstrating a strong correlation with human assessments. DynBench thus enables scalable, repeatable, and memorization-resistant dataset generation—providing a foundation for sustainable and fair evaluation of LLM-based KGQA systems.

Authors

Publication Details

Journal
International Journal of Semantic Computing
Published
2026-09-25
DOI
https://doi.org/10.1142/s1793351x26450054
Primary Topic
Advanced Graph Neural Networks
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Automated Robust Dynamic Generation of KGQA Benchmark Datasets

Maria Eltsova, Aleksandr Perevalov, Aleksandr Gashkov, Andreas Both
International Journal of Semantic Computing
Advanced Graph Neural Networks
article

Automated Robust Dynamic Generation of KGQA Benchmark Datasets

Maria Eltsova, Aleksandr Perevalov, Aleksandr Gashkov, Andreas Both
article en

Abstract

Knowledge Graph Question Answering (KGQA) systems transform natural-language questions into structured SPARQL queries to retrieve information from Knowledge Graphs (KGs). However, existing KGQA benchmarks are static, prone to obsolescence as KGs evolve, and increasingly unreliable for evaluating systems based on Large Language Models (LLMs) due to memorization effects (i.e., a model has reproduced the correct output from existing training data). This paper presents DynBench, a fully automated and robust framework for the dynamic generation of KGQA benchmark datasets. Unlike prior approaches, DynBench is enabled to automatically generate an arbitrary number of new benchmark datasets while preserving the natural-language surface forms and structural complexity of the original SPARQL queries in comparison to the given original KGQA dataset. Using entity and property substitutions within the same KG, DynBench produces semantically consistent questionquery pairs and automatically validates them without human intervention. We introduce an automatic validation mechanism that ensures semantic and syntactic integrity through backtransformation and metric-based validation. Two dynamic benchmarks were generated from the well-known datasets QALD-9-plus as well as LC-QuAD 2.0 and evaluated using both human experts and automated metrics to show DynBench’s datasetagnostic capabilities. Given our results, we can infer that the Levenshtein distance serves as the most reliable automatic validation measure, achieving a precision of up to 0.96 and demonstrating a strong correlation with human assessments. DynBench thus enables scalable, repeatable, and memorization-resistant dataset generation—providing a foundation for sustainable and fair evaluation of LLM-based KGQA systems.

International Journal of Semantic Computing
Responsible consumption and production
Openalex Percentile: Top 9%
Advanced Graph Neural Networks
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Automated Robust Dynamic Generation of KGQA Benchmark Datasets — Maria Eltsova, Aleksandr Perevalov, et al. · International Journal of Semantic Computing (2026) | TGRS Research Map | TGRS