SPARQL-LLM: Real-Time SPARQL Query Generation from Natural Language Questions

The advent of large language models is contributing to the emergence of novel approaches that promise to better tackle the challenge of generating structured queries, such as SPARQL queries, from natural language. However, these new approaches mostly focus on response accuracy while ignoring other evaluation criteria, such as runtime and cost to generate SPARQL queries. Consequently, they are often not production-ready or easy to deploy over real-world knowledge graphs with good accuracy. To mitigate these issues, in this paper, we describe and systematically evaluate SPARQL-LLM, an open-source and triplestore-agnostic approach, powered by lightweight metadata, that generates SPARQL queries from natural language text. First, we describe its architecture, which consists of dedicated components for metadata indexing, prompt building, and query generation and execution. Then, we evaluate it based on a state-of-the-art challenge with multilingual questions, and a collection of questions from three of the most prevalent knowledge graphs within the field of bioinformatics. Our results demonstrate a substantial improvement of up to \\(59\\% \\) in F1 score over the second-best system participating in the challenge, adaptability to high-resource languages such as English, Spanish, and German, as well as ability to form complex bioinformatics queries. Furthermore, our results show that our system is up to 27 × faster than the second-best system participating in the challenge, while costing a maximum of $0.01 per question, making it suitable for real-time, low-cost text-to-SPARQL applications. SPARQL-LLM is publicly released as an open-source project at https://github.com/sib-swiss/sparql-llm and is currently deployed over real-world decentralized knowledge graphs at https://www.expasy.org/chat.

Authors

Institutions

Publication Details

Journal
ACM Transactions on the Web
Published
2026-09-15
DOI
https://doi.org/10.1145/3847195
Citations
3
Primary Topic
Biomedical Text Mining and Ontologies
Type
article
Field-Weighted Citation Impact
8.28

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

SPARQL-LLM: Real-Time SPARQL Query Generation from Natural Language Questions

Tarcisio Mendes de Farias, Ana-Claudia Sima, Vincent Emonet, Panayiotis Smeros et al.
3 citations
ACM Transactions on the Web
Biomedical Text Mining and Ontologies
8.28
article

SPARQL-LLM: Real-Time SPARQL Query Generation from Natural Language Questions

Tarcisio Mendes de Farias, Ana-Claudia Sima, Vincent Emonet, Panayiotis Smeros, Ruijie Wang
article en
3 citations

Abstract

The advent of large language models is contributing to the emergence of novel approaches that promise to better tackle the challenge of generating structured queries, such as SPARQL queries, from natural language. However, these new approaches mostly focus on response accuracy while ignoring other evaluation criteria, such as runtime and cost to generate SPARQL queries. Consequently, they are often not production-ready or easy to deploy over real-world knowledge graphs with good accuracy. To mitigate these issues, in this paper, we describe and systematically evaluate SPARQL-LLM, an open-source and triplestore-agnostic approach, powered by lightweight metadata, that generates SPARQL queries from natural language text. First, we describe its architecture, which consists of dedicated components for metadata indexing, prompt building, and query generation and execution. Then, we evaluate it based on a state-of-the-art challenge with multilingual questions, and a collection of questions from three of the most prevalent knowledge graphs within the field of bioinformatics. Our results demonstrate a substantial improvement of up to \(59\% \) in F1 score over the second-best system participating in the challenge, adaptability to high-resource languages such as English, Spanish, and German, as well as ability to form complex bioinformatics queries. Furthermore, our results show that our system is up to 27 × faster than the second-best system participating in the challenge, while costing a maximum of $0.01 per question, making it suitable for real-time, low-cost text-to-SPARQL applications. SPARQL-LLM is publicly released as an open-source project at https://github.com/sib-swiss/sparql-llm and is currently deployed over real-world decentralized knowledge graphs at https://www.expasy.org/chat.

ACM Transactions on the Web
SIB Swiss Institute of Bioinformatics (CH), University of Zurich (CH), University of Lausanne (CH)
Schweizerischer Nationalfonds zur Förderung der Wissenschaftlichen Forschung
Openalex Percentile: Top 4%
Biomedical Text Mining and Ontologies
8.28
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.