From policy to practice: clause‑grounded answers with retrieval‑augmented generation

Retrieval-augmented generation (RAG) supports policy question answering by grounding responses in institutional documents. This study describes a campus assistant developed with Azure OpenAI GPT-4 and an Azure Cosmos DB knowledge store containing university regulations. The evaluation covers 51 regulation documents organised into five categories and a question-answer set comprising 1,742 items. Human reviewers classified 1,644 system replies as correct, yielding an overall response accuracy of 94.37%. Response accuracy is an aggregate human-judged correctness measure, not exact match. Because the study did not retain item-level evaluation materials, it does not report separate measures of retrieval quality, citation fidelity, unsupported content, latency, or user experience.

Authors

Institutions

Publication Details

Journal
Journal of Experimental & Theoretical Artificial Intelligence
Published
2026-09-21
DOI
https://doi.org/10.1080/0952813x.2026.2729304
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

From policy to practice: clause‑grounded answers with retrieval‑augmented generation

Wen-Lung Tsai, Ren-Qi Huang
Journal of Experimental & Theoretical Artificial Intelligence
Topic Modeling
article

From policy to practice: clause‑grounded answers with retrieval‑augmented generation

Wen-Lung Tsai, Ren-Qi Huang
article en

Abstract

Retrieval-augmented generation (RAG) supports policy question answering by grounding responses in institutional documents. This study describes a campus assistant developed with Azure OpenAI GPT-4 and an Azure Cosmos DB knowledge store containing university regulations. The evaluation covers 51 regulation documents organised into five categories and a question-answer set comprising 1,742 items. Human reviewers classified 1,644 system replies as correct, yielding an overall response accuracy of 94.37%. Response accuracy is an aggregate human-judged correctness measure, not exact match. Because the study did not retain item-level evaluation materials, it does not report separate measures of retrieval quality, citation fidelity, unsupported content, latency, or user experience.

Journal of Experimental & Theoretical Artificial Intelligence
National Taipei University of Business (TW)
Quality Education
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.