Empirical Post-Generation Correctness Assessment for Natural-Language Access to Urban PostGIS Databases

Natural-language interfaces make urban spatial databases accessible, but generated structured query language (SQL) can execute successfully while returning a semantically incorrect answer. We evaluate a post-generation correctness-ranking layer on 18,900 candidates from Nanjing, Wuhan, and Shenzhen. The original city split is a controlled, highly template-aligned benchmark rather than a structure-disjoint deployment test. Under strict source-side separation of fitting, isotonic calibration, threshold selection, and target evaluation, Spatial-safe improves executed non-empty area under the receiver operating characteristic curve (AUROC) in all nine transfer directions by 0.062 on average. Grouped-threshold area under the risk–coverage curve (AURC) is favorable in 8/9 directions, and calibration is mixed. Source-selected thresholds raise mean conditional target coverage within the executed non-empty operating subset from 0.457 to 0.669, while mean empirical risk rises from 0.110 to 0.136. Unseen-template and SQL-skeleton-disjoint effects are smaller and mixed. To determine whether the existing benchmark could support the adequately powered structure-disjoint comparison, we applied a predeclared power/data-sufficiency feasibility gate. With only 77 independent SQL-skeleton groups, 42/60 required cells fail the gate; therefore, the present benchmark cannot support an adequately powered structure-disjoint performance comparison without additional skeleton-diverse questions. We do not substitute an underpowered estimate for that missing evidence. The evidence therefore supports empirical correctness-ranking improvement in this controlled benchmark, not universal structure-disjoint generalization, formal SQL verification, benchmark-wide human validation of all correctness labels, or target-domain risk guarantees.

Authors

Institutions

Publication Details

Journal
ISPRS International Journal of Geo-Information
Published
2026-09-22
DOI
https://doi.org/10.3390/ijgi15100436
Primary Topic
Geographic Information Systems Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Empirical Post-Generation Correctness Assessment for Natural-Language Access to Urban PostGIS Databases

Jiqiu Deng, Zhiyong Guo, C L Zhang, Hui Zhang et al.
ISPRS International Journal of Geo-Information
Geographic Information Systems Studies
article

Empirical Post-Generation Correctness Assessment for Natural-Language Access to Urban PostGIS Databases

Jiqiu Deng, Zhiyong Guo, C L Zhang, Hui Zhang, Xiao Ma, Liji Sun, Longbo Li
article en

Abstract

Natural-language interfaces make urban spatial databases accessible, but generated structured query language (SQL) can execute successfully while returning a semantically incorrect answer. We evaluate a post-generation correctness-ranking layer on 18,900 candidates from Nanjing, Wuhan, and Shenzhen. The original city split is a controlled, highly template-aligned benchmark rather than a structure-disjoint deployment test. Under strict source-side separation of fitting, isotonic calibration, threshold selection, and target evaluation, Spatial-safe improves executed non-empty area under the receiver operating characteristic curve (AUROC) in all nine transfer directions by 0.062 on average. Grouped-threshold area under the risk–coverage curve (AURC) is favorable in 8/9 directions, and calibration is mixed. Source-selected thresholds raise mean conditional target coverage within the executed non-empty operating subset from 0.457 to 0.669, while mean empirical risk rises from 0.110 to 0.136. Unseen-template and SQL-skeleton-disjoint effects are smaller and mixed. To determine whether the existing benchmark could support the adequately powered structure-disjoint comparison, we applied a predeclared power/data-sufficiency feasibility gate. With only 77 independent SQL-skeleton groups, 42/60 required cells fail the gate; therefore, the present benchmark cannot support an adequately powered structure-disjoint performance comparison without additional skeleton-diverse questions. We do not substitute an underpowered estimate for that missing evidence. The evidence therefore supports empirical correctness-ranking improvement in this controlled benchmark, not universal structure-disjoint generalization, formal SQL verification, benchmark-wide human validation of all correctness labels, or target-domain risk guarantees.

ISPRS International Journal of Geo-InformationVol. 15(10)
Central South University (CN), Ministry of Natural Resources (CN), Zhaotong University (CN), Department of Science and Technology of Hunan Province (CN)
Sustainable cities and communities
Openalex Percentile: Top 3%
Geographic Information Systems Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.