Separating Runtime Acceptance from Offline Contract Alignment in LLM-Orchestrated Flood GIS Workflows: A Dual-Case Evaluation

Flood assessments require reproducible links between spatial products and the measurements ultimately reported to analysts, yet a large language model (LLM) agent may complete a geographic information system (GIS) workflow without generating the required product or traceable claim. We present GeoFloodAgent, an LLM-orchestrated flood-GIS workflow that combines typed planning and deterministic GIS tools with a separate evaluation of runtime disposition and post-run alignment with task-specific product, lineage, and action contracts. We evaluated 80 Chinese-language tasks from the Zhengzhou (2021) and Forlì (2023) flood cases under seven configurations and three repetitions, yielding a registry-remediated composite of 1680 assignments. In the Full configuration, 212 assignments (88.3%) were accepted and contract-aligned, whereas 16 accepted assignments did not align with their registered contracts; all 16 concerned clarification or refusal actions. Three runs were rejected despite retained contract-aligned evidence. The cluster-weighted Full-Tool Agent differences were −1.15 percentage points for accepted contract mismatch and +1.07 percentage points for contract-aligned acceptance, with empirical intervals spanning benefit and harm. Under frozen Zhengzhou inputs, deterministic outputs were sensitive to the minimum SAR object-size setting: relative to 80 pixels, candidate flood extent changed by +105.6% at 40 pixels and −30.9% at 120 pixels, whereas a 20 m simplification tolerance reproducibly failed during geometric union. These configuration-specific results do not validate hydrological accuracy or identify universal optimal defaults. These findings show that completion rate alone is inadequate for evaluating LLM-orchestrated flood GIS: runtime disposition and evidence alignment should be reported jointly. The study evaluates a finite, study-authored benchmark and does not establish hydrological accuracy, operational decision validity, or open-world error detection.

Authors

Institutions

Publication Details

Journal
ISPRS International Journal of Geo-Information
Published
2026-10-01
DOI
https://doi.org/10.3390/ijgi15100447
Primary Topic
Flood Risk Assessment and Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Separating Runtime Acceptance from Offline Contract Alignment in LLM-Orchestrated Flood GIS Workflows: A Dual-Case Evaluation

Hongyun Zhang, Fang Wang, Jin Liu, Yahong Zhao et al.
ISPRS International Journal of Geo-Information
Flood Risk Assessment and Management
article

Separating Runtime Acceptance from Offline Contract Alignment in LLM-Orchestrated Flood GIS Workflows: A Dual-Case Evaluation

Hongyun Zhang, Fang Wang, Jin Liu, Yahong Zhao, Ye Xuan
article en

Abstract

Flood assessments require reproducible links between spatial products and the measurements ultimately reported to analysts, yet a large language model (LLM) agent may complete a geographic information system (GIS) workflow without generating the required product or traceable claim. We present GeoFloodAgent, an LLM-orchestrated flood-GIS workflow that combines typed planning and deterministic GIS tools with a separate evaluation of runtime disposition and post-run alignment with task-specific product, lineage, and action contracts. We evaluated 80 Chinese-language tasks from the Zhengzhou (2021) and Forlì (2023) flood cases under seven configurations and three repetitions, yielding a registry-remediated composite of 1680 assignments. In the Full configuration, 212 assignments (88.3%) were accepted and contract-aligned, whereas 16 accepted assignments did not align with their registered contracts; all 16 concerned clarification or refusal actions. Three runs were rejected despite retained contract-aligned evidence. The cluster-weighted Full-Tool Agent differences were −1.15 percentage points for accepted contract mismatch and +1.07 percentage points for contract-aligned acceptance, with empirical intervals spanning benefit and harm. Under frozen Zhengzhou inputs, deterministic outputs were sensitive to the minimum SAR object-size setting: relative to 80 pixels, candidate flood extent changed by +105.6% at 40 pixels and −30.9% at 120 pixels, whereas a 20 m simplification tolerance reproducibly failed during geometric union. These configuration-specific results do not validate hydrological accuracy or identify universal optimal defaults. These findings show that completion rate alone is inadequate for evaluating LLM-orchestrated flood GIS: runtime disposition and evidence alignment should be reported jointly. The study evaluates a finite, study-authored benchmark and does not establish hydrological accuracy, operational decision validity, or open-world error detection.

ISPRS International Journal of Geo-InformationVol. 15(10)
Liaoning Technical University (CN), Wuhan University (CN), Institute of Disaster Prevention (CN), State Key Laboratory of Information Engineering in Surveying Mapping and Remote Sensing (CN), Department of Disaster Prevention and Mitigation (TH)
Openalex Percentile: Top 15%
Flood Risk Assessment and Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.