Quality gate framework for auditing narrative rendering of evacuation simulation logs by multimodal retrieval augmented generation
Abstract Large language models (LLMs) can render agent-based simulation logs as natural-language narratives, but uncontrolled generation may omit facts, reorder events, leak internal codes, or introduce unsupported subjective content. We propose a quality-gated multimodal retrieval-augmented generation (RAG) framework that renders tsunami-evacuation simulation logs as verifiable candidate narratives. The framework separates fixed event-layer facts, context-layer labels derived from the simulation, and a variable value-layer rendering. The layered pipeline references movement logs, synthetic-population attributes, route-map images, and disaster-prevention map information, and evaluates each output through explicit quality gates. To isolate generation variability, we fixed the reference corpus, vector store, search query, and retrieved documents, and evaluated 100 iterative generations per condition. For three agents following the same route, location coverage was high (95%–99%) and sequence inconsistency was low (0%–1%), whereas timestamp coverage was comparatively low (83%–91%). Prohibited terms or internal representations appeared in 7%–23% of outputs, and the quality-gate pass rate ranged from 62–81% across the three agents, indicating that timestamp completeness and output consistency remain the main bottlenecks. Varying household composition within the separated value layer—replacing a household that includes a teenager with one that includes a member in their 90s instead, and substituting an all-elderly household—shifted lexical indicators for child-related, elderly-related, mutual-aid, and household expressions in directions consistent with the inputs and produced no clear disruption of the factual skeleton under any tested condition. This study provides a diagnostic evaluation framework and empirical methodology for auditing factual-skeleton preservation, boundary management, and failure modes in LLM/RAG-based rendering of simulation outputs.
Authors
- Fumihiro Sakahira (ORCID: https://orcid.org/0000-0002-1228-4069)
- Yuto Hirahata
Institutions
- Osaka Institute of Technology (JP)
Publication Details
- Journal
- Discover Artificial Intelligence
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1007/s44163-026-02075-5
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00