Quality gate framework for auditing narrative rendering of evacuation simulation logs by multimodal retrieval augmented generation

Abstract Large language models (LLMs) can render agent-based simulation logs as natural-language narratives, but uncontrolled generation may omit facts, reorder events, leak internal codes, or introduce unsupported subjective content. We propose a quality-gated multimodal retrieval-augmented generation (RAG) framework that renders tsunami-evacuation simulation logs as verifiable candidate narratives. The framework separates fixed event-layer facts, context-layer labels derived from the simulation, and a variable value-layer rendering. The layered pipeline references movement logs, synthetic-population attributes, route-map images, and disaster-prevention map information, and evaluates each output through explicit quality gates. To isolate generation variability, we fixed the reference corpus, vector store, search query, and retrieved documents, and evaluated 100 iterative generations per condition. For three agents following the same route, location coverage was high (95%–99%) and sequence inconsistency was low (0%–1%), whereas timestamp coverage was comparatively low (83%–91%). Prohibited terms or internal representations appeared in 7%–23% of outputs, and the quality-gate pass rate ranged from 62–81% across the three agents, indicating that timestamp completeness and output consistency remain the main bottlenecks. Varying household composition within the separated value layer—replacing a household that includes a teenager with one that includes a member in their 90s instead, and substituting an all-elderly household—shifted lexical indicators for child-related, elderly-related, mutual-aid, and household expressions in directions consistent with the inputs and produced no clear disruption of the factual skeleton under any tested condition. This study provides a diagnostic evaluation framework and empirical methodology for auditing factual-skeleton preservation, boundary management, and failure modes in LLM/RAG-based rendering of simulation outputs.

Authors

Institutions

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-10-06
DOI
https://doi.org/10.1007/s44163-026-02075-5
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Quality gate framework for auditing narrative rendering of evacuation simulation logs by multimodal retrieval augmented generation

Fumihiro Sakahira, Yuto Hirahata
Discover Artificial Intelligence
Topic Modeling
article

Quality gate framework for auditing narrative rendering of evacuation simulation logs by multimodal retrieval augmented generation

Fumihiro Sakahira, Yuto Hirahata
article en

Abstract

Abstract Large language models (LLMs) can render agent-based simulation logs as natural-language narratives, but uncontrolled generation may omit facts, reorder events, leak internal codes, or introduce unsupported subjective content. We propose a quality-gated multimodal retrieval-augmented generation (RAG) framework that renders tsunami-evacuation simulation logs as verifiable candidate narratives. The framework separates fixed event-layer facts, context-layer labels derived from the simulation, and a variable value-layer rendering. The layered pipeline references movement logs, synthetic-population attributes, route-map images, and disaster-prevention map information, and evaluates each output through explicit quality gates. To isolate generation variability, we fixed the reference corpus, vector store, search query, and retrieved documents, and evaluated 100 iterative generations per condition. For three agents following the same route, location coverage was high (95%–99%) and sequence inconsistency was low (0%–1%), whereas timestamp coverage was comparatively low (83%–91%). Prohibited terms or internal representations appeared in 7%–23% of outputs, and the quality-gate pass rate ranged from 62–81% across the three agents, indicating that timestamp completeness and output consistency remain the main bottlenecks. Varying household composition within the separated value layer—replacing a household that includes a teenager with one that includes a member in their 90s instead, and substituting an all-elderly household—shifted lexical indicators for child-related, elderly-related, mutual-aid, and household expressions in directions consistent with the inputs and produced no clear disruption of the factual skeleton under any tested condition. This study provides a diagnostic evaluation framework and empirical methodology for auditing factual-skeleton preservation, boundary management, and failure modes in LLM/RAG-based rendering of simulation outputs.

Discover Artificial IntelligenceVol. 6(1)
Osaka Institute of Technology (JP)
Openalex Percentile: Top 11%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Quality gate framework for auditing narrative rendering of evacuation simulation logs by multimodal retrieval augmented generation — Fumihiro Sakahira, Yuto Hirahata · Discover Artificial Intelligence (2026) | TGRS Research Map | TGRS