Uncertainty-Aware Multimodal AI: Detecting Ambiguity Instead of Guessing in Vision-Language Systems
Modern vision-language models (VLMs) can generate image captions, answer visual ques-tions, and follow natural-language instructions with high fluency, but they are typically trainedto always produce an answer, even when the input is genuinely ambiguous or the model itselfis uncertain. This can produce confident-sounding but incorrect outputs, which is especiallyproblematic in accessibility applications and human-robot interaction, where a wrong guess car-ries real cost. In this work, we present a prototype multimodal AI pipeline that detects itsown uncertainty by asking the same question twice with reworded phrasing and comparing thetwo answers for consistency; when the answers disagree, the system asks a clarifying questioninstead of committing to a guess. We evaluate this approach on an image captioning and visualquestion answering (VQA) pipeline with spoken output, and separately extend the same idea toa simulated robot instruction-following task. We report honest, mixed results: the method suc-cessfully flags several genuinely ambiguous cases, but we also identify two concrete limitations –oversensitivity of simple heuristics, and run-to-run inconsistency caused by inherent randomnessin language model sampling – that we argue are important open problems for uncertainty-awaremultimodal systems rather than issues to conceal
Authors
- Varshini R
Institutions
- Lynn University (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-14
- DOI
- https://doi.org/10.5281/zenodo.22750881
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00