Uncertainty-Aware Multimodal AI: Detecting Ambiguity Instead of Guessing in Vision-Language Systems

Modern vision-language models (VLMs) can generate image captions, answer visual ques-tions, and follow natural-language instructions with high fluency, but they are typically trainedto always produce an answer, even when the input is genuinely ambiguous or the model itselfis uncertain. This can produce confident-sounding but incorrect outputs, which is especiallyproblematic in accessibility applications and human-robot interaction, where a wrong guess car-ries real cost. In this work, we present a prototype multimodal AI pipeline that detects itsown uncertainty by asking the same question twice with reworded phrasing and comparing thetwo answers for consistency; when the answers disagree, the system asks a clarifying questioninstead of committing to a guess. We evaluate this approach on an image captioning and visualquestion answering (VQA) pipeline with spoken output, and separately extend the same idea toa simulated robot instruction-following task. We report honest, mixed results: the method suc-cessfully flags several genuinely ambiguous cases, but we also identify two concrete limitations –oversensitivity of simple heuristics, and run-to-run inconsistency caused by inherent randomnessin language model sampling – that we argue are important open problems for uncertainty-awaremultimodal systems rather than issues to conceal

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-14
DOI
https://doi.org/10.5281/zenodo.22750880
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Uncertainty-Aware Multimodal AI: Detecting Ambiguity Instead of Guessing in Vision-Language Systems

Varshini R
Zenodo (CERN European Organization for Nuclear Research)
Multimodal Machine Learning Applications
article

Uncertainty-Aware Multimodal AI: Detecting Ambiguity Instead of Guessing in Vision-Language Systems

Varshini R
article en

Abstract

Modern vision-language models (VLMs) can generate image captions, answer visual ques-tions, and follow natural-language instructions with high fluency, but they are typically trainedto always produce an answer, even when the input is genuinely ambiguous or the model itselfis uncertain. This can produce confident-sounding but incorrect outputs, which is especiallyproblematic in accessibility applications and human-robot interaction, where a wrong guess car-ries real cost. In this work, we present a prototype multimodal AI pipeline that detects itsown uncertainty by asking the same question twice with reworded phrasing and comparing thetwo answers for consistency; when the answers disagree, the system asks a clarifying questioninstead of committing to a guess. We evaluate this approach on an image captioning and visualquestion answering (VQA) pipeline with spoken output, and separately extend the same idea toa simulated robot instruction-following task. We report honest, mixed results: the method suc-cessfully flags several genuinely ambiguous cases, but we also identify two concrete limitations –oversensitivity of simple heuristics, and run-to-run inconsistency caused by inherent randomnessin language model sampling – that we argue are important open problems for uncertainty-awaremultimodal systems rather than issues to conceal

Zenodo (CERN European Organization for Nuclear Research)
Lynn University (US)
Quality Education
Openalex Percentile: Top 13%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Uncertainty-Aware Multimodal AI: Detecting Ambiguity Instead of Guessing in Vision-Language Systems — Varshini R · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS