Post-Hoc Conformal Prediction for Hallucination Detection in Vision-Language Models: A Distribution-Free Coverage Guarantee on POPE-Adversarial

Vision-Language Models (VLMs) frequently produce confident but factually incorrect answers about image content, a failure mode commonly termed hallucination. Standard softmax confidence scores are not calibrated: a model can assign high probability to a wrong answer with no statistical guarantee that this probability reflects the true likelihood of correctness. We apply split-conformal prediction as a post-hoc, model-agnostic wrapper around a frozen LLaVA-1.5-7B model, converting its raw output into a prediction set ({Yes}, {No}, or {Yes, No}) with a distribution-free marginal coverage guarantee: the true label falls inside the set with probability at least 1 − α, for a user-chosen α, under no assumption on the model or data distribution beyond exchangeability. We evaluate on the POPE-Adversarial benchmark (COCO val2014), the hardest of the three standard POPE splits, using 1,000 calibration and 2,000 test samples at a target coverage of 90% (α = 0.10). The calibrated system achieves 91.45% empirical coverage, consistent with the theoretical guarantee, at an average prediction-set size of 1.195 labels: 80.5% of test questions receive a confident singleton set, while 19.5% are flagged as uncertain (both labels retained), directly identifying the subset of answers most likely to be unreliable. Raw point-prediction accuracy on the same data is 83.35%, meaning conformal wrapping does not improve raw accuracy but instead provides calibrated, statistically-grounded uncertainty on top of it. We report the full pipeline, a theoretical sanity check of the observed coverage against its expected sampling distribution, and the limitations of a single-split, single-model evaluation. Code, results and figures: https://github.com/sreejatamaity/Conformal-vlm-hallucination This is a preprint and has not been peer reviewed. AI assistance: Generative AI (Claude, Anthropic) was used for drafting and editing certain sections of the manuscript text and for code debugging. The author conducted all experiments, independently verified all reported results against the raw output files, and takes responsibility for the accuracy and content of the work.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-06
DOI
https://doi.org/10.5281/zenodo.23180933
Primary Topic
Multimodal Machine Learning Applications
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Post-Hoc Conformal Prediction for Hallucination Detection in Vision-Language Models: A Distribution-Free Coverage Guarantee on POPE-Adversarial

Sreejata Maity
Zenodo (CERN European Organization for Nuclear Research)
Multimodal Machine Learning Applications
preprint

Post-Hoc Conformal Prediction for Hallucination Detection in Vision-Language Models: A Distribution-Free Coverage Guarantee on POPE-Adversarial

Sreejata Maity
preprint en

Abstract

Vision-Language Models (VLMs) frequently produce confident but factually incorrect answers about image content, a failure mode commonly termed hallucination. Standard softmax confidence scores are not calibrated: a model can assign high probability to a wrong answer with no statistical guarantee that this probability reflects the true likelihood of correctness. We apply split-conformal prediction as a post-hoc, model-agnostic wrapper around a frozen LLaVA-1.5-7B model, converting its raw output into a prediction set ({Yes}, {No}, or {Yes, No}) with a distribution-free marginal coverage guarantee: the true label falls inside the set with probability at least 1 − α, for a user-chosen α, under no assumption on the model or data distribution beyond exchangeability. We evaluate on the POPE-Adversarial benchmark (COCO val2014), the hardest of the three standard POPE splits, using 1,000 calibration and 2,000 test samples at a target coverage of 90% (α = 0.10). The calibrated system achieves 91.45% empirical coverage, consistent with the theoretical guarantee, at an average prediction-set size of 1.195 labels: 80.5% of test questions receive a confident singleton set, while 19.5% are flagged as uncertain (both labels retained), directly identifying the subset of answers most likely to be unreliable. Raw point-prediction accuracy on the same data is 83.35%, meaning conformal wrapping does not improve raw accuracy but instead provides calibrated, statistically-grounded uncertainty on top of it. We report the full pipeline, a theoretical sanity check of the observed coverage against its expected sampling distribution, and the limitations of a single-split, single-model evaluation. Code, results and figures: https://github.com/sreejatamaity/Conformal-vlm-hallucination This is a preprint and has not been peer reviewed. AI assistance: Generative AI (Claude, Anthropic) was used for drafting and editing certain sections of the manuscript text and for code debugging. The author conducted all experiments, independently verified all reported results against the raw output files, and takes responsibility for the accuracy and content of the work.

Zenodo (CERN European Organization for Nuclear Research)
PES University (IN)
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Post-Hoc Conformal Prediction for Hallucination Detection in Vision-Language Models: A Distribution-Free Coverage Guarantee on POPE-Adversarial — Sreejata Maity · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS