Beyond Accuracy: Evaluating Evidence Grounding in Multimodal Large Language Models for Newspaper Understanding

Although multimodal Large Language Models (LLMs) perform well on document question answering tasks, Answer Correctness, i.e., whether a model’s answer is correct with respect to the question, does not fully describe model behavior because it does not show whether answers are supported by relevant evidence. This paper presents a framework for multimodal document understanding that jointly evaluates Answer Correctness and Evidence Grounding, i.e., whether the evidence cited by the model alongside its answer is present, relevant and consistent with that answer. We evaluate two LLMs on an initial benchmark of 10 Italian newspaper front pages published on the same day, where information is spread across text, images and page layout. The benchmark combines human evaluation of model answers with assessment of the supporting evidence. On this benchmark, both models achieve high Answer Correctness, with GPT-5.5 reaching 92.9% and Claude Sonnet 5 reaching 88.8%. Most outputs are both fully correct and fully supported by the evidence (87.69% and 83.46%, respectively). However, Answer Correctness and Evidence Grounding differ in 11.15% of GPT-5.5 outputs and 15.38% of Claude Sonnet 5 outputs. Importantly, 5.00% of GPT-5.5 outputs and 6.54% of Claude Sonnet 5 outputs combine fully grounded evidence with an incorrect answer. These results on this limited benchmark suggest that the two measures capture complementary aspects of model behavior and that evaluating both can help distinguish errors in evidence grounding from errors in answer generation and interpretation.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-09-25
DOI
https://doi.org/10.3390/app16199535
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Beyond Accuracy: Evaluating Evidence Grounding in Multimodal Large Language Models for Newspaper Understanding

Roberto Zanoli, Alberto Lavelli
Applied Sciences
Topic Modeling
article

Beyond Accuracy: Evaluating Evidence Grounding in Multimodal Large Language Models for Newspaper Understanding

Roberto Zanoli, Alberto Lavelli
article en

Abstract

Although multimodal Large Language Models (LLMs) perform well on document question answering tasks, Answer Correctness, i.e., whether a model’s answer is correct with respect to the question, does not fully describe model behavior because it does not show whether answers are supported by relevant evidence. This paper presents a framework for multimodal document understanding that jointly evaluates Answer Correctness and Evidence Grounding, i.e., whether the evidence cited by the model alongside its answer is present, relevant and consistent with that answer. We evaluate two LLMs on an initial benchmark of 10 Italian newspaper front pages published on the same day, where information is spread across text, images and page layout. The benchmark combines human evaluation of model answers with assessment of the supporting evidence. On this benchmark, both models achieve high Answer Correctness, with GPT-5.5 reaching 92.9% and Claude Sonnet 5 reaching 88.8%. Most outputs are both fully correct and fully supported by the evidence (87.69% and 83.46%, respectively). However, Answer Correctness and Evidence Grounding differ in 11.15% of GPT-5.5 outputs and 15.38% of Claude Sonnet 5 outputs. Importantly, 5.00% of GPT-5.5 outputs and 6.54% of Claude Sonnet 5 outputs combine fully grounded evidence with an incorrect answer. These results on this limited benchmark suggest that the two measures capture complementary aspects of model behavior and that evaluating both can help distinguish errors in evidence grounding from errors in answer generation and interpretation.

Applied SciencesVol. 16(19)
Fondazione Bruno Kessler (IT)
Quality Education
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.