Beyond Accuracy: Evaluating Evidence Grounding in Multimodal Large Language Models for Newspaper Understanding
Although multimodal Large Language Models (LLMs) perform well on document question answering tasks, Answer Correctness, i.e., whether a model’s answer is correct with respect to the question, does not fully describe model behavior because it does not show whether answers are supported by relevant evidence. This paper presents a framework for multimodal document understanding that jointly evaluates Answer Correctness and Evidence Grounding, i.e., whether the evidence cited by the model alongside its answer is present, relevant and consistent with that answer. We evaluate two LLMs on an initial benchmark of 10 Italian newspaper front pages published on the same day, where information is spread across text, images and page layout. The benchmark combines human evaluation of model answers with assessment of the supporting evidence. On this benchmark, both models achieve high Answer Correctness, with GPT-5.5 reaching 92.9% and Claude Sonnet 5 reaching 88.8%. Most outputs are both fully correct and fully supported by the evidence (87.69% and 83.46%, respectively). However, Answer Correctness and Evidence Grounding differ in 11.15% of GPT-5.5 outputs and 15.38% of Claude Sonnet 5 outputs. Importantly, 5.00% of GPT-5.5 outputs and 6.54% of Claude Sonnet 5 outputs combine fully grounded evidence with an incorrect answer. These results on this limited benchmark suggest that the two measures capture complementary aspects of model behavior and that evaluating both can help distinguish errors in evidence grounding from errors in answer generation and interpretation.
Authors
- Roberto Zanoli (ORCID: https://orcid.org/0000-0003-0870-0872)
- Alberto Lavelli (ORCID: https://orcid.org/0000-0002-7175-6804)
Institutions
- Fondazione Bruno Kessler (IT)
Publication Details
- Journal
- Applied Sciences
- Published
- 2026-09-25
- DOI
- https://doi.org/10.3390/app16199535
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00