Conflict-aware uncertainty alignment for tool-augmented document-centric multimodal intelligence
Vision-language systems increasingly use external expert tools, such as optical character recognition (OCR) engines and document layout parsers, to improve document understanding. However, external tool outputs may be inaccurate or inconsistent with visual evidence, and conventional tool-augmented pipelines often incorporate them without explicitly modeling their reliability. This can result in error propagation, unnecessary tool invocation, and degraded performance under challenging document conditions. This paper presents a conflict-aware uncertainty alignment framework for tool-augmented document-centric multimodal intelligence. The framework integrates spatial visual uncertainty estimation, spatially aligned tool uncertainty, conflict estimation, and policy-based tool arbitration within a unified decision process. Tool confidence is converted into a semantically consistent uncertainty representation and aligned with visual uncertainty in a common spatial coordinate system. Their localized disagreement is then used to guide decisions to rely on internal visual reasoning, invoke an external tool, or reject an unreliable tool output. A reinforcement learning-based policy, optimized using Proximal Policy Optimization (PPO), jointly considers prediction correctness, tool-use cost, and correct rejection of unreliable external evidence. PPO is employed as the policy-optimization mechanism rather than as a methodological contribution in itself. Experiments on four document-understanding benchmarks, namely DocVQA, FUNSD, SROIE, and PubLayNet, demonstrate consistent improvements over representative baseline methods. The framework achieves absolute performance gains of 5.8% on DocVQA, 4.9 F1 points on FUNSD, and 4.9% on SROIE, while reducing average external tool invocation by more than 50% compared with unconditional tool-use pipelines. Robustness experiments under controlled document degradations, including held-out corruption conditions, further show improved arbitration under unreliable OCR and layout-analysis outputs. Statistical analysis and qualitative visualizations provide additional evidence that spatial conflict maps can localize disagreement between visual and tool uncertainty. Overall, the findings indicate that uncertainty-guided conflict estimation and selective tool arbitration provide a practical approach for improving the accuracy–efficiency trade-off and handling imperfect external evidence in document-centric multimodal systems.
Authors
- Kandhasamy Saritha
- Rathinasamy Muthusami (ORCID: https://orcid.org/0000-0001-7322-8653)
Institutions
- CMR University (IN)
Publication Details
- Journal
- Discover Artificial Intelligence
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1007/s44163-026-02409-3
- Primary Topic
- Handwritten Text Recognition Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00