PKSF: A Task-Aware Prior Knowledge Selection and Fusion Framework for TextVQA
Text-based visual question answering (TextVQA) requires reasoning over images containing rich textual content, often involving knowledge beyond what is directly observable. Existing methods fuse visual objects and OCR tokens but struggle when questions require external knowledge. Moreover, naively incorporating retrieved knowledge often introduces irrelevant or misleading information, which may hinder reasoning rather than support it. To address these challenges, we propose a TextVQA framework that integrates external prior knowledge to support multimodal reasoning. Given an image and question, a task-aware knowledge retrieval module selects relevant candidates, which are then filtered and verified by a knowledge verification module leveraging large language models. The verified knowledge and question are compressed into compact embeddings via a perceiver-based semantic resampler and jointly processed with visual and OCR features in a multimodal reasoning module. Experiments on the TextVQA and ST-VQA datasets demonstrate that our approach effectively leverages external knowledge to improve performance on knowledge-intensive questions.
Authors
- Shuangjiao Zhai (ORCID: https://orcid.org/0000-0002-4934-3640)
- Jia Qin (ORCID: https://orcid.org/0009-0001-6216-7602)
- Jianchao Zeng (ORCID: https://orcid.org/0000-0001-7755-7550)
- Suzhen Lin (ORCID: https://orcid.org/0000-0001-8644-454X)
- Yanxia Jin (ORCID: https://orcid.org/0009-0009-2542-388X)
- Zanxia Jin
- Pinle Qin
Institutions
- North University of China (CN)
Publication Details
- Journal
- Electronics
- Published
- 2026-09-14
- DOI
- https://doi.org/10.3390/electronics15184168
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Natural Science Foundation of China