Multimodal LLM Pipeline for Automated Analysis of Technical Graphical Documents
Automating the evaluation of technical documents, such as laboratory reports, requires balancing accurate multimodal information extraction with efficient downstream processing. This paper proposes a vision–language model (VLM) processing pipeline that extracts textual and graphical content into structured intermediate representations (JSON), enabling subsequent document assessment without redundant image retransmission. We evaluate the approach across a corpus of laboratory reports, investigating the impact of prompting strategies on text extraction and document analysis. The evaluation reveals distinct performance trade-offs across prompting configurations: Prompt 10 with Llama 4 Maverick achieved the highest textual overlap (Dice = 0.7402), whereas Prompt 5 yielded the highest observed TF-IDF cosine similarity and normalized Levenshtein similarity scores. Furthermore, downstream validation using structured text representations provides a viable architectural approach to avoid repeated visual-input transmission in multi-stage analysis workflows. These findings demonstrate the feasibility of structured intermediate representations for technical document parsing, while highlighting how specific prompt designs trade off different extraction quality metrics.
Authors
- Mariusz Maczka (ORCID: https://orcid.org/0000-0003-4137-4829)
- Szymon Studniarz
Institutions
- Rzeszów University of Technology (PL)
Publication Details
- Journal
- AI
- Published
- 2026-10-04
- DOI
- https://doi.org/10.3390/ai7100403
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00