Multimodal LLM Pipeline for Automated Analysis of Technical Graphical Documents

Automating the evaluation of technical documents, such as laboratory reports, requires balancing accurate multimodal information extraction with efficient downstream processing. This paper proposes a vision–language model (VLM) processing pipeline that extracts textual and graphical content into structured intermediate representations (JSON), enabling subsequent document assessment without redundant image retransmission. We evaluate the approach across a corpus of laboratory reports, investigating the impact of prompting strategies on text extraction and document analysis. The evaluation reveals distinct performance trade-offs across prompting configurations: Prompt 10 with Llama 4 Maverick achieved the highest textual overlap (Dice = 0.7402), whereas Prompt 5 yielded the highest observed TF-IDF cosine similarity and normalized Levenshtein similarity scores. Furthermore, downstream validation using structured text representations provides a viable architectural approach to avoid repeated visual-input transmission in multi-stage analysis workflows. These findings demonstrate the feasibility of structured intermediate representations for technical document parsing, while highlighting how specific prompt designs trade off different extraction quality metrics.

Authors

Institutions

Publication Details

Journal
AI
Published
2026-10-04
DOI
https://doi.org/10.3390/ai7100403
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Multimodal LLM Pipeline for Automated Analysis of Technical Graphical Documents

Mariusz Maczka, Szymon Studniarz
AI
Multimodal Machine Learning Applications
article

Multimodal LLM Pipeline for Automated Analysis of Technical Graphical Documents

Mariusz Maczka, Szymon Studniarz
article en

Abstract

Automating the evaluation of technical documents, such as laboratory reports, requires balancing accurate multimodal information extraction with efficient downstream processing. This paper proposes a vision–language model (VLM) processing pipeline that extracts textual and graphical content into structured intermediate representations (JSON), enabling subsequent document assessment without redundant image retransmission. We evaluate the approach across a corpus of laboratory reports, investigating the impact of prompting strategies on text extraction and document analysis. The evaluation reveals distinct performance trade-offs across prompting configurations: Prompt 10 with Llama 4 Maverick achieved the highest textual overlap (Dice = 0.7402), whereas Prompt 5 yielded the highest observed TF-IDF cosine similarity and normalized Levenshtein similarity scores. Furthermore, downstream validation using structured text representations provides a viable architectural approach to avoid repeated visual-input transmission in multi-stage analysis workflows. These findings demonstrate the feasibility of structured intermediate representations for technical document parsing, while highlighting how specific prompt designs trade off different extraction quality metrics.

AIVol. 7(10)
Rzeszów University of Technology (PL)
Openalex Percentile: Top 14%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.