A Multilevel Visual and Textual Framework for Near-Duplicate Diagram Detection in Electronic Documents

Near-duplicate diagram detection in electronic documents is challenging because diagram identity depends on graphical structure, spatial composition, and textual labels, while reused images may undergo compression, cropping, rotation, photometric changes, or perspective distortion. This study proposes a cascaded multimodal framework combining perceptual hashing, Siamese Residual Network with 18 layers (Siamese ResNet18), Distillation with No Labels, ver. 2 (DINOv2)visual representations, and a text-similarity classifier. A controlled benchmark was constructed from Artificial Intelligence 2D Diagram Dataset (AI2D) using Light, Medium, and Hard transformations, with source-grouped splitting by base_id to prevent leakage across training, validation, and test sets. Perceptual hashing achieved Area Under the Receiver Operating Characteristic Curve (ROC-AUC) = 0.7483, while Siamese ResNet18 increased ROC-AUC to 0.8599. DINOv2 provided the strongest visual performance, achieving Accuracy = 0.9933, F1-score = 0.9933, ROC-AUC = 0.9992, and Average Precision = 0.9994; F1-score remained 0.9901 for Hard transformations. Visual fusion increased ROC-AUC to 0.9995, and full multimodal fusion reached ROC-AUC = 0.9998. At an early-exit threshold of 0.95, 27.6% of pairs were resolved at the hashing level. These results support the coarse-to-fine design on the constructed AI2D-derived benchmark. A targeted hard-negative stress test revealed substantially higher false-positive rates under deliberately matched spatial layouts, with an overall False Positive Rate (FPR) of 0.48 for DINOv2 and 0.16 for full multimodal fusion. Generalization to naturally reused or redrawn diagrams, larger and more diverse hard-negative collections, and Optical Character Recognition (OCR)-derived text remains to be evaluated.

Authors

Institutions

Publication Details

Journal
Information
Published
2026-09-15
DOI
https://doi.org/10.3390/info17090897
Primary Topic
Handwritten Text Recognition Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Multilevel Visual and Textual Framework for Near-Duplicate Diagram Detection in Electronic Documents

Yurii Andrashko, Oleksandr Kuchanskyi, Zhan Amangeldiyev, Dina Kantayeva et al.
Information
Handwritten Text Recognition Techniques
article

A Multilevel Visual and Textual Framework for Near-Duplicate Diagram Detection in Electronic Documents

Yurii Andrashko, Oleksandr Kuchanskyi, Zhan Amangeldiyev, Dina Kantayeva, Svitlana Biloshchytska, Myroslava Tovt-Kuchanska
article en

Abstract

Near-duplicate diagram detection in electronic documents is challenging because diagram identity depends on graphical structure, spatial composition, and textual labels, while reused images may undergo compression, cropping, rotation, photometric changes, or perspective distortion. This study proposes a cascaded multimodal framework combining perceptual hashing, Siamese Residual Network with 18 layers (Siamese ResNet18), Distillation with No Labels, ver. 2 (DINOv2)visual representations, and a text-similarity classifier. A controlled benchmark was constructed from Artificial Intelligence 2D Diagram Dataset (AI2D) using Light, Medium, and Hard transformations, with source-grouped splitting by base_id to prevent leakage across training, validation, and test sets. Perceptual hashing achieved Area Under the Receiver Operating Characteristic Curve (ROC-AUC) = 0.7483, while Siamese ResNet18 increased ROC-AUC to 0.8599. DINOv2 provided the strongest visual performance, achieving Accuracy = 0.9933, F1-score = 0.9933, ROC-AUC = 0.9992, and Average Precision = 0.9994; F1-score remained 0.9901 for Hard transformations. Visual fusion increased ROC-AUC to 0.9995, and full multimodal fusion reached ROC-AUC = 0.9998. At an early-exit threshold of 0.95, 27.6% of pairs were resolved at the hashing level. These results support the coarse-to-fine design on the constructed AI2D-derived benchmark. A targeted hard-negative stress test revealed substantially higher false-positive rates under deliberately matched spatial layouts, with an overall False Positive Rate (FPR) of 0.48 for DINOv2 and 0.16 for full multimodal fusion. Generalization to naturally reused or redrawn diagrams, larger and more diverse hard-negative collections, and Optical Character Recognition (OCR)-derived text remains to be evaluated.

InformationVol. 17(9)
Kyiv National University of Construction and Architecture (UA), National Technical University of Ukraine “Igor Sikorsky Kyiv Polytechnic Institute” (UA), Astana Medical University (KZ), Uzhhorod National University (UA), Bogomolets National Medical University (UA)
Openalex Percentile: Top 13%
Handwritten Text Recognition Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.