AI-Enabled Intelligent Document Processing and Multi-Class Classification for Electronic Document Management Systems Using Deep Feature Ensembles

Abstract- Modern enterprise workflows encounter an enormous influx of physical and digitized documentation across administrative, financial, legal, and operational domains. Conventional electronic document management systems (EDMS) depend on manual clerical sorting, rigid coordinate template parsing, and basic OCR engines. Consequently, such pipelines exhibit severe operational fragility, high latency, and pronounced vulnerability to scanning degradation such as rotational skew, non-uniform illumination, and bleed-through artifacts. To overcome these systemic bottlenecks, this paper presents an end-to-end, multi-layer deep learning-assisted framework engineered for automated enterprise document intelligence. The proposed architecture incorporates a Radon projection-based angular deskewing algorithm operating across a continuous [-45°, +45°] search space, coupled with adaptive local Sauvola binarization validated through Peak Signal-to-Noise Ratio (PSNR >= 28 dB) and Structural Similarity Index (SSIM >= 0.85) quality verification gates. Deep visual representations extracted via a truncated ResNet-50 backbone (2048-dimensional bottleneck embeddings) are fused with sub-word TF-IDF lexical distributions to drive a calibrated Histogram-based Gradient Boosting (HistGB) and Multi-Layer Perceptron (MLP) softvoting ensemble. Neural text transcription is executed via a Character Region Awareness for Text Detection (CRAFT) network coupled with a Bidirectional Gated Recurrent Unit (Bi-GRU) sequence decoder trained under Connectionist Temporal Classification (CTC) loss. A contextual entity extraction module leverages regular expressions and token classification to redact high-risk Personally IdentifiableInformation (PII), while full-text search is accelerated via SQLite FTS5 BM25 inverted indexing. Rigorous experimental evaluationconducted on the standard RVL-CDIP benchmark (400,000 document images across 16 categories) demonstrates that the proposed framework achieves a 98.5% classification accuracy, an 88.5% OCR word recognition rate, an 89.0% entity extraction F1-score, and an average per-page execution latency of 0.45 seconds, substantially outperforming existing enterprise pipelines.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-08
DOI
https://doi.org/10.5281/zenodo.23236910
Primary Topic
Handwritten Text Recognition Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

AI-Enabled Intelligent Document Processing and Multi-Class Classification for Electronic Document Management Systems Using Deep Feature Ensembles

Dr. S. Jhansi Rani K. Sravani Reddy
Zenodo (CERN European Organization for Nuclear Research)
Handwritten Text Recognition Techniques
article

AI-Enabled Intelligent Document Processing and Multi-Class Classification for Electronic Document Management Systems Using Deep Feature Ensembles

Dr. S. Jhansi Rani K. Sravani Reddy
article en

Abstract

Abstract- Modern enterprise workflows encounter an enormous influx of physical and digitized documentation across administrative, financial, legal, and operational domains. Conventional electronic document management systems (EDMS) depend on manual clerical sorting, rigid coordinate template parsing, and basic OCR engines. Consequently, such pipelines exhibit severe operational fragility, high latency, and pronounced vulnerability to scanning degradation such as rotational skew, non-uniform illumination, and bleed-through artifacts. To overcome these systemic bottlenecks, this paper presents an end-to-end, multi-layer deep learning-assisted framework engineered for automated enterprise document intelligence. The proposed architecture incorporates a Radon projection-based angular deskewing algorithm operating across a continuous [-45°, +45°] search space, coupled with adaptive local Sauvola binarization validated through Peak Signal-to-Noise Ratio (PSNR >= 28 dB) and Structural Similarity Index (SSIM >= 0.85) quality verification gates. Deep visual representations extracted via a truncated ResNet-50 backbone (2048-dimensional bottleneck embeddings) are fused with sub-word TF-IDF lexical distributions to drive a calibrated Histogram-based Gradient Boosting (HistGB) and Multi-Layer Perceptron (MLP) softvoting ensemble. Neural text transcription is executed via a Character Region Awareness for Text Detection (CRAFT) network coupled with a Bidirectional Gated Recurrent Unit (Bi-GRU) sequence decoder trained under Connectionist Temporal Classification (CTC) loss. A contextual entity extraction module leverages regular expressions and token classification to redact high-risk Personally IdentifiableInformation (PII), while full-text search is accelerated via SQLite FTS5 BM25 inverted indexing. Rigorous experimental evaluationconducted on the standard RVL-CDIP benchmark (400,000 document images across 16 categories) demonstrates that the proposed framework achieves a 98.5% classification accuracy, an 88.5% OCR word recognition rate, an 89.0% entity extraction F1-score, and an average per-page execution latency of 0.45 seconds, substantially outperforming existing enterprise pipelines.

Zenodo (CERN European Organization for Nuclear Research)
Andhra University (IN)
Openalex Percentile: Top 15%
Handwritten Text Recognition Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.