AI-Enabled Intelligent Document Processing and Multi-Class Classification for Electronic Document Management Systems Using Deep Feature Ensembles
Abstract- Modern enterprise workflows encounter an enormous influx of physical and digitized documentation across administrative, financial, legal, and operational domains. Conventional electronic document management systems (EDMS) depend on manual clerical sorting, rigid coordinate template parsing, and basic OCR engines. Consequently, such pipelines exhibit severe operational fragility, high latency, and pronounced vulnerability to scanning degradation such as rotational skew, non-uniform illumination, and bleed-through artifacts. To overcome these systemic bottlenecks, this paper presents an end-to-end, multi-layer deep learning-assisted framework engineered for automated enterprise document intelligence. The proposed architecture incorporates a Radon projection-based angular deskewing algorithm operating across a continuous [-45°, +45°] search space, coupled with adaptive local Sauvola binarization validated through Peak Signal-to-Noise Ratio (PSNR >= 28 dB) and Structural Similarity Index (SSIM >= 0.85) quality verification gates. Deep visual representations extracted via a truncated ResNet-50 backbone (2048-dimensional bottleneck embeddings) are fused with sub-word TF-IDF lexical distributions to drive a calibrated Histogram-based Gradient Boosting (HistGB) and Multi-Layer Perceptron (MLP) softvoting ensemble. Neural text transcription is executed via a Character Region Awareness for Text Detection (CRAFT) network coupled with a Bidirectional Gated Recurrent Unit (Bi-GRU) sequence decoder trained under Connectionist Temporal Classification (CTC) loss. A contextual entity extraction module leverages regular expressions and token classification to redact high-risk Personally IdentifiableInformation (PII), while full-text search is accelerated via SQLite FTS5 BM25 inverted indexing. Rigorous experimental evaluationconducted on the standard RVL-CDIP benchmark (400,000 document images across 16 categories) demonstrates that the proposed framework achieves a 98.5% classification accuracy, an 88.5% OCR word recognition rate, an 89.0% entity extraction F1-score, and an average per-page execution latency of 0.45 seconds, substantially outperforming existing enterprise pipelines.
Authors
- Dr. S. Jhansi Rani K. Sravani Reddy
Institutions
- Andhra University (IN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-08
- DOI
- https://doi.org/10.5281/zenodo.23236910
- Primary Topic
- Handwritten Text Recognition Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00