Character-Level Visual Guidance for Arbitrarily Shaped Scene Text Detection and Recognition

Curved, blurred, low-contrast, and cluttered scene text remains difficult to localize and transcribe because visual boundaries and character identities are often ambiguous. Most detectors refine spatial regions from visual and positional features, while many recognizers depend on implicit attention between global image tokens and language context. This separation weakens character-level evidence in both localization and text decoding. To address this issue, this paper studies explicit character-level guidance for scene text detection and recognition. For arbitrary-shaped text detection, the character-guided adaptive detector (CADet) combines a text-enhancement network (TENet), a character information adaptive guidance module (CIA), and a position and classification compensation module (COMP). Then, character semantics can participate in boundary-query construction before final detection. For scene text recognition (STR), serialized image embeddings for text recognition (SIETR) introduce local visual embeddings into autoregressive decoding via the character local image embedding module (CLIE) and permutation language modeling (PLM). On ArT, Total-Text, and CTW1500, CADet obtains F-measures of 79.5%, 89.4%, and 89.2%, respectively, outperforming representative Transformer-based detectors with only a small increase in computation. For recognition, SIETR reaches 95.6% sample-size-weighted average accuracy with 23.8 M parameters and 3.2 G FLOPs and improves most irregular-text benchmarks over PARSeq with fewer FLOPs. The results indicate that character semantics for detection and local visual evidence for decoding are an effective way to improve robustness in difficult scene text.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-09-11
DOI
https://doi.org/10.3390/electronics15184120
Primary Topic
Handwritten Text Recognition Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Character-Level Visual Guidance for Arbitrarily Shaped Scene Text Detection and Recognition

Lijia Chen, Dan Chen, Hu Lin
Electronics
Handwritten Text Recognition Techniques
article

Character-Level Visual Guidance for Arbitrarily Shaped Scene Text Detection and Recognition

Lijia Chen, Dan Chen, Hu Lin
article en

Abstract

Curved, blurred, low-contrast, and cluttered scene text remains difficult to localize and transcribe because visual boundaries and character identities are often ambiguous. Most detectors refine spatial regions from visual and positional features, while many recognizers depend on implicit attention between global image tokens and language context. This separation weakens character-level evidence in both localization and text decoding. To address this issue, this paper studies explicit character-level guidance for scene text detection and recognition. For arbitrary-shaped text detection, the character-guided adaptive detector (CADet) combines a text-enhancement network (TENet), a character information adaptive guidance module (CIA), and a position and classification compensation module (COMP). Then, character semantics can participate in boundary-query construction before final detection. For scene text recognition (STR), serialized image embeddings for text recognition (SIETR) introduce local visual embeddings into autoregressive decoding via the character local image embedding module (CLIE) and permutation language modeling (PLM). On ArT, Total-Text, and CTW1500, CADet obtains F-measures of 79.5%, 89.4%, and 89.2%, respectively, outperforming representative Transformer-based detectors with only a small increase in computation. For recognition, SIETR reaches 95.6% sample-size-weighted average accuracy with 23.8 M parameters and 3.2 G FLOPs and improves most irregular-text benchmarks over PARSeq with fewer FLOPs. The results indicate that character semantics for detection and local visual evidence for decoding are an effective way to improve robustness in difficult scene text.

ElectronicsVol. 15(18)
Fujian Business University (CN), Fuzhou University (CN)
Peace, Justice and strong institutions
Openalex Percentile: Top 13%
Handwritten Text Recognition Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.