Hierarchical Cascaded Tokenization and Dynamic Language State Tracking
This report presents a hardware-aligned tokenization framework that prevents vocabulary explosion without relying on cross-lingual semantic pivots. The framework partitions the token index space into four bounded regions: control identifiers, primitive bytes, surface lexicons, and a universal grapheme pool. Input text is processed through a deterministic three-stage fallback: localized trie prefix matching, canonical graphemic decomposition, and primitive byte mapping. This structure guarantees zero out-of-vocabulary conditions within a compact, hardware-aligned index space. During inference, a finite-state decoder tracks script transitions dynamically to prevent cross-orthographic drift.
Authors
- Young Jin Oh (ORCID: https://orcid.org/0000-0001-7335-0149)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-15
- DOI
- https://doi.org/10.5281/zenodo.22771283
- Primary Topic
- Natural Language Processing Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00