A Unified Neural Framework for Punctuation and Capitalization Restoration Using XLM-RoBERTa–BiLSTM
Accurate punctuation and capitalization are essential for the readability, interpretability, and structural coherence of machine-generated text. Their absence is particularly problematic in automatic speech recognition outputs and other forms of unstructured text, where missing punctuation and incorrect capitalization reduce both human readability and the effectiveness of downstream natural language processing tasks. This study proposes a hybrid XLM-RoBERTa–BiLSTM model for joint punctuation restoration and text capitalization on English-language data. The proposed architecture combines contextual representations produced by the multilingual pre-trained XLM-RoBERTa encoder with the sequential modeling capabilities of a bidirectional long short-term memory layer, followed by token-level classification in a unified label space. The model was trained and evaluated on a dataset derived from the IWSLT 2012 TED Talks corpus. Experimental evaluation on a held-out random test subset demonstrates strong performance. Excluding the dominant no-punctuation class, the model achieves an accuracy of 0.929, precision of 0.892, recall of 0.914, and an F1-score of 0.903. Including the dominant no-punctuation class increases these values to 0.961, 0.927, 0.919, and 0.923, respectively, reflecting the pronounced class imbalance in the dataset. Class-wise analysis shows high effectiveness for frequent punctuation classes and reliable capitalization prediction, whereas rare punctuation–capitalization categories remain more challenging because of their limited representation in the training and test subsets. Overall, the proposed hybrid XLM-RoBERTa–BiLSTM architecture achieves strong performance in joint punctuation restoration and text capitalization and represents an effective approach to improving the readability and structural quality of automatic speech recognition transcripts and other machine-generated text.
Authors
- Volodymyr Shymkovych (ORCID: https://orcid.org/0000-0003-4014-2786)
- Grzegorz Nowakowski (ORCID: https://orcid.org/0000-0002-3086-0947)
- Sergii Telenyk (ORCID: https://orcid.org/0000-0001-9202-9406)
- Artem Kramov
Institutions
- Taras Shevchenko National University of Kyiv (UA)
- National Technical University of Ukraine “Igor Sikorsky Kyiv Polytechnic Institute” (UA)
- Cracow University of Technology (PL)
Publication Details
- Journal
- Applied Sciences
- Published
- 2026-09-16
- DOI
- https://doi.org/10.3390/app16189176
- Primary Topic
- Natural Language Processing Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00