Bosnian CORE NLP Standard (BCS-compatible): Text Normalization & Tokenization — v1.1-LTS

**Bosnian CORE NLP Standard (BCS-compatible) v1.1-LTS** is a deterministic and audit-ready specification for the normalization, sentence segmentation, tokenization, and quantitative measurement of Bosnian Latin-script text. The release combines a formal specification with an executable reference implementation (`bcscore`) written in pure Python ≥ 3.9 and a byte-exact conformance suite containing 118 test cases. Version 1.1-LTS supersedes v1.0-LTS and resolves a set of specification ambiguities and inconsistencies, including handling of ZWJ characters, dashes, abbreviation lists, case folding, segmentation order, and n-gram reset behavior. The release introduces explicit character-stream profiles (`CHAR-LS`, `CHAR-L`, `CHAR-NWS`, `CHAR-FULL`) and machine-readable measurement profiles intended to make quantitative NLP and information-theoretic results reproducible and comparable. Annex M defines the supported quantitative measures, including Shannon entropy, Miller–Madow correction, Onicescu energy, Rényi entropy, Gini–Simpson index, HHI, Jensen–Shannon divergence, Zipf analysis, and Heaps analysis, together with required sampling diagnostics. Annex P documents the ENT-2025 measurement profile used for the published information-theoretic analysis of Bosnian. The release also includes reproducibility mechanisms such as content-derived `run_id` values, `SOURCE_DATE_EPOCH`, deterministic summation, JSON Schemas, and `verify-run`. **Supersedes:** v1.0-LTS — https://doi.org/10.5281/zenodo.18570562 **Author:** Hasan Kahrimanović Hyper Efficient System LLC ORCID: https://orcid.org/0009-0005-1746-4498 **Licensing:** Specification and specification assets: CC BY 4.0. Reference implementation: MIT License.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.23049395
Primary Topic
Natural Language Processing Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Bosnian CORE NLP Standard (BCS-compatible): Text Normalization & Tokenization — v1.1-LTS

Hasan Kahrimanovic
Zenodo (CERN European Organization for Nuclear Research)
Natural Language Processing Techniques
article

Bosnian CORE NLP Standard (BCS-compatible): Text Normalization & Tokenization — v1.1-LTS

Hasan Kahrimanovic
article en

Abstract

**Bosnian CORE NLP Standard (BCS-compatible) v1.1-LTS** is a deterministic and audit-ready specification for the normalization, sentence segmentation, tokenization, and quantitative measurement of Bosnian Latin-script text. The release combines a formal specification with an executable reference implementation (`bcscore`) written in pure Python ≥ 3.9 and a byte-exact conformance suite containing 118 test cases. Version 1.1-LTS supersedes v1.0-LTS and resolves a set of specification ambiguities and inconsistencies, including handling of ZWJ characters, dashes, abbreviation lists, case folding, segmentation order, and n-gram reset behavior. The release introduces explicit character-stream profiles (`CHAR-LS`, `CHAR-L`, `CHAR-NWS`, `CHAR-FULL`) and machine-readable measurement profiles intended to make quantitative NLP and information-theoretic results reproducible and comparable. Annex M defines the supported quantitative measures, including Shannon entropy, Miller–Madow correction, Onicescu energy, Rényi entropy, Gini–Simpson index, HHI, Jensen–Shannon divergence, Zipf analysis, and Heaps analysis, together with required sampling diagnostics. Annex P documents the ENT-2025 measurement profile used for the published information-theoretic analysis of Bosnian. The release also includes reproducibility mechanisms such as content-derived `run_id` values, `SOURCE_DATE_EPOCH`, deterministic summation, JSON Schemas, and `verify-run`. **Supersedes:** v1.0-LTS — https://doi.org/10.5281/zenodo.18570562 **Author:** Hasan Kahrimanović Hyper Efficient System LLC ORCID: https://orcid.org/0009-0005-1746-4498 **Licensing:** Specification and specification assets: CC BY 4.0. Reference implementation: MIT License.

Zenodo (CERN European Organization for Nuclear Research)
Quality Education
Openalex Percentile: Top 9%
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Bosnian CORE NLP Standard (BCS-compatible): Text Normalization & Tokenization — v1.1-LTS — Hasan Kahrimanovic · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS