Human-Reviewed Uzbek Legal Named Entity Recognition Dataset
This article describes a human-reviewed Uzbek legal-domain named entity recognition (NER) dataset developed as a reusable resource for low-resource legal NLP. The release contains 12 entity categories: PER, ORG, LOC, DATE, MONEY, POSITION, DOCNO, LAW, COURT, BANK, TIN, and CADASTRE. The dataset is provided in XLSX, CSV, JSON, and JSONL formats and is structured into two complementary layers: a core subset of manually reviewable source-grounded records and an extended augmented subset used to support lower-frequency labels in training-oriented settings. The package also includes supporting documentation, split guidance, a data dictionary, and review-related metadata, including provenance, verification status, and quality flags. Character-level start and end offsets are included where recoverable. The release is intended to facilitate Uzbek legal NER research, resource curation, and transparent reuse under provenance-aware conditions.
Authors
- Nasiba B. Azizova (ORCID: https://orcid.org/0000-0001-8579-197X)
- Gulnoza Narkabilova
- Bobur R. Saidov (ORCID: https://orcid.org/0009-0000-5540-2013)
- Zilolakhon Ruzmetova
- Feruzakhon Rustamova
- Firuza Halimova
- Umida Bazarova
Institutions
- Andijan State University (UZ)
- Urgench State University (UZ)
- Kurgan State University (RU)
- The Future University (SD)
- Karshi State University (UZ)
- Navoi State University (UZ)
- Andijan Institute of Agricultural (UZ)
- Samarkand State Institute of Foreign Languages (UZ)
- Navoi State Mining Institute (UZ)
- Andijan State Medical Institute (UZ)
Publication Details
- Journal
- F1000Research
- Published
- 2026-09-15
- DOI
- https://doi.org/10.12688/f1000research.180408.2
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00