YILDO-LLM-100M-Tri Base-v1.0: From-Scratch Training, Technical Report, and Public Release Record
YILDO-LLM-100M-Tri Base-v1.0 is a compact trilingual decoder-only Transformer language model developed for Arabic, English, and Turkish. The model contains 95,371,008 trainable parameters, uses a project-trained 32,000-token byte-level BPE vocabulary, and supports a 512-token context window. The tokenizer, model implementation, model weights, data-preparation workflow, and training pipeline were developed within the YILDO AI Research Project. Model weights were initialized randomly, and no external pretrained language-model weights were used. This technical note documents the Base-v1.0 frozen release, including corpus preparation, tokenizer characteristics, training record, multilingual evaluation, standardized benchmark results, reproducibility controls, and checkpoint integrity. Implementation-level architectural and optimization details are intentionally withheld pending intellectual-property review. YILDO-LLM-100M-Tri Base-v1.0 is intended to provide a reproducible technical baseline for continued research, supervised fine-tuning, and future scaling of the YILDO-LLM family.
Authors
- YILDIRIM SALAHALDIN HUSSEIN (ORCID: https://orcid.org/0009-0007-7411-1503)
Institutions
- University of Kirkuk (IQ)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23166553
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00