YILDO-LLM-100M-Tri Base-v1.0: From-Scratch Training, Technical Report, and Public Release Record

YILDO-LLM-100M-Tri Base-v1.0 is a compact trilingual decoder-only Transformer language model developed for Arabic, English, and Turkish. The model contains 95,371,008 trainable parameters, uses a project-trained 32,000-token byte-level BPE vocabulary, and supports a 512-token context window. The tokenizer, model implementation, model weights, data-preparation workflow, and training pipeline were developed within the YILDO AI Research Project. Model weights were initialized randomly, and no external pretrained language-model weights were used. This technical note documents the Base-v1.0 frozen release, including corpus preparation, tokenizer characteristics, training record, multilingual evaluation, standardized benchmark results, reproducibility controls, and checkpoint integrity. Implementation-level architectural and optimization details are intentionally withheld pending intellectual-property review. YILDO-LLM-100M-Tri Base-v1.0 is intended to provide a reproducible technical baseline for continued research, supervised fine-tuning, and future scaling of the YILDO-LLM family.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23166554
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

YILDO-LLM-100M-Tri Base-v1.0: From-Scratch Training, Technical Report, and Public Release Record

YILDIRIM SALAHALDIN HUSSEIN
Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
article

YILDO-LLM-100M-Tri Base-v1.0: From-Scratch Training, Technical Report, and Public Release Record

YILDIRIM SALAHALDIN HUSSEIN
article en

Abstract

YILDO-LLM-100M-Tri Base-v1.0 is a compact trilingual decoder-only Transformer language model developed for Arabic, English, and Turkish. The model contains 95,371,008 trainable parameters, uses a project-trained 32,000-token byte-level BPE vocabulary, and supports a 512-token context window. The tokenizer, model implementation, model weights, data-preparation workflow, and training pipeline were developed within the YILDO AI Research Project. Model weights were initialized randomly, and no external pretrained language-model weights were used. This technical note documents the Base-v1.0 frozen release, including corpus preparation, tokenizer characteristics, training record, multilingual evaluation, standardized benchmark results, reproducibility controls, and checkpoint integrity. Implementation-level architectural and optimization details are intentionally withheld pending intellectual-property review. YILDO-LLM-100M-Tri Base-v1.0 is intended to provide a reproducible technical baseline for continued research, supervised fine-tuning, and future scaling of the YILDO-LLM family.

Zenodo (CERN European Organization for Nuclear Research)
University of Kirkuk (IQ)
Openalex Percentile: Top 10%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.