YILDIRIM SALAHALDIN HUSSEIN · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23166553
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
YILDO-LLM-100M-Tri Base-v1.0 is a compact trilingual decoder-only Transformer language model developed for Arabic, English, and Turkish. The model contains 95,371,008 trainable parameters, uses a project-trained 32,000-token byte-level BPE vocabulary, and supports a 512-token context window. The tokenizer, model implementation, model weights, data-preparation workflow, and training pipeline were developed within the YILDO AI Research Project. Model weights were initialized randomly, and no external pretrained language-model weights were used. This technical note documents the Base-v1.0 frozen release, including corpus preparation, tokenizer characteristics, training record, multilingual evaluation, standardized benchmark results, reproducibility controls, and checkpoint integrity. Implementation-level architectural and optimization details are intentionally withheld pending intellectual-property review. YILDO-LLM-100M-Tri Base-v1.0 is intended to provide a reproducible technical baseline for continued research, supervised fine-tuning, and future scaling of the YILDO-LLM family.
No comments yet — start the discussion below.