Hasan Kahrimanovic · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.18570561
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
**Bosnian CORE NLP Standard (BCS-compatible) v1.1-LTS** is a deterministic and audit-ready specification for the normalization, sentence segmentation, tokenization, and quantitative measurement of Bosnian Latin-script text. The release combines a formal specification with an executable reference implementation (`bcscore`) written in pure Python ≥ 3.9 and a byte-exact conformance suite containing 118 test cases. Version 1.1-LTS supersedes v1.0-LTS and resolves a set of specification ambiguities and inconsistencies, including handling of ZWJ characters, dashes, abbreviation lists, case folding, segmentation order, and n-gram reset behavior. The release introduces explicit character-stream profiles (`CHAR-LS`, `CHAR-L`, `CHAR-NWS`, `CHAR-FULL`) and machine-readable measurement profiles intended to make quantitative NLP and information-theoretic results reproducible and comparable. Annex M defines the supported quantitative measures, including Shannon entropy, Miller–Madow correction, Onicescu energy, Rényi entropy, Gini–Simpson index, HHI, Jensen–Shannon divergence, Zipf analysis, and Heaps analysis, together with required sampling diagnostics. Annex P documents the ENT-2025 measurement profile used for the published information-theoretic analysis of Bosnian. The release also includes reproducibility mechanisms such as content-derived `run_id` values, `SOURCE_DATE_EPOCH`, deterministic summation, JSON Schemas, and `verify-run`. **Supersedes:** v1.0-LTS — https://doi.org/10.5281/zenodo.18570562 **Author:** Hasan Kahrimanović Hyper Efficient System LLC ORCID: https://orcid.org/0009-0005-1746-4498 **Licensing:** Specification and specification assets: CC BY 4.0. Reference implementation: MIT License.
No comments yet — start the discussion below.