Nikhil Bhagat · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22916910
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Retrieval-augmented generation systems assume that newer documents supersede older ones. Statutory law violates this assumption: when an Act is amended, the earlier text remains authoritative for events preceding the amendment. Recent benchmarks show that retrieval over a static consolidated corpus returns real but inapplicable provisions, but they cover high-resource languages with strong retrievers. Nepali has neither: it has no standard information retrieval benchmark, and published results show lexical retrieval outperforming multilingual dense retrieval on Nepali legal text. Building a version-aware Nepali corpus appears blocked: Gazette notices are scanned images and official Act PDFs extract to corrupted text. We show that the amendment record is nonetheless recoverable, from numbered footnotes carried inline in consolidated Acts, which name the amendment and its operation (insertion, substitution, or repeal) at clause granularity. We present NEPVERSA, the first version-aware Nepali statutory benchmark, whose candidate gold is derived mechanically from these published statements and scored by deterministic nuggets rather than by a language-model judge. Footnote markers are read from the published HTML, where each of the 119 markers is either attached to a unit or reported; a Markdown-only reading of the same pages misses 26 amended units and invents 6. Our central finding is that an approximately correct parser yields a corpus that appears entirely correct, because a missed amendment is indistinguishable from a provision never amended. The release covers 5 Acts, 455 provisions and 181 items, each checked against its source footnote by the author; a legal-expert audit remains outstanding. Preprint; not peer reviewed. Code and data (release v2.0.0): https://github.com/NikeGunn/nyasathi-reserch
No comments yet — start the discussion below.