Zhongren Wang · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22913478
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This record contains English and Chinese versions of an exploratory research paper on directional-memory language modeling, together with its reproducibility package. The prototype assigns each token entity eight trainable low-dimensional direction codes. Shared low-rank factors aggregate information from a causal 32-token context into a single next-token distribution. It does not impose knowledge-graph constraints. On a frozen Chinese corpus with 12,855 training targets and an active vocabulary of 854 tokens, the eight-head models achieved test perplexities of 75.24 (rank 16) and 69.83 (rank 19). A nearly parameter-matched single-head model achieved 65.42, compared with 68.59 for a smoothed bigram baseline. These results do not demonstrate an eight-head advantage. The package includes both PDF manuscripts, experiment reports and JSON results, the frozen corpus, training and evaluation code, tests, trained experiment weights, and a SHA-256 manifest. The upstream BGE embedding checkpoint is not included; its pinned revision and setup instructions are provided. This is an exploratory, non-peer-reviewed study. Similar-direction soft deduplication is proposed but not implemented, no large-scale language model has been trained, and the test set has been used in multiple exploratory comparisons.
No comments yet — start the discussion below.