Manish Kumar Parihar · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22942524
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Problem:Transformer is called black box because we cannot see which embedding weights were touched by input and which output IDs were top. Idea like tokenizer:In tokenizer, word_12 gets fixed ID 12. During training embedding values update but ID 12 stays 12. Index fixed, value updates. Same for neurons and weights in this experiment:- Each neuron gets fixed N ID in sequence: N0,N1,N2... In experiment 4096 neurons -> N IDs 0 to 4095 fixed.- Each weight gets fixed S ID in sequence: S0,S1,S2... In experiment 1,179,648 weights -> S IDs 0 to 1179647 fixed.- During training values update but N and S ID numbers stay same. S_map example: N_embedding S 0-524287, W_Q S 524288-540671, W_K S 540672-557055 etc. Counting done once, IDs fixed. What is actually auditable and proven in fresh experiment: 1. Embedding trace has clean list — auditable:Input ["word_12","word_45","word_78","word_200"] -> N IDs [12,45,78,200]N 12 uses S 1536-1663 only (128 weights)N 45 uses S 5760-5887 onlyN 78 uses S 9984-10111 onlyN 200 uses S 25600-25727 onlyNot all weights, only these. This is logged and repeatable. 2. Output N IDs trackable:For query, activated N IDs top10 = [2895,209,3716,1615,3982,1957,2305,478,2521,2506]Same query again gives same list -> first == second True -> fixed and trackable. This uses threshold top-k which is standard in interpretability research. 3. Training keeps IDs fixed:Before training N count 4096 after 4096 same, S count 1179648 after 1179647 same -> IDs fixed while values updated. Proven in fresh_experiment_proof.txt. What is NOT claimed in this version:Middle layers W_Q,W_K,W_V,W_O,FFN are dense — almost every weight gets small continuous number, not simple on/off. So I do not claim full middle chain in simple ordered list. Full causal trace needs more work like activation patching. Answer to "just renaming indices" feedback:Yes S IDs are renaming flattened indices, like tokenizer IDs are renaming. But fixed address is needed to log and audit. Without fixed ID you cannot say which S IDs were touched. Answer to "model cannot inspect own weights" feedback:That is true for large production models, but not for small owned model where weights are accessible. This experiment runs on owned model with full weight access. Proof files included:fresh_experiment_proof.txt and auditable_trace.json show fixed N and S counts, embedding S trace, top10 N repeatability. Conclusion:Giving fixed N ID to each neuron and fixed S ID to each weight makes embedding touch auditable and output N top10 trackable. Works for thousands or lakhs of neurons as same counting method. Full middle layer dense trace is future work. Author: Manish Kumar Parihar
No comments yet — start the discussion below.