Panagiotis Gkilis · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22798415
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Twenty neural audio decoders measured under one frozen five-gate protocol on identical audio, each gate pre-registered before measurement and each instrument attacked before it was trusted. Only three of twenty arms truly stream. An arm counts as incremental only if stateful chunked decode reproduces full-context output (max|err| ≤ 1e-3) and state is load-bearing (stateless error ≥ 10× stateful). Three explicitly causal FocalCodec configurations pass, with C2 ratios of 3,530-36,740 against a threshold of 10; fifteen arms use stateless chunking. Non-causal configurations driven through the identical stateful code path give a ratio of exactly 1 - the negative control. The sole pre-registered cross-arm ranking metric failed its own validity instrument. Griffin-Lim, a zero-parameter phase-retrieval algorithm, scores best on the mel-cepstral distance in six states of six. The pre-registration declared in advance what that would mean, so it is reported as a property of the metric rather than a decoder ranking; a proposed explanation was tested and refuted. A second measurement sharing no implementation found the same three survivors. Speaker-embedding retention across three separately calibrated encoders: the causal configurations show no measured loss when streamed (offline-vs-streamed cosine 0.9999) while thirteen others lose -0.078 to -0.742. The two measurements are independent in construction but not in data - both run on the same reconstructions. Not measured: perceptual quality. No listening test was run, and nothing here says which decoder sounds best. The study passed four rounds of independent adversarial audit, which found three defects that changed conclusions; those defects and their repairs are reported in the paper. Not released: the source audio (1,512 private, rights-cleared recordings of seven consented speakers), the derived speaker embeddings, and the reconstructed audio. Speaker names are pseudonymised and private paths masked; every substitution is recorded with its source hash.
No comments yet — start the discussion below.