Randolph James Ferlic, Kimberly Kate Ferlic · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22968133
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The Acoustic Tier-0 Gate: A Deep, Broad, and Adversarially Stress-Tested Map of a Frozen Single-Token Always-On Audio Classifier and Machine-Health Monitor Randolph James Ferlic, M.D. and Kimberly Kate Ferlic — Fieldstone Analytics, LLC, Austin, TX, USA Preprint · Zenodo DOI: 10.5281/zenodo.22968134 · CC-BY 4.0 · Community: spiral-domain-encoder-campaign Abstract The microphone is the canonical always-on sensor, and the always-on audio layer — the logic that listens continuously and decides whether anything is worth a richer model's attention — is exactly the Tier-0 tier we have argued the 2026 agentic edge leaves unowned: an energy budget dominated not by the occasional expensive inference but by the perpetual cheap one. We study a frozen, deterministic, class-discriminant single-token encoder (a 128-dimensional log-mel front-end → a supervised linear-discriminant ⊕ principal-component subspace → a k-means codebook of at most 256 cells → an 8-bit token → a per-cell lookup-table decision; and, without labels, a per-machine k-means "self-twin" novelty detector) as that always-on acoustic layer. One decision costs on the order of 1.1 million operations for the shared time-frequency front-end plus ~12,000 for the token itself — about 2.2 microjoules in floating point, ~0.17 in fixed point, and roughly 220× below a deployed keyword-spotting network's inference — validated against a measured 0.774 ms per decision. Rather than merely demonstrate that the encoder "works," we build and adversarially stress a predictive, deployment-facing map of where it wins, pays, and fails across two capabilities and five public venues (Free Spoken Digit, ESC-50/ESC-10, Speech Commands, UrbanSound8K, and the MIMII industrial-machine corpus, all four machine types), through 34 pre-registered results. On classification, the token is near-parity with a strong gradient-boosted model on identical features where the task is few-class and clean (spoken digits +0.025; ten environmental classes +0.052) and pays an honest, priced tax where the task is hard — driven by difficulty (speaker variability, confusable words: keyword spotting −0.19) more than by class count alone. On label-free machine health, a per-machine self-twin gates anomalies with mean AUROC 0.906 across all four MIMII machine types (15 of 16 units), commissions from ~16 clips of normal sound, matches the DCASE-standard autoencoder (0.838 vs 0.834) at ~6–8× less inference and with deterministic, closed-form behavior, and — necessarily — must be built per machine (own 0.87–0.96 ≫ cross 0.42–0.65). We characterize the deployment surface an acquirer's diligence actually hits: a bounded but dial-able operating point (recall 0.29→0.56 at 1→10% false-positive rate), calibration (raw confidence needs recalibration), a complete noise story (a train/test noise mismatch crashes the codebook, and only commissioning-time augmentation repairs it — two of our own predictions, observation redundancy and a spectral-subtraction denoiser, were falsified by the data), and a self-critical methodology finding (a random split hides normal-operation non-stationarity that a temporal split exposes). We map a second-tier upgrade path (zero-label gate → free temporal persistence +0.146 → few-shot supervised confirm 0.77–0.93) and two platform properties (one core runs a classifier and a novelty head at ~0 tax; per-capability energy falls as ~1/M). The map is anchored by five honest negatives. The token is not an accuracy champion and does not pretend to be; it is a predictable, auditable, private, self-healing cost instrument — which is exactly what the always-on layer on the microphone should be. This is a characterization of previously described, filed methods; it discloses no new algorithmic subject matter, and the per-deployment selection of configuration is retained as trade secret. Highlights · The reframe — a Tier-0 layer on the microphone: the always-on audio tier (is this speech or noise? my wake word? a normal machine or a failing one?) is the acoustic instance of the sub-milliwatt layer beneath the NPU. One decision ≈ 1.1 M shared front-end operations + ~12 k token operations (~2.2 µJ, ~0.17 µJ fixed-point; the token's marginal arithmetic ~23.6 nJ, ~220× below a keyword-spotting inference), validated at 0.774 ms/decision, with the raw audio kept off the wire — one byte out. · A predictive classification map across five venues: the token is near-parity where the task is few-class and clean (Free Spoken Digit +0.025, ESC-10 +0.052) and pays a priced tax where it is hard (Speech Commands −0.19, UrbanSound8K +0.18, ESC-50 +0.155). The driver is task difficulty, not class count alone — a ten-class keyword benchmark taxes the token as much as a fifty-class one. · A label-free machine-health gate across ALL FOUR MIMII types: per-machine self-twin (k-means novelty on normal-only sound) at mean AUROC 0.906 (15/16 units) — valve 0.842, slider 0.893, pump 0.953, fan 0.937 — commissioned from ~16 clips, matching the DCASE-standard autoencoder (0.838 vs 0.834) at ~6–8× less inference and deterministically, and necessarily per machine (own 0.87–0.96 ≫ cross 0.42–0.65). · The deployment surface, characterized honestly: a bounded but dial-able operating point (recall 0.29 → 0.56 at 1 → 10% FPR); calibration (recalibrate before gating); a complete noise story — a train/test mismatch crashes the codebook and only commissioning-time augmentation repairs it (0.100 → 0.502), while two of our own predicted test-time fixes (observation redundancy; a spectral-subtraction denoiser) were falsified by the data; cross-mic transfer solved by per-device re-commissioning; bit-exact determinism (float32 = float64; ≤256-entry lookup table). · A self-critical methodology finding: a random split reports a 4.3% false-positive rate while a temporal, deploy-forward split exposes 80.8% — normal-operation non-stationarity that must be validated on a temporal split (threshold-free AUROC is robust; the fixed operating point is not). Likely a corpus artifact; the lesson stands. · A second-tier upgrade path: zero-label gate (bounded) → free temporal persistence (+0.146) → richer unsupervised models do not help (the ceiling is the label-free constraint) → a few-shot supervised confirm breaks it (5 labels → 0.77, 50 → 0.93). · Two platform properties: one shared token runs a classifier and a label-free novelty head at ~0 tax (open-set unknown-class rejection is honestly weak), and per-capability energy falls as ~1/M — one sub-milliwatt core, N capabilities, at roughly the energy of one. · Five anchoring negatives (difficulty tax; temporal-order front-end limit; weak open-set OOD; the noise mismatch with two falsified fixes; normal non-stationarity). Honesty as evidence: a characterization that never reports a failure has not been stressed. · All characterization of filed / published methods — no new algorithmic subject matter; the accuracy-difficulty boundary and the non-stationarity behavior are properties of the frozen method; the per-deployment configuration-selection and commissioning procedure is a trade secret. What this record contains · `Manuscript_Paper50.pdf` — the manuscript, nine figures embedded (the difficulty axis; the label-free gate; the per-machine matrix; the gating economics; the temporal-order front-end limit; the per-decision energy/op-count; the capability×energy frontier; the industrial four-type gate; the venue map), six tables, and 73 references; and `Manuscript_Paper50.docx`, the editable source. · `PAPER_50_ZENODO_ARCHIVE.zip` — the reproducibility archive (md5 in ARCHIVE_MD5.txt): the frozen pre-registration (A1/A2/P1/G1/B1/E1), the runners (six core-hypothesis; four reviewer-credibility; three acquirer-diligence; five second-tier; five robustness; five very-skeptical; two platform), the frozen 128-dim log-mel front-end (audio_feats.py), the frozen token encoder module, the figure builders, the two Modal apps, the per-experiment result records (JSON) and per-simulation notes (markdown) — including the industrial-breadth and UrbanSound8K result JSONs — the nine figures, the manuscript source, and a README. All datasets are public and not redistributed (fetched from their public sources at run time); all paths and identifiers are scrubbed (absolute paths → PATH_TO_DATA/PATH_TO_SCRATCH, any cloud handle → MODAL_USER) and leak-scanned. Cite as R. J. Ferlic and K. K. Ferlic, "The acoustic Tier-0 gate: a deep, broad, and adversarially stress-tested map of a frozen single-token always-on audio classifier and machine-health monitor," Zenodo, 2026, doi: 10.5281/zenodo.22968134. License and patent notice Released under CC-BY 4.0. Consistent with that license, no patent or IP right of the authors is licensed, waived, or conveyed by this deposit. This work characterizes previously-described methods and discloses no new algorithmic subject matter; gradient boosting, k-means / vector quantization, product quantization, Fisher discriminant analysis, log-mel and mel-frequency-cepstral front-ends, temperature scaling, spectral subtraction, autoencoder anomaly detection, and nearest-centroid novelty detection are established prior art, used only as tools. The methods characterized — the class-discriminant single-token codebook encoder and its nearest-centroid monitor, the multi-token / token-ladder and soft-readout mechanisms, the inference-time co-channel fusion, the foundation codebook, and the per-entity self-twin — are the subject of filed and pending U.S. patent applications held by the authors, including U.S. Provisional Application No. 64/095,354 (the encoder), the personalization / on-device-adaptation applications (priority U.S. Application No. 19/467,303 and its continuations), the multi-token / token-ladder application (No. 64/119,487), and the inference-time fusion application (No. 64/137,805). The observed accuracy-difficulty b
No comments yet — start the discussion below.