Anh Khoa Doan Ngoc · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22801505
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
System-one decision models — fast, non-generative networks that emit typed, calibrated probabilistic decisions for machine-to-machine pipelines, of which TypeSafe AI's Jev, trained by Reinforcement Learning for Calibrated Decisions (RLCD), is the announced instance — are audited by hop-level calibration on stationary held-out data. We show that this class omits two structural primitives: recursive belief over a latent state (the hidden-Markov primitive) and graded predicate membership (the fuzzy primitive). We prove two negative results. Lemma 1: when outcomes are coupled through a latent Markov regime, marginal calibration is not preserved under regime shift, trajectory error counts are overdispersed relative to the binomial baseline that hop-level calibration implies, with a closed-form inflation factor, and every permutation-invariant audit statistic — expected calibration error included — has exactly zero power to detect this. Lemma 2: thresholding a typed score makes the composed pipeline action discontinuous; an ε-perturbation of evidence flips a full action on a set of measure Θ(ε), every downstream hop then receives an O(1) input change so the conditional flip probability is Θ(1), and when several hops' thresholds cross at a common value of a shared latent — as when a harness re-thresholds a score the model has already thresholded — the all-hops-flip event has measure Θ(ε) rather than Θ(ε^H); a fuzzy membership composed by a t-norm and defuzzified once by a Lipschitz map keeps the action Lipschitz in the evidence. We state the Principle of Deferred Crispification and specify a reference architecture, Belief-State Fuzzy System-One (BSF-S1), whose output object carries a per-regime membership matrix, a calibrated distribution over the degree, and a scaled-likelihood vector over regimes, and collapses exactly once at the actuator at O(K^2+KM) cost per step. Five CPU-scale experiments, the two temporal ones run over five seeds against the exact filter as the achievable floor and a Bayes-optimal memoryless head as the strongest memoryless competitor, are consistent with each prediction of the lemmas and expose where a learned approximation falls short of them; two operationally defined audit metrics (trajectory calibration error and action-mass sensitivity) can be run against a closed model from the outside. The paper is architectural: Jev is closed, and nothing here is an empirical claim about its internals. Code, results and revision history: https://github.com/dnakhoa/jev-deferred-crispification
No comments yet — start the discussion below.