Zhimai Hou, Jinhuan Wu, Ying Geng, Shuo Lin, Guoliang Yang · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22918573
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Typed "System One" decision models claim to emit schema-valid actions with calibrated probabilities at millisecond latency, at a fraction of language-model cost---claims that currently have no third-party evaluation. We introduce SystemOne-UAV, a benchmark that makes each claim falsifiable without access to the model's weights or vendor cooperation. Core design: a UAV mission simulator in which a hidden battery-sensor bias (30% of episodes, U(8,25)%) renders identical observations with different optimal actions, so the Bayes-optimal predictor is forced to hedge---calibration error then has an absolute yardstick, not a proxy. The protocol ships (i) an oracle utility specification reproducing every label, (ii) three test regimes (IID, noise x3, events x3), (iii) a six-objective baseline matrix, and (iv) measurement rules that separate interface properties (type legality under a closed vocabulary) from learned properties (confidence quality). Baselines reproduce the RLVR over-confidence pathology at the decision-token level (five-seed accuracy-only RL: ECE 8.5x worse, .0418 vs .0049, paired t=20.7, accuracy statistically unchanged), show that "zero type errors" is an identity under closed-vocabulary readout, and place batched local inference at 0.08--0.19 us per decision---five to six orders of magnitude below the advertised hosted latency, isolating calibration, not speed, as the binding constraint for autonomy. We also report what is, to our knowledge, the first third-party evaluation of an open-weight System One model on a physical-control benchmark (concurrent third-party work probes open-weight typed models on text workflow decisions only [Sun and Xu, 2026]): Laya-421M [ConvAI Innovations, 2026] transferred zero-shot reaches .399 accuracy (majority class .366), Brier .751 and confidence-vs-error AUROC .655, and a conformal gate built on its confidence returns risk .585 at a nominal alpha=.05; a frozen-encoder, refit-head control attributes roughly 70% of that failure to representation rather than readout. Calibration does not transfer across task families, and an escalation contract is only as strong as the ranking quality of the confidence it consumes. An adapter slot for the hosted API reports pending rather than imputed numbers. All code, seeds, and run artifacts are released.
No comments yet — start the discussion below.