Said Chah Slaoui · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23025360
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Background Large language models often transform mathematical, logical, and symbolic inputs into representations that resemble canonical forms. We call this phenomenon implicit canonicalization. It is an effective mechanism for discovering structured candidates, but it does not establish that a candidate is equivalent to its input, canonical, or valid: probabilistic generation and formal decision are distinct functions. Methods We introduce S-AI-CCC, a controlled cognitive canonicalization framework within Sparse Artificial Intelligence (S-AI). A language model proposes candidates; S-AI-Recursive stabilizes the cognitive trajectory through a hormonal controller built on two antagonistic signals, Clarifine for convergence and Confusionin for residual uncertainty; and S-AI-RLM applies a total certificate that combines equivalence to the input, canonical-form membership, and validity in a decidable target language. Every admissible input receives a terminal verdict (Accept, Reject, or Abstain), and a canonical form is committed only on certified acceptance. The framework is evaluated on the public mixed Boolean-arithmetic (MBA) simplification benchmarks, in a reference configuration whose first layer is a declared surrogate proposer rather than a language model. Results Under explicit conditions on dissipation, coupling, emissions, and delays, the hormonal dynamics are exponentially stable in the deterministic regime and ultimately bounded in mean square under persistent noise, and the architecture reaches a terminal verdict in finite time. Dynamical convergence is shown to be distinct from validity. Over 11,950 public expressions, the system returns 11,400 Accepts, 329 certified Rejects, and 221 Abstains, with no false Accept detected by an independent check. An uncertified proposer is wrong in 94.9% of cases; with capped evidence, a converging recursion commits 92.55% wrong outputs without verification and none with it. Canonical outputs are unique in 17 of 17 equivalence classes. The evaluation also delimits the approach: certification beyond the linear fragment is rarely tractable, a specialized simplifier covers more inputs, and one contraction hypothesis fails during the discovery phase. Conclusions S-AI-CCC turns implicit canonicalization into a controlled process in which language models discover candidates while guarantees are carried by the layers able to support them: stability by the regulated dynamics, correctness by the total verifier. Certification prevents silent errors, at the price of declared abstention wherever proofs are not tractable.
No comments yet — start the discussion below.