Arin Agarwal · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22850169
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
We ask whether behavioral constraints acquired through post training remainexplicitly reportable. Using constrained recipe generation as a testbed, fivebanned ingredients enforced via LoRA fine tuning of Llama 3.1 8B Instruct wecompare supervised fine tuning (SFT) and Group Relative Policy Optimization(GRPO) against an untrained baseline on a four tier Constraint AwarenessBenchmark. Averaged over three seeds, both methods raise behavioral compliancefrom 4% to about 90% while reducing explicit constraint reporting below theuntrained model (0.48/5 to 0.16/5 for SFT, 0.07/5 for GRPO) and erodingretained third person knowledge (93% to 36% for SFT, 14% for GRPO; p less than0.01 between methods). Contrary to our initial hypothesis, the reward basedsignal is the more destructive of the two: a reward that penalizes bannedingredient tokens regardless of framing learns a context independentsuppression rather than a self directed constraint. A context conditionedreward designed to teach the self to other distinction fails, collapsing towardinclusion in both framings. Probing prompt time hidden states recovers peringredient avoidance at 83.8% (layer 24 MLP), but only 6.4 points above a peringredient base rate predictor (77.4%), and the model's own verbal self reportis more accurate still (87.8%). A positive control adding explicit selfdescription examples does not restore reporting. The failure is thereforespecific to enumerating constraints on request, not a general loss of access tothem.
No comments yet — start the discussion below.