Minseong Sim · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22859044
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Reinforcement learning with verifiable rewards (RLVR) evaluates decoded outcomes, whereas policy-gradient estimators operate on concrete trajectories and batch-dependent training state. These two levels need not respect the same equivalence relation. We formalize estimator homomorphism compatibility: when several concrete trajectories represent the same reward- or environment-level outcome, the stochastic estimator should not introduce an expected update component that distinguishes representatives solely through estimator state. For exchangeable score estimators with batch-coupled scalar coefficients, we derive an exact quotient/fiber decomposition in which the within-class component is governed by the covariance between a representative-conditioned effective coefficient and the conditional fiber score. Exact enumeration validates the identity. Controlled Transformer experiments, a frozen Hugging Face TRL GRPOTrainer, whole-model training, and cross-family tests show measurable estimator-induced representation pressure. An equal-length control demonstrates that the measured compatibility gap is not reducible to response-length weighting: two one-token strings accepted as the same answer can enter different clipping states after policy lag and yield different optimizer updates. Lower-learning-rate whole-model and LoRA follow-ups reproduce selective clipping in additional settings. We do not show that the effect necessarily harms downstream reward. The contribution is a quotient-conditioned estimator diagnostic, a decomposition that organizes multiple known estimator mechanisms, and direct measurements in LLM RLVR training.
No comments yet — start the discussion below.