Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Language models can learn the right intermediate judgement and the right final action without using the former to produce the latter. This raises a training question more specific than whether a cognition is represented: when, and under what objective pressure, does represented information acquire causal control over downstream policy? We study the acquisition of judgement-like control routes in scientific result triage. At the endpoint of ordinary joint supervision, raw task context directly drives a nonlinear action policy, while hidden states that strongly control the model’s own consequence report are nearly inert for action at the tested sites. We then train matched replicas from the same base model under ordinary judgement-plus-action supervision (T1) and a dependency-oriented objective that additionally rewards judgement-sensitive action (T3), saving dense checkpoints throughout training. With raw scene text fixed, changing only supplied consequence judgement becomes action-effective substantially earlier under T3 and reaches a much larger final gain: the correct high-minus-low intervention is positive by step 48 under T3 and reaches +3.48 pp, versus a late, weak +0.54 pp effect under T1. The learned route is not a clean ( field, item, value ) variable. At T3 step 240, a wrong-item value retains +2.63 pp of active-action effect, while a wrong-field value yields −0.71 pp and redirects probability toward measurement and evidence- collection actions. We therefore propose a control-routing view of training: optimization does not merely learn what to output; it determines which information is routed to which downstream reader, how much gain that route receives, and how precisely the route is typed. Functional field routing can emerge before exact item binding, while endogenous report states may remain outside the action path. These results separate representation learning, routing learning, and binding learning as distinct stages of policy formation.
No comments yet — start the discussion below.