Yibo Chen · Figshare 2026 · 2026
DOI: 10.6084/m9.figshare.34046073
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
When a language model assesses a scientific result, it is natural to imagine that it first forms stable internal variables–for example, how well established the result is and how consequential it would be if true–and then reads those variables out into a report. We find a more layered and reader-dependent computation. First, an orthogonal seven-factor study varies raw evidence, methodological quality, replication, verification, prior plausibility, conditional consequence, and target relevance. Structured reports preserve all seven directed factor effects on held-out scenes, ruling out a simple account in which complete scientific assessment is exhausted by two fixed report variables. Second, low-dimensional hidden geometry is nevertheless real: a nuisance-controlled late assessment span is stable within its source population, while new evidence/consequence coordinates systematically rotate across reordered, paraphrased, and free-form interfaces. Third, exact-state interventions identify a clause-local causal assessment signal. Consequence-associated perturbations transfer across scenes and paraphrases, but their semantic field is not self-contained in the donor state. A perturbation derived from another factor can produce nearly the same consequence effect when inserted into the recipient consequence slot, whereas the same consequence perturbation is much weaker when inserted into a different slot. The donor directions themselves are only moderately aligned. These results support a reader-assigned assessment computation: upstream states preserve factor-sensitive signals and partially portable signed/value components; the recipient slot and local reader assign much of the field semantics; and the report interface determines how these components are mixed and amplified into settlement, consequence, verification, relevance, and other outputs. The earlier low-dimensional assessment geometry is therefore better understood as a reader-specific pooled summary than as a universal fixed semantic coordinate system. Scientific judgement in language models is not simply stored and then reported; it is assembled at readout
No comments yet — start the discussion below.