Pranay M. Mahendrakar · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23026299
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0. Work on trust between language-model agents produces two kinds of object. Protocol work produces identity, attestation, stake and constraint, all bound at the transport layer before any content reaches a model. Behavioural work produces a per-source score: a reputation, a credibility, a reliability weight. Statements of motivation across the security literature suggest that the score has nowhere to go, because verified and unverified content arrive in the same undifferentiated context window and the model has no way to discount the low-trust part. This paper argues that this suggestion, taken literally, is false, and that it is true in a narrower form that matters more. It is false because at least four consumers of a trust value exist and have published results: admission and routing before the model, annotation written into the prompt, modulation inside the forward pass, and gates on actions after the model. Credibility annotations and attention scaling do move model outputs, and models already weigh source labels they were never asked to weigh. It is true in a narrower form because every graded discount the model reads that was located for this paper, whether written into the prompt or applied inside the forward pass, was evaluated against sources that err or against attacks fixed in advance, never against an attacker who adapts to the discount; the one graded router located that was attacked by a source writing its own evidence was captured. The label, the count of copies and the evidence behind a score are each writable by an attacker, and published attacks write all three: forged role tags, repeated low-credibility text, fabricated episodes that launder reputation. The trust defences that report guarantees against attacking sources consume trust as a discrete label or capability enforced by code outside the model, never as a graded weight the model reads, although their guarantees rest mostly on proofs and on fixed attacks rather than on adaptive evaluation. Against sources that err, by contrast, a graded discount may do better than exclusion when scores are noisy. The paper sets out the four-consumer partition, analyses which inputs of a graded consumer an attacker can write, argues that trust scored per source cannot follow influence that arrives per token, consolidates the measurements in one table, states the routing decision as an algorithm, and names eight studies that would settle the open part. No experiments are reported here.
No comments yet — start the discussion below.