Luis Marcos Vidal, Jay J. Van Bavel · · 2026
DOI: 10.31234/osf.io/4axhg_v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Trustworthy AI is defined by properties built into a system, but trust depends on how users perceive it, so a system can meet every requirement and still fail to be trusted. Human and artificial agents that justify decisions in contractualist terms, by appeal to what affected parties could reasonably accept, are trusted more than those that appeal to aggregate welfare or fixed rules, irrespective of the action selected, but why is unexplained. We propose that a contractualist justification discloses individualized regard, evidence that the affected party's standpoint entered the decision, which is what trust, as a judgment that an agent is willing to act well, is sensitive to. Perceived regard is the mediating inference and raises trust through two pathways: warmth, and the expectation that the agent's future decisions will take the perceiver's own stakes into account. The intentional stance is a gating condition rather than a mediator, deciding whether an agent can be the object of this inference at all, so the mechanism is shared by human and artificial agents and needs no genuine intentions. Such justification is the observable trace of virtual bargaining, a theory-of-mind process, and does not require agents with genuine mental states. We state five falsifiable hypotheses that discriminate the account from warmth, anthropomorphism, and mere-transparency alternatives, and identify moderators and a possible moral uncanny valley as further tests.
No comments yet — start the discussion below.