
S. Aasaithambi · Qeios 2026 · 2026
DOI: 10.32388/cq9g1l
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Enterprise skills graphs, which connect skills to the people who hold them and the work that requires them, have become the data substrate on which agentic HR systems recommend hiring, mobility, reskilling, and redeployment. A consensus has formed among HR analysts that the value of these systems depends less on the sophistication of the agents than on the trustworthiness of the graph beneath them. That trustworthiness, however, is discussed almost entirely in qualitative terms. Where it is quantified at all, it is reduced to a single proxy: coverage, the percentage of roles or people for which some skill data exists. Coverage answers how much data an organisation holds. It is silent on whether that data can be trusted enough to let an autonomous agent act. This paper develops the missing measurement science. Working within a design-science method and drawing on established data-quality standards (ISO/IEC 25012, ISO 8000, the FAIR principles) and on machine-learning calibration theory, it defines a reliability construct for skills graphs along five dimensions: Freshness, Accuracy of stated confidence, Contestability, Traceability, and Stability, which form the mnemonic FACTS. The dimensions are derived from the conditions under which an automated decision is defensible rather than assembled as a list, and each is operationalised as a metric computable on a graph's nodes and edges. They are combined into a dashboard-ready Skills Graph Reliability Index (SGRI) through a gated, deny-overrides aggregation rather than a weighted mean, so that weakness on any one dimension cannot be concealed by strength on another. SGRI resolves to three decision bands: act, human-gate, or do not automate. The paper makes five contributions. To begin with, it argues that unreliability in a skills graph compound rather than averages out under agentic use, because agent inferences written back without provenance contaminate the evidence against which later inferences would be validated, a mechanism named here evidential circularity. Second, it demonstrates that mechanism computationally. In a simulation of 20,000 edges over twelve cycles, unlabelled write-back caused measured calibration error to improve while true calibration error degraded, and the divergence increased with the write-back rate. At a write-back rate of 0.50 the contaminated graph reported a lower calibration error than a fully labelled control while its true error was approximately 4.8 times worse, so an organisation comparing the two on measured calibration would have preferred the corrupted graph. Third, the paper specifies the SGRI aggregation and decision bands and defends the deny-overrides rule on grounds of asymmetric loss. Fourth, it states the boundary of the construct explicitly: SGRI measures the trustworthiness of the evidence base, not the quality of inference drawn from it. Fifth, it confronts the ground-truth problem, since calibration in this domain depends on an oracle that is itself biased, and proposes triangulation and mandatory oracle-provenance disclosure as partial remedies. Fairness is treated as a stratification requirement across all five dimensions rather than as a sixth dimension. A reference architecture, a worked illustration, and a set of falsifiable propositions complete the paper. No claim is made that the index predicts decision quality; that relationship is stated as a falsifiable proposition and remains untested.
No comments yet — start the discussion below.