Keshav Likhar, Harsh Mandalgi, Chaturvedi Mohit · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23142444
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Operational evidence can contain the facts needed to diagnose a failure and identifiers or temporaldetails that need not be disclosed to an external model. We study diagnostic-invariant privacy-oriented transformation: changing the representation sent to a large language model while preservingthe relationships required for a root-cause-analysis (RCA) task. Three synthetic incident familiesare designed around different dependencies: identifier equality and cross-source linkage, external-reference shape, and causal temporal order. Fixed redaction, HMAC, tokenization, format-preservingpseudonymization, relative-time conversion, and temporal ablation provide controlled comparisons.A separate versioned exposure vector measures which literal, linkage, shape, cardinality, and temporalproperties remain visible; it is not an anonymity score. Development observations are directionallyconsistent with task-specific structure mattering: the format-sensitive family loses trigger/mechanismaccuracy under opaque identifiers, and paired temporal-mechanism discrimination declines from 3/4under raw evidence to 2/4 under relative time and 0/4 under temporal ablation. Evidence-supportscoring separately asks whether cited visible records actually substantiate a diagnosis. These resultsare limited to selected synthetic development cases, a small number of calls from one model family,and no formal privacy guarantee or held-out evaluation.
No comments yet — start the discussion below.