DiagPriv: Preserving Diagnostic Structure While Reducing Sensitive Data Exposure in LLM Root-Cause Analysis
Operational evidence can contain the facts needed to diagnose a failure and identifiers or temporaldetails that need not be disclosed to an external model. We study diagnostic-invariant privacy-oriented transformation: changing the representation sent to a large language model while preservingthe relationships required for a root-cause-analysis (RCA) task. Three synthetic incident familiesare designed around different dependencies: identifier equality and cross-source linkage, external-reference shape, and causal temporal order. Fixed redaction, HMAC, tokenization, format-preservingpseudonymization, relative-time conversion, and temporal ablation provide controlled comparisons.A separate versioned exposure vector measures which literal, linkage, shape, cardinality, and temporalproperties remain visible; it is not an anonymity score. Development observations are directionallyconsistent with task-specific structure mattering: the format-sensitive family loses trigger/mechanismaccuracy under opaque identifiers, and paired temporal-mechanism discrimination declines from 3/4under raw evidence to 2/4 under relative time and 0/4 under temporal ablation. Evidence-supportscoring separately asks whether cited visible records actually substantiate a diagnosis. These resultsare limited to selected synthetic development cases, a small number of calls from one model family,and no formal privacy guarantee or held-out evaluation.
Authors
- Chaturvedi Mohit
- Harsh Mandalgi (ORCID: https://orcid.org/0009-0000-9956-3636)
- Keshav Likhar (ORCID: https://orcid.org/0009-0000-1473-8274)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-04
- DOI
- https://doi.org/10.5281/zenodo.23142445
- Primary Topic
- Privacy-Preserving Technologies in Data
- Type
- preprint