Privacy-Constrained Distributionally Robust Detection and Collaborative Attribution of LLM-Generated Text
With the widespread use of large language models, distinguishing human-written text from LLM-generated text has become increasingly important for content authenticity, academic integrity, and digital forensics. However, existing detectors remain vulnerable to paraphrase attacks, domain shift, and privacy restrictions that prevent institutions from sharing raw text data. To address these challenges, this paper proposes FedRAT, a privacy-preserving federated framework that couples paired paraphrase consistency, generator-family auxiliary supervision, differentially private client updates, and loss-dependent aggregation within a unified detection setting. FedRAT jointly learns binary AI-text detection and weak attribution to a predefined generator family through a shared encoder. Experiments on multi-domain and multi-generator benchmarks show that FedRAT consistently outperforms local training, standard federated learning, paraphrase-augmented federated learning, and representative detection baselines under clean and paraphrased settings. The ablation results evaluate attribution learning, paraphrase consistency, risk-aware aggregation, and differential privacy, while Expected Calibration Error and Brier Score provide diagnostic measures of confidence quality. The results support FedRAT as an empirically evaluated integration for privacy-constrained LLM-generated text detection and coarse generator-family attribution under the tested rewriting conditions.
Authors
- Zihao Zhang (ORCID: https://orcid.org/0009-0006-4208-462X)
Institutions
- Twitter (United States) (US)
Publication Details
- Journal
- International Journal of Pattern Recognition and Artificial Intelligence
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1142/s0218001426400616
- Primary Topic
- Authorship Attribution and Profiling
- Type
- article
- Field-Weighted Citation Impact
- 0.00