Projected Primal–Dual Gradient Preference Optimization Under Category-Wise Split-Conformal Safety Constraints for Large Language Model Post-Training
Preference optimisation improves helpfulness, but a single aggregate safety penalty lets frequent harm categories dominate the objective while rare ones bear the residual risk. This paper models post-training as a vector-constrained programme: the pairwise preference loss is minimised under one risk tolerance per harm category. The solver is a projected primal–dual gradient recursion that prices each category by its own residual violation. For a nonconvex objective we bound three horizon-averaged residuals—primal stationarity, dual regret and the average constraint value—for the unbiased, unclipped recursion and a uniformly drawn iterate; the implemented recursion is compared empirically with a no-EMA, no-clipping form of itself. Certification stays separate from optimisation: on an untouched calibration split, a multiplicity-adjusted fixed-sequence test gives release thresholds valid for all categories at once under correct category assignment, with an additive bound for learned routers. It certifies the scoring and filtering of corpus-distributed candidate responses, not text the trained policy generates. In one exploratory setting—one Qwen2.5-0.5B-Instruct backbone, PKU-SafeRLHF, three seeds—the recursion meets the per-category tolerance that direct preference optimisation and scalar-constrained schemes violate, cutting the worst-category miss rate at the training threshold from 0.176 to 0.057 at unchanged preference accuracy. The certified operating point barely moves: most of the constraint’s effect is a score translation shared by the harmful and benign responses of a prompt, which calibration provably absorbs. Whether a training constraint holds and whether the deployed gate is safer are different questions.
Authors
- Yulin Tu (ORCID: https://orcid.org/0000-0002-3165-6278)
- 昊辰 姜
Institutions
- Yonsei University (KR)
- Kyung Hee University (KR)
- Dalian Polytechnic University (CN)
Publication Details
- Journal
- Mathematics
- Published
- 2026-10-08
- DOI
- https://doi.org/10.3390/math14193638
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00