Projected Primal–Dual Gradient Preference Optimization Under Category-Wise Split-Conformal Safety Constraints for Large Language Model Post-Training

Preference optimisation improves helpfulness, but a single aggregate safety penalty lets frequent harm categories dominate the objective while rare ones bear the residual risk. This paper models post-training as a vector-constrained programme: the pairwise preference loss is minimised under one risk tolerance per harm category. The solver is a projected primal–dual gradient recursion that prices each category by its own residual violation. For a nonconvex objective we bound three horizon-averaged residuals—primal stationarity, dual regret and the average constraint value—for the unbiased, unclipped recursion and a uniformly drawn iterate; the implemented recursion is compared empirically with a no-EMA, no-clipping form of itself. Certification stays separate from optimisation: on an untouched calibration split, a multiplicity-adjusted fixed-sequence test gives release thresholds valid for all categories at once under correct category assignment, with an additive bound for learned routers. It certifies the scoring and filtering of corpus-distributed candidate responses, not text the trained policy generates. In one exploratory setting—one Qwen2.5-0.5B-Instruct backbone, PKU-SafeRLHF, three seeds—the recursion meets the per-category tolerance that direct preference optimisation and scalar-constrained schemes violate, cutting the worst-category miss rate at the training threshold from 0.176 to 0.057 at unchanged preference accuracy. The certified operating point barely moves: most of the constraint’s effect is a score translation shared by the harmful and benign responses of a prompt, which calibration provably absorbs. Whether a training constraint holds and whether the deployed gate is safer are different questions.

Authors

Institutions

Publication Details

Journal
Mathematics
Published
2026-10-08
DOI
https://doi.org/10.3390/math14193638
Primary Topic
Adversarial Robustness in Machine Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Projected Primal–Dual Gradient Preference Optimization Under Category-Wise Split-Conformal Safety Constraints for Large Language Model Post-Training

Yulin Tu, 昊辰 姜
Mathematics
Adversarial Robustness in Machine Learning
article

Projected Primal–Dual Gradient Preference Optimization Under Category-Wise Split-Conformal Safety Constraints for Large Language Model Post-Training

Yulin Tu, 昊辰 姜
article en

Abstract

Preference optimisation improves helpfulness, but a single aggregate safety penalty lets frequent harm categories dominate the objective while rare ones bear the residual risk. This paper models post-training as a vector-constrained programme: the pairwise preference loss is minimised under one risk tolerance per harm category. The solver is a projected primal–dual gradient recursion that prices each category by its own residual violation. For a nonconvex objective we bound three horizon-averaged residuals—primal stationarity, dual regret and the average constraint value—for the unbiased, unclipped recursion and a uniformly drawn iterate; the implemented recursion is compared empirically with a no-EMA, no-clipping form of itself. Certification stays separate from optimisation: on an untouched calibration split, a multiplicity-adjusted fixed-sequence test gives release thresholds valid for all categories at once under correct category assignment, with an additive bound for learned routers. It certifies the scoring and filtering of corpus-distributed candidate responses, not text the trained policy generates. In one exploratory setting—one Qwen2.5-0.5B-Instruct backbone, PKU-SafeRLHF, three seeds—the recursion meets the per-category tolerance that direct preference optimisation and scalar-constrained schemes violate, cutting the worst-category miss rate at the training threshold from 0.176 to 0.057 at unchanged preference accuracy. The certified operating point barely moves: most of the constraint’s effect is a score translation shared by the harmful and benign responses of a prompt, which calibration provably absorbs. Whether a training constraint holds and whether the deployed gate is safer are different questions.

MathematicsVol. 14(19)
Yonsei University (KR), Kyung Hee University (KR), Dalian Polytechnic University (CN)
Openalex Percentile: Top 12%
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Projected Primal–Dual Gradient Preference Optimization Under Category-Wise Split-Conformal Safety Constraints for Large Language Model Post-Training — Yulin Tu, 昊辰 姜 · Mathematics (2026) | TGRS Research Map | TGRS