Pairwise Classification as a Unified Framework for Offline Reinforcement Learning and Large-Language-Model Alignment
Offline reinforcement learning (offline RL) and large-language-model (LLM) alignment are typically studied as independent domains and each has developed its own pairwise comparison technique. Prior studies have validated pairwise classification exclusively within offline RL, leaving open the question of whether it generalizes beyond a single domain. This study reconceptualizes pairwise classification not as a technique confined to offline RL, but as a domain-agnostic structural principle for determining which of the two candidates scores higher when a context is shared. We instantiate this principle as a pairwise-Decision Transformer (pairwise-DT), which extends the Decision Transformer (DT) by comparing the Q-values of two transitions in offline RL, and pairwise fine-tuning (pairwise-FT), which compares the preference scores of two responses in the LLM alignment by applying the same pairwise loss across both domains. Through experiments on three MuJoCo environments and the anthropic Helpful and Harmless Reinforcement Learning from Human Feedback (HH-RLHF) dataset, pairwise-FT achieved higher reward accuracy than direct-preference optimization (DPO) across five seeds in approximately half the training time, without a reference model. The pairwise DT outperformed the standard DT in all three MuJoCo environments. Sensitivity analysis over a range of temperature values confirms that this advantage is not an artifact of a favorably chosen hyperparameter; this robustness holds strongly on the LLM side and, to a lesser extent, on the MuJoCo side. On the LLM side, this advantage reflects preference-scoring/reward-modeling performance specifically, rather than full generation-quality alignment.
Authors
- Chayoung Kim (ORCID: https://orcid.org/0000-0002-4186-5882)
Institutions
- Hankyong National University (KR)
Publication Details
- Journal
- Applied Sciences
- Published
- 2026-09-14
- DOI
- https://doi.org/10.3390/app16189098
- Primary Topic
- Reinforcement Learning in Robotics
- Type
- article
- Field-Weighted Citation Impact
- 0.00