Rough Quantum-Torsional Preference & Verification Optimization (RQ-TVO): Super-Elevating RLVR, GRPO, and DPO
The paradigm of Large Language Model (LLM) post-training is bounded by the intrinsic limitations of Next-Token Prediction (NTP) and classical RLHF. While Direct Preference Optimization (DPO) [5], Group Relative Policy Optimization (GRPO) [6], and Reinforcement Learning with Verifiable Rewards (RLVR) offer significant advancements by eliminating the Critic network and mitigating reward hacking, they inherently rely on smooth scalar spaces. In this paper, we super-elevate this framework by embedding it within Universal Rough Operator Algebra (UROA) [3], Seonggil Theory of Complex Torsion (STCT) [1], andRough Quantum Information Geometry (RQIG) [4], leveraging foundational rough manifold bounds [2].
Authors
- Seonggil Lee
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23186256
- Primary Topic
- Reinforcement Learning in Robotics
- Type
- preprint