Reinforcement Learning Post-Training for Reasoning Large Language Models: Methods, Systems, and Evaluation
Reinforcement learning (RL) has become a central post-training approach for reasoning and agentic large language models (LLMs), particularly when task outcomes can be verified automatically. Comparisons across this literature remain difficult because a reported gain may combine changes to the learning signal, policy constraint, response granularity, data reuse, rollout system, and inference budget. This survey develops a mechanism-driven framework for separating these effects. It organizes methods along three axes— learning-signal type, policy/data regime, and optimization granularity—and decomposes their training stacks into reusable motifs spanning advantage construction, drift control, dense feedback, off-policy reuse, and rollout–update dataflow. We use this framework to synthesize representative method families, analyze cross-motif interactions, and conduct source-bounded case studies of how multicomponent recipes and asynchronous pipelines should be interpreted. We further distinguish literature-established diagnostics from survey-defined reporting primitives and provide a matched-budget reporting checklist. Finally, we discuss how these interfaces may transfer to vision-language-action (VLA) models, World Action Models (WAMs), and embodied-agent post-training for unmanned systems, where reports should describe action representation, world-model error, interaction cost, and safety constraints. The corpus emphasizes mechanism clarity and records the maturity of evidence from recent preprints and system reports.
Authors
- Zezhi Tang (ORCID: https://orcid.org/0000-0002-0182-6010)
- Liuhaichen Yang (ORCID: https://orcid.org/0009-0008-4353-3612)
- Yunqi Huang
- Ningwei Bai
- Hanbo Ma
- Zhengyang Zhong
- Hanshang Zhu
- Xinyu Tan
- Zhengtao Ding
Publication Details
- Journal
- Unmanned Systems
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1142/s2301385028300065
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00