Reinforcement Learning Post-Training for Reasoning Large Language Models: Methods, Systems, and Evaluation

Reinforcement learning (RL) has become a central post-training approach for reasoning and agentic large language models (LLMs), particularly when task outcomes can be verified automatically. Comparisons across this literature remain difficult because a reported gain may combine changes to the learning signal, policy constraint, response granularity, data reuse, rollout system, and inference budget. This survey develops a mechanism-driven framework for separating these effects. It organizes methods along three axes— learning-signal type, policy/data regime, and optimization granularity—and decomposes their training stacks into reusable motifs spanning advantage construction, drift control, dense feedback, off-policy reuse, and rollout–update dataflow. We use this framework to synthesize representative method families, analyze cross-motif interactions, and conduct source-bounded case studies of how multicomponent recipes and asynchronous pipelines should be interpreted. We further distinguish literature-established diagnostics from survey-defined reporting primitives and provide a matched-budget reporting checklist. Finally, we discuss how these interfaces may transfer to vision-language-action (VLA) models, World Action Models (WAMs), and embodied-agent post-training for unmanned systems, where reports should describe action representation, world-model error, interaction cost, and safety constraints. The corpus emphasizes mechanism clarity and records the maturity of evidence from recent preprints and system reports.

Authors

Publication Details

Journal
Unmanned Systems
Published
2026-09-30
DOI
https://doi.org/10.1142/s2301385028300065
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Reinforcement Learning Post-Training for Reasoning Large Language Models: Methods, Systems, and Evaluation

Zezhi Tang, Liuhaichen Yang, Yunqi Huang, Ningwei Bai et al.
Unmanned Systems
Multimodal Machine Learning Applications
article

Reinforcement Learning Post-Training for Reasoning Large Language Models: Methods, Systems, and Evaluation

Zezhi Tang, Liuhaichen Yang, Yunqi Huang, Ningwei Bai, Hanbo Ma, Zhengyang Zhong, Hanshang Zhu, Xinyu Tan, Zhengtao Ding
article en

Abstract

Reinforcement learning (RL) has become a central post-training approach for reasoning and agentic large language models (LLMs), particularly when task outcomes can be verified automatically. Comparisons across this literature remain difficult because a reported gain may combine changes to the learning signal, policy constraint, response granularity, data reuse, rollout system, and inference budget. This survey develops a mechanism-driven framework for separating these effects. It organizes methods along three axes— learning-signal type, policy/data regime, and optimization granularity—and decomposes their training stacks into reusable motifs spanning advantage construction, drift control, dense feedback, off-policy reuse, and rollout–update dataflow. We use this framework to synthesize representative method families, analyze cross-motif interactions, and conduct source-bounded case studies of how multicomponent recipes and asynchronous pipelines should be interpreted. We further distinguish literature-established diagnostics from survey-defined reporting primitives and provide a matched-budget reporting checklist. Finally, we discuss how these interfaces may transfer to vision-language-action (VLA) models, World Action Models (WAMs), and embodied-agent post-training for unmanned systems, where reports should describe action representation, world-model error, interaction cost, and safety constraints. The corpus emphasizes mechanism clarity and records the maturity of evidence from recent preprints and system reports.

Unmanned Systems
Openalex Percentile: Top 14%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.