Liuhaichen Yang, Hanshang Zhu, Zhengyang Zhong, Xinyu Tan, Ningwei Bai, Yunqi Huang, Hanbo Ma, Zhengtao Ding, Zezhi Tang · Unmanned Systems 2026 · 2026
DOI: 10.1142/s2301385028300065
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Reinforcement learning (RL) has become a central post-training approach for reasoning and agentic large language models (LLMs), particularly when task outcomes can be verified automatically. Comparisons across this literature remain difficult because a reported gain may combine changes to the learning signal, policy constraint, response granularity, data reuse, rollout system, and inference budget. This survey develops a mechanism-driven framework for separating these effects. It organizes methods along three axes— learning-signal type, policy/data regime, and optimization granularity—and decomposes their training stacks into reusable motifs spanning advantage construction, drift control, dense feedback, off-policy reuse, and rollout–update dataflow. We use this framework to synthesize representative method families, analyze cross-motif interactions, and conduct source-bounded case studies of how multicomponent recipes and asynchronous pipelines should be interpreted. We further distinguish literature-established diagnostics from survey-defined reporting primitives and provide a matched-budget reporting checklist. Finally, we discuss how these interfaces may transfer to vision-language-action (VLA) models, World Action Models (WAMs), and embodied-agent post-training for unmanned systems, where reports should describe action representation, world-model error, interaction cost, and safety constraints. The corpus emphasizes mechanism clarity and records the maturity of evidence from recent preprints and system reports.
No comments yet — start the discussion below.