Wei Zhu, Jinyin Bai, Rui Tang, Zehao Pang, Mingxi Wang, Chengjie Lu, Tianjin Ni, Hang Liu, Xiangchen Wang, Jinji Zhou, Yanlin Wu, Yongjun Peng, Zongzhe Nie, Shiluo Guo, Qinglin Xu, KaiYang Kou, Yihao Zhong · Preprints.org 2026 · 2026
DOI: 10.20944/preprints202608.0779.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Reward functions determine what reinforcement learning agents ultimately optimize, yet reward design for complex tasks has traditionally relied on extensive domain expertise and iterative engineering. Recent large language models and vision–language foundation models have introduced new mechanisms for interpreting task intent, synthesizing reward programs, evaluating states and trajectories, and refining rewards through policy feedback. This review organizes the emerging literature along three complementary directions: reward program synthesis, multimodal feedback, and feedback-driven reward optimization. We further propose a five-level trustworthiness framework spanning format validity, execution validity, semantic validity, behavioral validity, and structural assurance. Existing evidence shows that foundation models substantially broaden how rewards can be represented and acquired, but do not eliminate grounding errors, proxy misalignment, reward hacking, selection bias, or reward-search costs. We therefore examine the field from an end-to-end perspective that jointly considers policy performance, reward fidelity, trustworthiness, computational and human cost, and transfer. Finally, we identify verifiable reward representations, process reward models, budget-aware reward search, and transferable reward knowledge across tasks and multi-agent systems as key directions for future research.
No comments yet — start the discussion below.