Jeonggyu Huh · arXiv (Cornell University) 2026 · 2026
DOI: 10.48550/arxiv.2609.35012
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Closed-loop planning accounts for future observation-dependent actions but can be expensive to repeat at deployment. Bellman-gradient (BG) refinement differentiates conditional rollouts through an existing actor, corrects the current action, and stores the result in an executable policy. A backward sweep reuses deployed future feedback; full-horizon rollouts in an identified Gaussian belief model need no learned value critic. A nonlinear error recursion links quadratic continuation error, local correction, and policy storage. A controlled non-LQG example exhibits second-order policy accuracy with consistent storage, while ideal affine LQG admits exact backward recovery. In nonquadratic thrust, BG reduces actor cost by 2.47-3.20% and remains within 0.13-0.53% of the tested feedback MPC; original policies execute in 5-6 microseconds in a seed-0 native audit. With matched storage, BG attains competitive plant costs at about one ninth (arm) and one sixtieth (docking) of feedback-teacher-plus-student construction time on one GPU, excluding shared learning and node preparation. Richer common maps substantially narrow some student-BG gaps. These results expose the roles of learned feedback, local correction, and storage in executable control.
No comments yet — start the discussion below.