Malyala Gayatri G. Brindha · Natural Resources for Human Health 2026 · 2026
DOI: 10.53365/nrfhh.1755
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Sequential decision-making under uncertainty has traditionally been formalized through two largely separate mathematical traditions: stochastic optimal control, which derives policies from explicit dynamics models and value functions, and reinforcement learning (RL), which learns policies directly from interaction with an environment. Generative artificial intelligence, and diffusion and flow-based models specifically, has recently emerged as a bridging technology that reframes both traditions around a shared computational substrate: treating decision-making, planning, and control-policy generation as instances of generative sequence modeling. This paper reviews the rapidly growing literature at the intersection of generative AI, stochastic optimal control, and reinforcement learning, synthesizing foundational work reformulating diffusion denoising as a sequential Markov decision process with control-theoretic treatments of stochastic optimal control via deep learning and RL-based approaches to nonlinear stochastic system stabilization. Particular attention is given to the specific mechanism connecting generative modeling to adaptive decision-making: the reviewed literature converges on the finding that framing multi-step denoising, trajectory generation, or policy sampling as a Markov decision process allows policy-gradient reinforcement learning to directly optimize non-differentiable, black-box, or human-preference-derived reward objectives that classical gradient-based diffusion training cannot address. Comparative tables map generative-RL frameworks to their algorithmic mechanism and application domain, cross-reference the specific decision-making tasks, robotic control, autonomous driving, text-to-image alignment, energy-system management, examined against the architectures and reported outcomes each study documents, and set the control-theoretic stability guarantees available for classical stochastic optimal control against the comparatively weaker guarantees available for generative policy architectures. The review concludes that while generative-AI-based decision-making frameworks have demonstrated substantial empirical gains in expressiveness and multi-modal action modeling relative to classical Gaussian policy classes, formal stability and safety guarantees comparable to those long established in stochastic optimal control theory remain the field's central unresolved research challenge.
No comments yet — start the discussion below.