Zen Revista, 10 IA · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22957211
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This article offers a historical development review of reinforcement learning, the branch of machine learning that trains agents to act by the consequences of their actions, and whose development from the checkers programs of the 1950s through the temporal difference revolution, the Q-learning consolidation, and the deep learning merger to the superhuman game-playing systems of the 2010s is the story of a simple hypothesis, that reward prediction learning suffices for intelligence's acquisition, approaching its empirical test. The review reconstructs the development across five phases: the early programs, Samuel's checkers learner whose self-play evaluation function learning anticipated the field's architecture, and the dynamic programming's and trial-and-error psychology's foundations; the temporal difference revolution, Sutton's TD methods and Watkins's Q-learning, whose prediction-by-difference and off-policy convergence composed the field's mathematical core; the consolidation era, the Barto-Sutton textbook's canonization, the policy gradient family's development, and the field's Markov decision process formalization; the deep merger, the DQN's 2015 Nature result whose convolutional network learned Atari from pixels, and the AlphaGo and AlphaZero lineage whose self-play superhumanity the Go, chess, and shogi domains confirmed; and the present frontier, the multi-agent and the non-stationary programs, the sample efficiency's and the reward specification's open problems, and the extension from games to the balloons, chips, and language agents whose applications the field's industrial programs pursue. Three synthetic claims are advanced. First, reinforcement learning's history is the redemption of an idea the AI mainstream twice abandoned, the reward hypothesis whose early-era marginalization the field's later triumphs reversed, and the redemption's lesson is that the marginality's patience, the subfield's two decades of mathematical preparation before the deep merger, is the discipline's structure of readiness. Second, the deep merger's meaning is architectural rather than conceptual, since the deep networks supplied the function approximation whose absence the classical era's scaling failures reflected, and the merger's recipe, experience replay and target networks stabilizing the bootstrapping, is engineering's contribution to the idea's vindication. Third, the reward specification problem, the design of objectives whose pursuit the agent's competence misgeneralizes, has become the field's central difficulty as the agents' competence grew, and the problem's study, the reward hacking and the alignment research it seeded, connects the field's technical history to the safety program's agenda. An agenda is proposed spanning generalization, sample efficiency, and the alignment frontier.
No comments yet — start the discussion below.