Father of Reinforcement Learning · Turing Award 2024Father of Reinforcement Learning · Turing Award 2024
rlagents
Born 1956 in the USA. Earned a BA in Psychology from Stanford (1978), then a PhD in Computer Science from the University of Massachusetts Amherst (1984) under Andrew Barto -- together they founded the field of reinforcement learning. Joined the University of Alberta (2003) where he became a Distinguished Research Professor and iCORE Chair. Co-authored 'Reinforcement Learning: An Introduction' (1998, 2nd ed. 2018) with Barto -- THE standard textbook that trained virtually every RL researcher alive. Invented temporal-difference (TD) learning, policy gradient methods, and the Dyna architecture. Served as Distinguished Research Scientist at DeepMind Alberta. His 2019 essay 'The Bitter Lesson' -- arguing that general methods leveraging computation always win over hand-crafted approaches -- became one of the most cited philosophical pieces in modern AI. Shared the 2024 ACM Turing Award with Barto. Remains RL's contrarian conscience: argues LLMs lack true world understanding, that intelligence must be grounded in runtime experience ('The Era of Experience', 2025, with David Silver), and keeps pursuing a simple general agent architecture (the Alberta Plan).
Born 1956 in the USA. Earned a BA in Psychology from Stanford (1978), then a PhD in Computer Science from the University of Massachusetts Amherst (1984) under Andrew Barto -- together they founded the field of reinforcement learning. Joined the University of Alberta (2003) where he became a Distinguished Research Professor and iCORE Chair. Co-authored 'Reinforcement Learning: An Introduction' (1998, 2nd ed. 2018) with Barto -- THE standard textbook that trained virtually every RL researcher alive. Invented temporal-difference (TD) learning, policy gradient methods, and the Dyna architecture. Served as Distinguished Research Scientist at DeepMind Alberta. His 2019 essay 'The Bitter Lesson' -- arguing that general methods leveraging computation always win over hand-crafted approaches -- became one of the most cited philosophical pieces in modern AI. Shared the 2024 ACM Turing Award with Barto. Remains RL's contrarian conscience: argues LLMs lack true world understanding, that intelligence must be grounded in runtime experience ('The Era of Experience', 2025, with David Silver), and keeps pursuing a simple general agent architecture (the Alberta Plan).