What is reinforcement learning?
Reinforcement learning is one of the core approaches used for training AI systems. Unlike supervised learning, which relies on labeled examples, or unsupervised learning, which identifies patterns in data without guidance, this approach trains a model through trial and error, using feedback from its own actions.
Like a person learning a new skill, an AI system trained this way improves by receiving rewards for successful outcomes and penalties for less successful outcomes — adjusting its future actions accordingly. For example, a robot might be rewarded for taking fewer steps to reach an object; over many trials, it learns to achieve its goal based on experience rather than explicit pre-calculated instructions.
The foundations of learning under uncertainty
Early NSF-supported advances in mathematics, decision-making, adaptive learning and neuroscience helped lay the foundations of reinforcement learning.
The mathematics of machine learning
Beginning in the 1950s, NSF supported foundational research in probability theory and random (stochastic) processes at universities across the United States, laying the mathematical groundwork for future advances in machine learning and AI.
At the same time, researchers at the RAND Corporation, a think tank, developed a mathematical method called dynamic programming that showed how complex decisions could be broken into smaller steps, helping machines make better choices in uncertain situations. This mathematical framework became one of the most influential advances in modern decision science and remains a cornerstone of reinforcement learning and many other fields.
Throughout the 1960s and 1970s, NSF supported research in adaptive systems, control theory, computer science and applied mathematics that examined how decisions can be optimized through feedback and experience. Together, this work created clear, reliable ways to understand and work with randomness, including tools such as Markov processes, which describe how situations change step by step according to probabilities.
By advancing mathematical tools for understanding uncertainty, feedback and decision-making, this work helped pave the way for the emergence of reinforcement learning as a distinct field of AI.
Early demonstrations of learning from experience
NSF-supported research in the 1980s advanced the understanding of how intelligent systems can improve their performance over time. In one early example, researchers trained computers to balance a pole on a moving cart through repeated feedback with rewards and penalties.
This work helped establish key principles that would later define reinforcement learning, showing that computers could improve their decisions by learning which actions lead to successful outcomes.
Credit: Kiel Mutschelknaus, Columbia University
Bridging brains and machines
By the 1990s, researchers were connecting reinforcement learning theory with discoveries about how the brain learns from experience. Building on NSF-supported advances in temporal-difference learning — a method that improves predictions by comparing expected and actual outcomes over time — studies showed that the same reward-based computational principles used in reinforcement learning algorithms also help explain how dopamine neurons respond to rewards and prediction errors in the brain.
Revealing shared principles of learning in brains and machines deepened scientific understanding of intelligence and guided future advances in AI.
Credit: University of Washington
Turning discoveries into a foundational field guide
As research on decision-making and adaptation advanced, NSF-supported research at the Autonomous Learning Laboratory explored how learning through experience gives rise to increasingly complex behaviors. This cross-disciplinary research facility investigated sensorimotor learning: the process of linking sensory information with movement and actions through experience.
Researchers studied how this type of learning can give rise to higher cognitive abilities in animals, humans and robots. For example, researchers studied how robots could learn to navigate their surroundings and adapt their movements through experience, much like a child learning to walk or reach for an object.
Pulling from this work, researchers at the lab published Reinforcement Learning: An Introduction in 1998, with a second edition in 2018. Written by Richard Sutton and Andrew Barto, the book became one of the field's defining texts, helping generations of researchers understand how computational systems can learn to optimize decisions through feedback. Professors Sutton and Barto shared the 2024 A.M. Turing Award.
Entering the mainstream
Credit: Lee Sedol via Wikimedia (CC BY 4.0)
Decades of NSF investments in foundational research helped transform early theories of learning into powerful tools that drive modern AI technologies.
A major milestone came in 2016 with the development of AlphaGo. This AI system was designed to play the ancient strategy board game Go, a game known for its enormous number of possible moves and complexity. AlphaGo showed the world how reinforcement learning, combined with deep learning and large-scale computing, could enable machines to make sophisticated, human-like decisions, paving the way for many of today's intelligent technologies.
Today, reinforcement learning powers a wide range of technologies used across daily life, including:
Conversational AI
Large language models can be refined from human feedback, helping AI assistants produce responses aligned with human preferences.
Robotics
Robots can combine visual information, language instructions and experience to learn new tasks and adapt their actions over time.
Chip design
AI systems can optimize the layout of microprocessors, improving their performance and efficiency.
Personalized recommendations
Streaming and shopping services like Netflix and YouTube can tailor suggestions to users' tastes.
Smart-home technology
Devices such as thermostats and voice assistants adapt to user habits, improving comfort, energy efficiency and automation over time.
Smart manufacturing
Intelligent manufacturing systems can learn from data and experience to improve efficiency, reliability and adaptability.
Autonomous vehicles
Self-driving systems learn to navigate safely through ever-changing traffic conditions.
Supply-chain optimization
Companies can predict demand and move goods more efficiently.
Education
Educational software can adapt lessons and feedback to individual students, helping personalize learning and improve outcomes.
Investing in the future of adaptive, intelligent technologies
The NSF Energy, Power, Control, and Learning program provided early support for the pioneering reinforcement learning research of Andrew Barto and Richard Sutton and continues to fund foundational advances in learning and adaptive systems.
Today, NSF is expanding its investment in AI through its NSF-led National AI Research Institutes, bringing together universities, industry and government to tackle major challenges in the field.
Complemented by programs such as NSF Robust Intelligence, NSF Foundational Research in Robotics, NSF Collaborative Research in Computational Neuroscience, NSF Mathematical Foundations of Artificial Intelligence and NSF Manufacturing Systems Integration, these efforts support cross-disciplinary teams exploring how systems can learn, reason and work together to achieve desired goals.