A three panel image: On left, a robot arm attached to a chair picks up some fruit with a fork. At center, a close-up of hands typing on a smart phone. At right, a computer screen displaying chemical structures.

Reinforcement Learning: The Engine Powering Today's AI Revolution

NSF-funded research helped lay the foundation for AI systems powering generative AI, robotics and autonomous technologies.

How do machines learn to walk, drive cars or chat like humans? These advances rely on a range of artificial intelligence techniques, including reinforcement learning, a branch of AI inspired by the way animals and humans learn from positive and negative feedback.

Decades before chatbots and autonomous robots captured public attention, the U.S. National Science Foundation supported pioneering research that helped build the field of reinforcement learning from the ground up.

NSF's visionary investments helped turn ambitious research into one of the foundations of modern AI, advancing technologies that drive economic growth, shape new industries and inspire a new generation of talent.

What is reinforcement learning?

Reinforcement learning is one of the core approaches used for training AI systems. Unlike supervised learning, which relies on labeled examples, or unsupervised learning, which identifies patterns in data without guidance, this approach trains a model through trial and error, using feedback from its own actions.

Like a person learning a new skill, an AI system trained this way improves by receiving rewards for successful outcomes and penalties for less successful outcomes — adjusting its future actions accordingly. For example, a robot might be rewarded for taking fewer steps to reach an object; over many trials, it learns to achieve its goal based on experience rather than explicit pre-calculated instructions.

Learn more about artificial intelligence.

Image:
Animated grid-world showing a reinforcement learning agent learning through trial and error. Early in training it wanders and hits hazards; by the end it efficiently avoids hazards and reaches the goal.

A reinforcement learning agent learns through trial and error. As it gains experience, it learns to avoid hazards and reach its goal more efficiently.

U.S. National Science Foundation

The foundations of learning under uncertainty

Early NSF-supported advances in mathematics, decision-making, adaptive learning and neuroscience helped lay the foundations of reinforcement learning.

Image:
Animation of a Markov process with states connected by arrows labeled with probabilities, illustrating how a system changes from one state to another over time.

A Markov process models how a system changes from one state to another according to probabilities. These mathematical models became an important foundation for reinforcement learning and many other AI methods.

U.S. National Science Foundation

The mathematics of machine learning

Beginning in the 1950s, NSF supported foundational research in probability theory and random (stochastic) processes at universities across the United States, laying the mathematical groundwork for future advances in machine learning and AI.

At the same time, researchers at the RAND Corporation, a think tank, developed a mathematical method called dynamic programming that showed how complex decisions could be broken into smaller steps, helping machines make better choices in uncertain situations. This mathematical framework became one of the most influential advances in modern decision science and remains a cornerstone of reinforcement learning and many other fields.

Throughout the 1960s and 1970s, NSF supported research in adaptive systems, control theory, computer science and applied mathematics that examined how decisions can be optimized through feedback and experience. Together, this work created clear, reliable ways to understand and work with randomness, including tools such as Markov processes, which describe how situations change step by step according to probabilities.

By advancing mathematical tools for understanding uncertainty, feedback and decision-making, this work helped pave the way for the emergence of reinforcement learning as a distinct field of AI.

Image:
A short video showing a small pole balanced on a cart, moving back and forth across a track. Early in the video, the pole does not balance well. Late in the video the pole remains balanced even when disturbed by someone's hand.

The "cart-pole" task is a classic demonstration of reinforcement learning. The AI agent's objective is to balance a pole on a cart by moving the cart left or right along a track.

Cheng-Yueh Liu via YouTube/CC by 4.0

Early demonstrations of learning from experience

NSF-supported research in the 1980s advanced the understanding of how intelligent systems can improve their performance over time. In one early example, researchers trained computers to balance a pole on a moving cart through repeated feedback with rewards and penalties.

This work helped establish key principles that would later define reinforcement learning, showing that computers could improve their decisions by learning which actions lead to successful outcomes.

Visualization of the AI Institute ARNI
In the 1990s, researchers found that brains and artificial networks use similar algorithms to learn through trial and error.

Credit: Kiel Mutschelknaus, Columbia University

Bridging brains and machines

By the 1990s, researchers were connecting reinforcement learning theory with discoveries about how the brain learns from experience. Building on NSF-supported advances in temporal-difference learning — a method that improves predictions by comparing expected and actual outcomes over time — studies showed that the same reward-based computational principles used in reinforcement learning algorithms also help explain how dopamine neurons respond to rewards and prediction errors in the brain.

Revealing shared principles of learning in brains and machines deepened scientific understanding of intelligence and guided future advances in AI.

An assistive-feeding robot skewers a piece of fruit during a demonstration.
Reinforcement learning has been key to helping robotics, like this robotic arm designed to assist those with upper-body mobility impairments, perform complex tasks.

Credit: University of Washington

Turning discoveries into a foundational field guide

As research on decision-making and adaptation advanced, NSF-supported research at the Autonomous Learning Laboratory explored how learning through experience gives rise to increasingly complex behaviors. This cross-disciplinary research facility investigated sensorimotor learning: the process of linking sensory information with movement and actions through experience.

Researchers studied how this type of learning can give rise to higher cognitive abilities in animals, humans and robots. For example, researchers studied how robots could learn to navigate their surroundings and adapt their movements through experience, much like a child learning to walk or reach for an object.

Pulling from this work, researchers at the lab published Reinforcement Learning: An Introduction in 1998, with a second edition in 2018. Written by Richard Sutton and Andrew Barto, the book became one of the field's defining texts, helping generations of researchers understand how computational systems can learn to optimize decisions through feedback. Professors Sutton and Barto shared the 2024 A.M. Turing Award.

Entering the mainstream

An illustration of a game of Go
The 2016 match between AlphaGo and champion Go player Lee Sedol is considered a watershed moment for AI.

Credit: Lee Sedol via Wikimedia (CC BY 4.0)

Decades of NSF investments in foundational research helped transform early theories of learning into powerful tools that drive modern AI technologies.

A major milestone came in 2016 with the development of AlphaGo. This AI system was designed to play the ancient strategy board game Go, a game known for its enormous number of possible moves and complexity. AlphaGo showed the world how reinforcement learning, combined with deep learning and large-scale computing, could enable machines to make sophisticated, human-like decisions, paving the way for many of today's intelligent technologies.

Today, reinforcement learning powers a wide range of technologies used across daily life, including:

Conversational AI

Large language models can be refined from human feedback, helping AI assistants produce responses aligned with human preferences.

Robotics

Robots can combine visual information, language instructions and experience to learn new tasks and adapt their actions over time.

Chip design

AI systems can optimize the layout of microprocessors, improving their performance and efficiency.

Personalized recommendations

Streaming and shopping services like Netflix and YouTube can tailor suggestions to users' tastes.

Smart-home technology

Devices such as thermostats and voice assistants adapt to user habits, improving comfort, energy efficiency and automation over time.

Smart manufacturing

Intelligent manufacturing systems can learn from data and experience to improve efficiency, reliability and adaptability.

Autonomous vehicles

Self-driving systems learn to navigate safely through ever-changing traffic conditions.

Supply-chain optimization

Companies can predict demand and move goods more efficiently.

Education

Educational software can adapt lessons and feedback to individual students, helping personalize learning and improve outcomes.

Investing in the future of adaptive, intelligent technologies

The NSF Energy, Power, Control, and Learning program provided early support for the pioneering reinforcement learning research of Andrew Barto and Richard Sutton and continues to fund foundational advances in learning and adaptive systems.

Today, NSF is expanding its investment in AI through its NSF-led National AI Research Institutes, bringing together universities, industry and government to tackle major challenges in the field.

Complemented by programs such as NSF Robust Intelligence, NSF Foundational Research in Robotics, NSF Collaborative Research in Computational Neuroscience, NSF Mathematical Foundations of Artificial Intelligence and NSF Manufacturing Systems Integration, these efforts support cross-disciplinary teams exploring how systems can learn, reason and work together to achieve desired goals.