Who Cited It

Reinforcement Learning: A Survey

1996 · Journal of Artificial Intelligence Research · 8,953 citations · 12 from inside this corpus

Leslie Pack Kaelbling low, Michael L. Littman, Andrew Moore

This paper surveys the field of reinforcement learning from a computer-science perspective. It is written to be accessible to researchers familiar with machine learning. Both the historical basis of the field and a broad selection of current work are summarized. Reinforcement learning is the problem faced by an agent that learns behavior through trial-and-error interactions with a dynamic environment. The work described here has a resemblance to work in psychology, but differs considerably in the details and in the use of the word ``reinforcement.'' The paper discusses central issues of reinforcement learning, including trading off exploration and exploitation, establishing the foundations of the field via Markov decision theory, learning from delayed reinforcement, constructing empirical models to accelerate learning, making use of generalization and hierarchy, and coping with hidden state. It concludes with a survey of some implemented systems and an assessment of the practical utility of current methods for reinforcement learning.

Reinforcement Learning: A Survey (1996)Reinforcement Learning: A Sur…Genetic algorithms in search, optimization, and machine learning (1989)Genetic algorithms in search,…Adaptation in Natural and Artificial Systems (1992)Adaptation in Natural and Art…Genetic Algorithms in Search, Optimization and Machine Learning (1988)Genetic Algorithms in Search,…Q-learning (1992)Q-learningMarkov Decision Processes: Discrete Stochastic Dynamic Programming. (1995)Markov Decision Processes: Di…Simple statistical gradient-following algorithms for connectionist reinforcement learning (1992)Simple statistical gradient-f…Learning from Delayed Rewards (1989)Learning from Delayed RewardsSome Studies in Machine Learning Using the Game of Checkers (1959)Some Studies in Machine Learn…Learning to Predict by the Methods of Temporal Differences (1988)Learning to Predict by the Me…Dynamic Programming and Markov Processes. (1961)Dynamic Programming and Marko…Markov games as a framework for multi-agent reinforcement learning (1994)Markov games as a framework f…A New Approach to Manipulator Control: The Cerebellar Model Articulation Controller (CMAC) (1975)A New Approach to Manipulator…Markov Decision Processes: Discrete Stochastic Dynamic Programming (1995)Markov Decision Processes: Di…Learning Automata: An Introduction (1989)Learning Automata: An Introdu…Some studies in machine learning using the game of checkers (2000)Some studies in machine learn…Classifier Fitness Based on Accuracy (1995)Classifier Fitness Based on A…Integrated Architectures for Learning, Planning, and Reacting Based on Approximating Dyna… (1990)Integrated Architectures for …Deep learning in neural networks: An overview (2014)Deep learning in neural netwo…Ant colony system: a cooperative learning approach to the traveling salesman problem (1997)Ant colony system: a cooperat…Deep Reinforcement Learning with Double Q-Learning (2016)Deep Reinforcement Learning w…Machine Learning: Algorithms, Real-World Applications and Research Directions (2021)Machine Learning: Algorithms,…Deep Reinforcement Learning with Double Q-Learning (2016)Deep Reinforcement Learning w…Evolving Neural Networks through Augmenting Topologies (2002)Evolving Neural Networks thro…Reinforcement learning in robotics: A survey (2013)Reinforcement learning in rob…Deep Learning: A Comprehensive Overview on Techniques, Taxonomy, Applications and Researc… (2021)Deep Learning: A Comprehensiv…A Comprehensive Survey of Multiagent Reinforcement Learning (2008)A Comprehensive Survey of Mul…Transfer Learning for Reinforcement Learning Domains: A Survey (2009)Transfer Learning for Reinfor…A Survey on Deep Learning for Named Entity Recognition (2020)A Survey on Deep Learning for…Adaptive Computation and Machine Learning (2012)Adaptive Computation and Mach…
29 of 29 neighbouring works in this corpus. Blue is what this paper cites; orange is what cites it, and a dashed line is one neighbour citing another. Only the largest labels are drawn — every node carries its full title on hover.
this paper works it cites works citing it node size = global citations · hover for the full title

What this paper cites, inside the corpus

What cites it, inside the corpus

Topics

Reinforcement Learning in RoboticsComputer Science
Evolutionary Algorithms and ApplicationsComputer Science
Metaheuristic Optimization Algorithms ResearchComputer Science

Is this record sound?

complete

Nothing in this record contradicts itself and no field we check is missing.

  • supports3 author record(s) attached.
  • supports189 reference(s) recorded.
  • neutralThe DOI carries no year to check against.
  • supportsA title is present.

Provenance

Everything above was read from one stored OpenAlex payload, fetched 2026-09-04T03:58:43+00:00.

sha256 5cad55ac5d4d41e1…