Reinforcement learning in continuous state and action spaces

Hasselt, H. P. (Hado) van

Reinforcement learning in continuous state and action spaces

Authors: H. P. (Hado) van Hasselt
Publication date: 1 April 2012
Publisher: Springer Berlin Heidelberg

Abstract

Many traditional reinforcement-learning algorithms have been designed for problems with small finite state and action spaces. Learning in such discrete problems can been difficult, due to noise and delayed reinforcements. However, many real-world problems have continuous state or action spaces, which can make learning a good decision policy even more involved. In this chapter we discuss how to automatically find good decision policies in continuous domains. Because analytically computing a good policy from a continuous model can be infeasible, in this chapter we mainly focus on methods that explicitly update a representation of a value function, a policy or both. We discuss considerations in choosing an appropriate representation for these functions and discuss gradient-based and gradient-free ways to update the parameters. We show how to apply these methods to reinforcement-learning problems and discuss many specific algorithms. Amongst others, we cover gradient-based temporal-difference learning, evolutionary strategies, policy-gradient algorithms and actor-critic methods. We discuss the advantages of different approaches and compare the performance of a state-of-the-art actor-critic method and a state-of-the-art evolutionary strategy empirically

Similar works

Full text

Open in the Core reader

Download PDF

Available Versions

CWI's Institutional Repository

oai:cwi.nl:19689

Last time updated on 18/04/2020