Search CORE

122,385 research outputs found

Distral: Robust Multitask Reinforcement Learning

Author: Bapst Victor
Czarnecki Wojciech Marian
Hadsell Raia
Heess Nicolas
Kirkpatrick James
Pascanu Razvan
Quan John
Teh Yee Whye
Publication venue
Publication date: 01/01/2017
Field of study

Most deep reinforcement learning algorithms are data inefficient in complex and rich environments, limiting their applicability to many scenarios. One direction for improving data efficiency is multitask learning with shared neural network parameters, where efficiency may be improved through transfer across related tasks. In practice, however, this is not usually observed, because gradients from different tasks can interfere negatively, making learning unstable and sometimes even less data efficient. Another issue is the different reward schemes between tasks, which can easily lead to one task dominating the learning of a shared model. We propose a new approach for joint training of multiple tasks, which we refer to as Distral (Distill & transfer learning). Instead of sharing parameters between the different workers, we propose to share a "distilled" policy that captures common behaviour across tasks. Each worker is trained to solve its own task while constrained to stay close to the shared policy, while the shared policy is trained by distillation to be the centroid of all task policies. Both aspects of the learning process are derived by optimizing a joint objective function. We show that our approach supports efficient transfer on complex 3D environments, outperforming several related methods. Moreover, the proposed learning process is more robust and more stable---attributes that are critical in deep reinforcement learning

arXiv.org e-Print Archive

Oxford University Research Archive

Explore, Exploit or Listen: Combining Human Feedback and Policy Model to Speed up Deep Reinforcement Learning in 3D Worlds

Author: Ashley Haines (4342138)
David Gauthier (1893841)
David Taylor (140886)
Donald Landers (4342165)
Hamish Small (4342156)
Jeffrey Shields (4342159)
John Hoenig (4342153)
John Swenarton (4342132)
Mark Matsche (4342147)
Matthew Smith (326340)
Maya Groner (459554)
Philip Sadler (4342126)
Roger Pradel (111017)
Rémi Choquet (111021)
Wolfgang Vogelbein (4342162)
Publication venue
Publication date: 14/06/2017
Field of study

We describe a method to use discrete human feedback to enhance the performance of deep learning agents in virtual three-dimensional environments by extending deep-reinforcement learning to model the confidence and consistency of human feedback. This enables deep reinforcement learning algorithms to determine the most appropriate time to listen to the human feedback, exploit the current policy model, or explore the agent's environment. Managing the trade-off between these three strategies allows DRL agents to be robust to inconsistent or intermittent human feedback. Through experimentation using a synthetic oracle, we show that our technique improves the training speed and overall performance of deep reinforcement learning in navigating three-dimensional environments using Minecraft. We further show that our technique is robust to highly innacurate human feedback and can also operate when no human feedback is given

arXiv.org e-Print Archive

Dryad Digital Repository (Duke University)

FigShare