Search CORE

52 research outputs found

Explore, Exploit or Listen: Combining Human Feedback and Policy Model to Speed up Deep Reinforcement Learning in 3D Worlds

Author: Ashley Haines (4342138)
David Gauthier (1893841)
David Taylor (140886)
Donald Landers (4342165)
Hamish Small (4342156)
Jeffrey Shields (4342159)
John Hoenig (4342153)
John Swenarton (4342132)
Mark Matsche (4342147)
Matthew Smith (326340)
Maya Groner (459554)
Philip Sadler (4342126)
Roger Pradel (111017)
Rémi Choquet (111021)
Wolfgang Vogelbein (4342162)
Publication venue
Publication date: 14/06/2017
Field of study

We describe a method to use discrete human feedback to enhance the performance of deep learning agents in virtual three-dimensional environments by extending deep-reinforcement learning to model the confidence and consistency of human feedback. This enables deep reinforcement learning algorithms to determine the most appropriate time to listen to the human feedback, exploit the current policy model, or explore the agent's environment. Managing the trade-off between these three strategies allows DRL agents to be robust to inconsistent or intermittent human feedback. Through experimentation using a synthetic oracle, we show that our technique improves the training speed and overall performance of deep reinforcement learning in navigating three-dimensional environments using Minecraft. We further show that our technique is robust to highly innacurate human feedback and can also operate when no human feedback is given

arXiv.org e-Print Archive

Dryad Digital Repository (Duke University)

FigShare

Appendix A. An explanation of the fitting of the cause-specific mortality rates model with the MARK program.

Author: Michael Schaub (80488)
Roger Pradel (111017)
Publication venue
Publication date
Field of study

An explanation of the fitting of the cause-specific mortality rates model with the MARK program

FigShare

Appendix B. The parameter index matrices of the cause-specific mortality rates model used in the MARK program.

Author: Michael Schaub (80488)
Roger Pradel (111017)
Publication venue
Publication date
Field of study

The parameter index matrices of the cause-specific mortality rates model used in the MARK program

FigShare

General pattern of transitions between breeding states.

Author: Arnaud Béchet (111025)
Roger Pradel (111017)
Rémi Choquet (111021)
Publication venue
Publication date
Field of study

<p>, breeder with previous experiences and , non-breeder with previous experiences for . The transition probabilities are expressed in terms of , the probability of surviving to the next breeding season, and , the probability, conditional on survival, of breeding the next season (the 's will be age-dependent in practice).</p

FigShare

Appendix B. Decomposition of the transition and event matrices of our model.

Author: Gilles Gauthier (148776)
Guillaume Souchay (2928141)
Roger Pradel (111017)
Publication venue
Publication date
Field of study

Decomposition of the transition and event matrices of our model

FigShare

Appendix F. Goodness-of-fit tests.

Author: Gilles Gauthier (148776)
Guillaume Souchay (2928141)
Roger Pradel (111017)
Publication venue
Publication date
Field of study

Goodness-of-fit tests

FigShare

Appendix D. Rank and identifiability of our multi-event models.

Author: Gilles Gauthier (148776)
Guillaume Souchay (2928141)
Roger Pradel (111017)
Publication venue
Publication date
Field of study

Rank and identifiability of our multi-event models

FigShare

Breeding probability as a function of age and experience for greater flamingos breeding in the Camargue, south of France.

Author: Arnaud Béchet (111025)
Roger Pradel (111017)
Rémi Choquet (111021)
Publication venue
Publication date
Field of study

<p>full squares: no previous breeding episode; full triangles: one previous breeding episode; full circles: 2 or more previous breeding episodes. The curve for inexperienced individuals is that obtained with a normal-year first-year survival of 0.632; the dashed curve below is for a value of 0.763 of the same parameter corresponding to an absence of emigration (see text for details).</p

FigShare

Prediction of the role of experience in the increase of breeding probability with age.

Author: Arnaud Béchet (111025)
Roger Pradel (111017)
Rémi Choquet (111021)
Publication venue
Publication date
Field of study

<p>Under a pure restraint hypothesis, breeding probability is hypothesized to increase as a response to the decline in residual reproductive value with age <a href="http://www.plosone.org/article/info:doi/10.1371/journal.pone.0051016#pone.0051016-Pianka1" target="_blank">[61]</a>. Under a pure constraint hypothesis, breeding probability increases through improved skills <a href="http://www.plosone.org/article/info:doi/10.1371/journal.pone.0051016#pone.0051016-Forslund1" target="_blank">[2]</a>.</p

FigShare

Appendix G. Complete results of model selection.

Author: Gilles Gauthier (148776)
Guillaume Souchay (2928141)
Roger Pradel (111017)
Publication venue
Publication date
Field of study

Complete results of model selection

FigShare