Search CORE

3,471 research outputs found

Equilibrium Propagation: Bridging the Gap Between Energy-Based Models and Backpropagation

Author: Bengio Yoshua
Scellier Benjamin
Publication venue
Publication date: 01/01/2017
Field of study

We introduce Equilibrium Propagation, a learning framework for energy-based models. It involves only one kind of neural computation, performed in both the first phase (when the prediction is made) and the second phase of training (after the target or prediction error is revealed). Although this algorithm computes the gradient of an objective function just like Backpropagation, it does not need a special computation or circuit for the second phase, where errors are implicitly propagated. Equilibrium Propagation shares similarities with Contrastive Hebbian Learning and Contrastive Divergence while solving the theoretical issues of both algorithms: our algorithm computes the gradient of a well defined objective function. Because the objective function is defined in terms of local perturbations, the second phase of Equilibrium Propagation corresponds to only nudging the prediction (fixed point, or stationary distribution) towards a configuration that reduces prediction error. In the case of a recurrent multi-layer supervised network, the output units are slightly nudged towards their target in the second phase, and the perturbation introduced at the output layer propagates backward in the hidden layers. We show that the signal 'back-propagated' during this second phase corresponds to the propagation of error derivatives and encodes the gradient of the objective function, when the synaptic update corresponds to a standard form of spike-timing dependent plasticity. This work makes it more plausible that a mechanism similar to Backpropagation could be implemented by brains, since leaky integrator neural computation performs both inference and error back-propagation in our model. The only local difference between the two phases is whether synaptic changes are allowed or not

arXiv.org e-Print Archive

Directory of Open Access Journals

Frontiers - Publisher Connector

Recommended from our members

Economic MPC of Nonlinear Processes via Recurrent Neural Networks Using Structural Process Knowledge

Author: Park Michael Jihyuck
Publication venue: eScholarship, University of California
Publication date: 01/01/2020
Field of study

This work discusses three methods that incorporate a priori process knowledge into recurrent neural network (RNN) modeling of nonlinear processes to get increased prediction accuracy and provide information on how the neural network models are structured. The first method proposes a hybrid model that integrates first-principles models and RNN models together. The second method proposes a partially-connected RNN model which its structure is based on a priori structural process knowledge. The third method proposes a weight-constrained RNN model that integrates weight constraints into the training of the RNN model. The proposed RNN models are used in an economic model predictive control system and then applied to a chemical process example to validate the improved approximation performance compared to a fully-connected RNN model that is treated as a black box model

eScholarship - University of California

Automating Vehicles by Deep Reinforcement Learning using Task Separation with Hill Climbing

Author: A Liniger
B Paden
C Urmson
CW Anderson
D Dolgov
D Wierstra
DQ Mayne
E Frazzoli
HT Siegelmann
J Xu
P Falcone
R Tedrake
T Schouwenaars
Publication venue
Publication date: 02/08/2018
Field of study

Within the context of autonomous driving a model-based reinforcement learning algorithm is proposed for the design of neural network-parameterized controllers. Classical model-based control methods, which include sampling- and lattice-based algorithms and model predictive control, suffer from the trade-off between model complexity and computational burden required for the online solution of expensive optimization or search problems at every short sampling time. To circumvent this trade-off, a 2-step procedure is motivated: first learning of a controller during offline training based on an arbitrarily complicated mathematical system model, before online fast feedforward evaluation of the trained controller. The contribution of this paper is the proposition of a simple gradient-free and model-based algorithm for deep reinforcement learning using task separation with hill climbing (TSHC). In particular, (i) simultaneous training on separate deterministic tasks with the purpose of encoding many motion primitives in a neural network, and (ii) the employment of maximally sparse rewards in combination with virtual velocity constraints (VVCs) in setpoint proximity are advocated.Comment: 10 pages, 6 figures, 1 tabl

arXiv.org e-Print Archive

Crossref