Towards Accelerating Training of Batch Normalization: A Manifold
  Perspective

Chen, Wei; Ma, Zhi-Ming; Meng, Qi; Yi, Mingyang

Towards Accelerating Training of Batch Normalization: A Manifold Perspective

Authors: Wei Chen
Zhi-Ming Ma
Qi Meng
Mingyang Yi
Publication date: 8 January 2021
Publisher

Abstract

Batch normalization (BN) has become a crucial component across diverse deep neural networks. The network with BN is invariant to positively linear re-scaling of weights, which makes there exist infinite functionally equivalent networks with various scales of weights. However, optimizing these equivalent networks with the first-order method such as stochastic gradient descent will converge to different local optima owing to different gradients across training. To alleviate this, we propose a quotient manifold \emph{PSI manifold}, in which all the equivalent weights of the network with BN are regarded as the same one element. Then, gradient descent and stochastic gradient descent on the PSI manifold are also constructed. The two algorithms guarantee that every group of equivalent weights (caused by positively re-scaling) converge to the equivalent optima. Besides that, we give the convergence rate of the proposed algorithms on PSI manifold and justify that they accelerate training compared with the algorithms on the Euclidean weight space. Empirical studies show that our algorithms can consistently achieve better performances over various experimental settings

Similar works

Full text

Available Versions

arXiv.org e-Print Archive

oai:arXiv.org:2101.02916

Last time updated on 02/03/2021