Byzantine Stochastic Gradient Descent

Alistarh, Dan-Adrian; Allen-Zhu, Zeyuan; Bengio, S.; Cesa-Bianchi, N.; Garnett, R.; Grauman, K.; Larochelle, H.; Li, Jerry; Wallach, H.

research

Byzantine Stochastic Gradient Descent

Authors: Dan-Adrian Alistarh
Zeyuan Allen-Zhu
S. Bengio
N. Cesa-Bianchi
R. Garnett
K. Grauman
H. Larochelle
Jerry Li
H. Wallach
Publication date: 1 January 2018
Publisher

Abstract

This paper studies the problem of distributed stochastic optimization in an adversarial setting where, out of the

m

machines which allegedly compute stochastic gradients every iteration, an

\alpha

-fraction are Byzantine, and can behave arbitrarily and adversarially. Our main result is a variant of stochastic gradient descent (SGD) which finds

\varepsilon

-approximate minimizers of convex functions in

T = \tilde{O}\big( \frac{1}{\varepsilon^2 m} + \frac{\alpha^2}{\varepsilon^2} \big)

iterations. In contrast, traditional mini-batch SGD needs

T = O\big( \frac{1}{\varepsilon^2 m} \big)

iterations, but cannot tolerate Byzantine failures. Further, we provide a lower bound showing that, up to logarithmic factors, our algorithm is information-theoretically optimal both in terms of sampling complexity and time complexity

Similar works

Full text

Available Versions

IST Austria: PubRep (Institute of Science and Technology)

oai:pub.research-explorer.app....

Last time updated on 15/12/2019