Search CORE

72,350 research outputs found

A Fast Anderson-Chebyshev Acceleration for Nonlinear Optimization

Author: Li Jian
Li Zhize
Publication venue
Publication date: 01/03/2020
Field of study

Anderson acceleration (or Anderson mixing) is an efficient acceleration method for fixed point iterations

x_{t+1}=G(x_t)

, e.g., gradient descent can be viewed as iteratively applying the operation

G(x) \triangleq x-\alpha\nabla f(x)

. It is known that Anderson acceleration is quite efficient in practice and can be viewed as an extension of Krylov subspace methods for nonlinear problems. In this paper, we show that Anderson acceleration with Chebyshev polynomial can achieve the optimal convergence rate

O(\sqrt{\kappa}\ln\frac{1}{\epsilon})

, which improves the previous result

O(\kappa\ln\frac{1}{\epsilon})

provided by (Toth and Kelley, 2015) for quadratic functions. Moreover, we provide a convergence analysis for minimizing general nonlinear problems. Besides, if the hyperparameters (e.g., the Lipschitz smooth parameter

L

) are not available, we propose a guessing algorithm for guessing them dynamically and also prove a similar convergence rate. Finally, the experimental results demonstrate that the proposed Anderson-Chebyshev acceleration method converges significantly faster than other algorithms, e.g., vanilla gradient descent (GD), Nesterov's Accelerated GD. Also, these algorithms combined with the proposed guessing algorithm (guessing the hyperparameters dynamically) achieve much better performance.Comment: To appear in AISTATS 202

arXiv.org e-Print Archive

Institutional Knowledge at Singapore Management University

A Simple Proximal Stochastic Gradient Method for Nonsmooth Nonconvex Optimization

Author: Li Jian
Li Zhize
Publication venue
Publication date: 01/12/2018
Field of study

We analyze stochastic gradient algorithms for optimizing nonconvex, nonsmooth finite-sum problems. In particular, the objective function is given by the summation of a differentiable (possibly nonconvex) component, together with a possibly non-differentiable but convex component. We propose a proximal stochastic gradient algorithm based on variance reduction, called ProxSVRG+. Our main contribution lies in the analysis of ProxSVRG+. It recovers several existing convergence results and improves/generalizes them (in terms of the number of stochastic gradient oracle calls and proximal oracle calls). In particular, ProxSVRG+ generalizes the best results given by the SCSG algorithm, recently proposed by [Lei et al., 2017] for the smooth nonconvex case. ProxSVRG+ is also more straightforward than SCSG and yields simpler analysis. Moreover, ProxSVRG+ outperforms the deterministic proximal gradient descent (ProxGD) for a wide range of minibatch sizes, which partially solves an open problem proposed in [Reddi et al., 2016b]. Also, ProxSVRG+ uses much less proximal oracle calls than ProxSVRG [Reddi et al., 2016b]. Moreover, for nonconvex functions satisfied Polyak-\L{}ojasiewicz condition, we prove that ProxSVRG+ achieves a global linear convergence rate without restart unlike ProxSVRG. Thus, it can \emph{automatically} switch to the faster linear convergence in some regions as long as the objective function satisfies the PL condition locally in these regions. ProxSVRG+ also improves ProxGD and ProxSVRG/SAGA, and generalizes the results of SCSG in this case. Finally, we conduct several experiments and the experimental results are consistent with the theoretical results.Comment: 32nd Conference on Neural Information Processing Systems (NeurIPS 2018

arXiv.org e-Print Archive

Institutional Knowledge at Singapore Management University

Recommended from our members

A method to take account of inhomogeneity in mechanical component reliability calculations

Author: Li Jian Ping
Li Jian-Ping
Thompson G.
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 01/03/2005
Field of study

YesThis paper proposes a method by which material inhomogeneity may be taken into account in a reliability calculation. The method employs Monte-Carlo simulation; and introduces a material strength index, and a standard deviation of material strength to model the variation in the strength of a component throughout its volume. The method is compared to conventional load-strength interference theory. The results are identical for the case of homogeneous material, but reliability is shown to reduce for the same load as the component volume increases. The case of a tensile bar is used to explore the variation of reliability with component volume

Bradford Scholars

The University of Manchester - Institutional Repository