Search CORE

114 research outputs found

Learning Mixtures of Linear Classifiers

Author: Ioannidis Stratis
Montanari Andrea
Sun Yuekai
Publication venue
Publication date: 30/07/2014
Field of study

We consider a discriminative learning (regression) problem, whereby the regression function is a convex combination of k linear classifiers. Existing approaches are based on the EM algorithm, or similar techniques, without provable guarantees. We develop a simple method based on spectral techniques and a `mirroring' trick, that discovers the subspace spanned by the classifiers' parameter vectors. Under a probabilistic assumption on the feature vector distribution, we prove that this approach has nearly optimal statistical efficiency

arXiv.org e-Print Archive

CiteSeerX

Score Function Features for Discriminative Learning: Matrix and Tensor Framework

Author: Anandkumar Anima
Janzamin Majid
Sedghi Hanie
Publication venue
Publication date: 08/12/2014
Field of study

Feature learning forms the cornerstone for tackling challenging learning problems in domains such as speech, computer vision and natural language processing. In this paper, we consider a novel class of matrix and tensor-valued features, which can be pre-trained using unlabeled samples. We present efficient algorithms for extracting discriminative information, given these pre-trained features and labeled samples for any related task. Our class of features are based on higher-order score functions, which capture local variations in the probability density function of the input. We establish a theoretical framework to characterize the nature of discriminative information that can be extracted from score-function features, when used in conjunction with labeled samples. We employ efficient spectral decomposition algorithms (on matrices and tensors) for extracting discriminative components. The advantage of employing tensor-valued features is that we can extract richer discriminative information in the form of an overcomplete representations. Thus, we present a novel framework for employing generative models of the input for discriminative learning.Comment: 29 page

arXiv.org e-Print Archive

eScholarship - University of California

Score Function Features for Discriminative Learning

Author: Anandkumar Anima
Janzamin Majid
Sedghi Hanie
Publication venue
Publication date: 19/12/2014
Field of study

arXiv.org e-Print Archive

eScholarship - University of California

Caltech Authors

Recommended from our members

Provable Tensor Methods for Learning Mixtures of Generalized Linear Models

Author: Anandkumar Anima
Janzamin Majid
Sedghi Hanie
Publication venue: PMLR
Publication date: 09/12/2014
Field of study

We consider the problem of learning mixtures of generalized linear models (GLM) which arise in classification and regression problems. Typical learning approaches such as expectation maximization (EM) or variational Bayes can get stuck in spurious local optima. In contrast, we present a tensor decomposition method which is guaranteed to correctly recover the parameters. The key insight is to employ certain feature transformations of the input, which depend on the input generative model. Specifically, we employ score function tensors of the input and compute their cross-correlation with the response variable. We establish that the decomposition of this tensor consistently recovers the parameters, under mild non-degeneracy conditions. We demonstrate that the computational and sample complexity of our method is a low order polynomial of the input and the latent dimensions

eScholarship - University of California

Caltech Authors

SQ Lower Bounds for Learning Mixtures of Linear Classifiers

Author: Diakonikolas Ilias
Kane Daniel M.
Sun Yuxin
Publication venue
Publication date: 18/10/2023
Field of study

We study the problem of learning mixtures of linear classifiers under Gaussian covariates. Given sample access to a mixture of

r

distributions on

\mathbb{R}^n

of the form

(\mathbf{x},y_{\ell})

\ell\in [r]

, where

\mathbf{x}\sim\mathcal{N}(0,\mathbf{I}_n)

and

y_\ell=\mathrm{sign}(\langle\mathbf{v}_\ell,\mathbf{x}\rangle)

for an unknown unit vector

\mathbf{v}_\ell

, the goal is to learn the underlying distribution in total variation distance. Our main result is a Statistical Query (SQ) lower bound suggesting that known algorithms for this problem are essentially best possible, even for the special case of uniform mixtures. In particular, we show that the complexity of any SQ algorithm for the problem is

n^{\mathrm{poly}(1/\Delta) \log(r)}

, where

\Delta

is a lower bound on the pairwise

\ell_2

-separation between the

\mathbf{v}_\ell

's. The key technical ingredient underlying our result is a new construction of spherical designs that may be of independent interest.Comment: To appear in NeurIPS 202

arXiv.org e-Print Archive