Search CORE

16,429 research outputs found

Universal Approximation Depth and Errors of Narrow Belief Networks with Discrete Units

Author: Montúfar Guido F.
Publication venue
Publication date: 01/01/2014
Field of study

We generalize recent theoretical work on the minimal number of layers of narrow deep belief networks that can approximate any probability distribution on the states of their visible units arbitrarily well. We relax the setting of binary units (Sutskever and Hinton, 2008; Le Roux and Bengio, 2008, 2010; Mont\'ufar and Ay, 2011) to units with arbitrary finite state spaces, and the vanishing approximation error to an arbitrary approximation error tolerance. For example, we show that a

q

-ary deep belief network with

L\geq 2+\frac{q^{\lceil m-\delta \rceil}-1}{q-1}

layers of width

n \leq m + \log_q(m) + 1

for some

m\in \mathbb{N}

can approximate any probability distribution on

\{0,1,\ldots,q-1\}^n

without exceeding a Kullback-Leibler divergence of

\delta

. Our analysis covers discrete restricted Boltzmann machines and na\"ive Bayes models as special cases.Comment: 19 pages, 5 figures, 1 tabl

arXiv.org e-Print Archive

CiteSeerX

eScholarship - University of California

Nondeterminism and an abstract formulation of Ne\v{c}iporuk's lower bound method

Author: Beame Paul
Grosshans Nathan
McKenzie Pierre
Segoufin Luc
Publication venue: 'Association for Computing Machinery (ACM)'
Publication date: 05/08/2016
Field of study

A formulation of "Ne\v{c}iporuk's lower bound method" slightly more inclusive than the usual complexity-measure-specific formulation is presented. Using this general formulation, limitations to lower bounds achievable by the method are obtained for several computation models, such as branching programs and Boolean formulas having access to a sublinear number of nondeterministic bits. In particular, it is shown that any lower bound achievable by the method of Ne\v{c}iporuk for the size of nondeterministic and parity branching programs is at most

O(n^{3/2}/\log n)

arXiv.org e-Print Archive

INRIA a CCSD electronic archive server

Combinatorially interpreting generalized Stirling numbers

Author: Engbers John
Galvin David
Hilyard Justin
Publication venue
Publication date: 01/01/2014
Field of study

Let

w

be a word in alphabet

\{x,D\}

with

m

x

's and

n

D

's. Interpreting "

x

" as multiplication by

x

, and "

D

" as differentiation with respect to

x

, the identity

wf(x) = x^{m-n}\sum_k S_w(k) x^k D^k f(x)

, valid for any smooth function

f(x)

, defines a sequence

(S_w(k))_k

, the terms of which we refer to as the {\em Stirling numbers (of the second kind)} of

w

. The nomenclature comes from the fact that when

w=(xD)^n

, we have

S_w(k)={n \brace k}

, the ordinary Stirling number of the second kind. Explicit expressions for, and identities satisfied by, the

S_w(k)

have been obtained by numerous authors, and combinatorial interpretations have been presented. Here we provide a new combinatorial interpretation that retains the spirit of the familiar interpretation of

{n \brace k}

as a count of partitions. Specifically, we associate to each

w

a quasi-threshold graph

G_w

, and we show that

S_w(k)

enumerates partitions of the vertex set of

G_w

into classes that do not span an edge of

G_w

. We also discuss some relatives of, and consequences of, our interpretation, including

q

-analogs and bijections between families of labelled forests and sets of restricted partitions.Comment: To appear in Eur. J. Combin., doi:10.1016/j.ejc.2014.07.00

arXiv.org e-Print Archive

CiteSeerX