Search CORE

20 research outputs found

Efficient LZ78 factorization of grammar compressed text

Author: A. Amir
A. Jeż
E. Ukkonen
E.M. McCreight
J. Jansson
J. Westbrook
J. Ziv
J. Ziv
K. Goto
K. Goto
M. Crochemore
M. Li
M. Li
M.A. Bender
O. Berkman
P. Weiner
R. Cilibrasi
T. Kida
V. Freschi
W. Rytter
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/01/2012
Field of study

We present an efficient algorithm for computing the LZ78 factorization of a text, where the text is represented as a straight line program (SLP), which is a context free grammar in the Chomsky normal form that generates a single string. Given an SLP of size

n

representing a text

S

of length

N

, our algorithm computes the LZ78 factorization of

T

O(n\sqrt{N}+m\log N)

time and

O(n\sqrt{N}+m)

space, where

m

is the number of resulting LZ78 factors. We also show how to improve the algorithm so that the

n\sqrt{N}

term in the time and space complexities becomes either

nL

, where

L

is the length of the longest LZ78 factor, or

(N - \alpha)

where

\alpha \geq 0

is a quantity which depends on the amount of redundancy that the SLP captures with respect to substrings of

S

of a certain length. Since

m = O(N/\log_\sigma N)

where

\sigma

is the alphabet size, the latter is asymptotically at least as fast as a linear time algorithm which runs on the uncompressed string when

\sigma

is constant, and can be more efficient when the text is compressible, i.e. when

m

and

n

are small.Comment: SPIRE 201

arXiv.org e-Print Archive

Crossref

Speeding-up $q$ -gram mining on grammar-based compressed texts

Author: Bannai Hideo
Goto Keisuke
Inenaga Shunuke
Takeda Masayuki
坂内英夫
後藤啓介
稲永俊介
竹田正幸
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 15/02/2012
Field of study

We present an efficient algorithm for calculating

q

-gram frequencies on strings represented in compressed form, namely, as a straight line program (SLP). Given an SLP

\mathcal{T}

of size

n

that represents string

T

, the algorithm computes the occurrence frequencies of all

q

-grams in

T

, by reducing the problem to the weighted

q

-gram frequencies problem on a trie-like structure of size

m = |T|-\mathit{dup}(q,\mathcal{T})

, where

\mathit{dup}(q,\mathcal{T})

is a quantity that represents the amount of redundancy that the SLP captures with respect to

q

-grams. The reduced problem can be solved in linear time. Since

m = O(qn)

, the running time of our algorithm is

O(\min\{|T|-\mathit{dup}(q,\mathcal{T}),qn\})

, improving our previous

O(qn)

algorithm when

q = \Omega(|T|/n)

arXiv.org e-Print Archive

Kyushu University Institutional Repository

Minimal Absent Words in Rooted and Unrooted Trees

Author: B Schieber
C Barton
D Belazzougui
D Belazzougui
F Mignosi
F Mignosi
F Mignosi
G Fici
G Fici
M Béal
M Béal
M Crochemore
M Crochemore
M Crochemore
M-P Béal
MA Bender
P Charalampopoulos
P Charalampopoulos
RM Silva
S Chairungsee
T Shibuya
Y Almirantis
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/01/2019
Field of study

We extend the theory of minimal absent words to (rooted and unrooted) trees, having edges labeled by letters from an alphabet of cardinality. We show that the set of minimal absent words of a rooted (resp. unrooted) tree T with n nodes has cardinality (resp.), and we show that these bounds are realized. Then, we exhibit algorithms to compute all minimal absent words in a rooted (resp. unrooted) tree in output-sensitive time (resp. assuming an integer alphabet of size polynomial in n

arXiv.org e-Print Archive

Crossref

Archivio istituzionale della ricerca - Università di Palermo

Computing Runs on a Trie

Author: Bannai Hideo
Inenaga Shunsuke
Nakashima Yuto
Sugahara Ryo
Takeda Masayuki
Publication venue: LIPIcs - Leibniz International Proceedings in Informatics. 30th Annual Symposium on Combinatorial Pattern Matching (CPM 2019)
Publication date: 01/01/2019
Field of study

A maximal repetition, or run, in a string, is a maximal periodic substring whose smallest period is at most half the length of the substring. In this paper, we consider runs that correspond to a path on a trie, or in other words, on a rooted edge-labeled tree where the endpoints of the path must be a descendant/ancestor of the other. For a trie with n edges, we show that the number of runs is less than n. We also show an O(n sqrt{log n}log log n) time and O(n) space algorithm for counting and finding the shallower endpoint of all runs. We further show an O(n log n) time and O(n) space algorithm for finding both endpoints of all runs. We also discuss how to improve the running time even more

arXiv.org e-Print Archive

Dagstuhl Research Online Publication Server

Compact q-gram Profiling of Compressed Strings

Author: E. Ukkonen
G. Paaß
J. Kärkkäinen
J. Ziv
J. Ziv
K. Goto
K. Goto
M. Charikar
R.M. Karp
T. Gärtner
T. Shibuya
W. Matsubara
W. Rytter
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/01/2013
Field of study

We consider the problem of computing the q-gram profile of a string \str of size

N

compressed by a context-free grammar with

n

production rules. We present an algorithm that runs in

O(N-\alpha)

expected time and uses O(n+q+\kq) space, where

N-\alpha\leq qn

is the exact number of characters decompressed by the algorithm and \kq\leq N-\alpha is the number of distinct q-grams in \str. This simultaneously matches the current best known time bound and improves the best known space bound. Our space bound is asymptotically optimal in the sense that any algorithm storing the grammar and the q-gram profile must use \Omega(n+q+\kq) space. To achieve this we introduce the q-gram graph that space-efficiently captures the structure of a string with respect to its q-grams, and show how to construct it from a grammar

arXiv.org e-Print Archive

CiteSeerX

Crossref

Online Research Database In Technology

Tying up the loose ends in fully LZW-compressed pattern matching

Author: Gawrychowski Pawel
Publication venue: LIPIcs - Leibniz International Proceedings in Informatics. 29th International Symposium on Theoretical Aspects of Computer Science (STACS 2012)
Publication date: 19/09/2011
Field of study

We consider a natural generalization of the classical pattern matching problem: given compressed representations of a pattern p[1..M] and a text t[1..N] of sizes m and n, respectively, does p occur in t? We develop an optimal linear time solution for the case when p and t are compressed using the LZW method. This improves the previously known O((n+m)log(n+m)) time solution of Gasieniec and Rytter, and essentially closes the line of research devoted to tudying LZW-compressed exact pattern matching

arXiv.org e-Print Archive

Dagstuhl Research Online Publication Server

MPG.PuRe