Search CORE

11,444 research outputs found

Composite repetition-aware data structures

Author: A Blumer
A Lempel
D Arroyuelo
D Belazzougui
DE Willard
J Radoszewski
J Sirén
J Ziv
M Crochemore
M Crochemore
M Raffinot
P Ferragina
S Kreft
T Gagie
V Mäkinen
V Mäkinen
W Rytter
Publication venue
Publication date: 01/01/2015
Field of study

In highly repetitive strings, like collections of genomes from the same species, distinct measures of repetition all grow sublinearly in the length of the text, and indexes targeted to such strings typically depend only on one of these measures. We describe two data structures whose size depends on multiple measures of repetition at once, and that provide competitive tradeoffs between the time for counting and reporting all the exact occurrences of a pattern, and the space taken by the structure. The key component of our constructions is the run-length encoded BWT (RLBWT), which takes space proportional to the number of BWT runs: rather than augmenting RLBWT with suffix array samples, we combine it with data structures from LZ77 indexes, which take space proportional to the number of LZ77 factors, and with the compact directed acyclic word graph (CDAWG), which takes space proportional to the number of extensions of maximal repeats. The combination of CDAWG and RLBWT enables also a new representation of the suffix tree, whose size depends again on the number of extensions of maximal repeats, and that is powerful enough to support matching statistics and constant-space traversal.Comment: (the name of the third co-author was inadvertently omitted from previous version

arXiv.org e-Print Archive

Archivio istituzionale della ricerca - Università degli Studi di Venezia Ca' Foscari

Fully dynamic data structure for LCE queries in compressed space

Author: Bannai Hideo
I Tomohiro
Inenaga Shunsuke
Nishimoto Takaaki
Takeda Masayuki
Publication venue
Publication date: 01/01/2016
Field of study

A Longest Common Extension (LCE) query on a text

T

of length

N

asks for the length of the longest common prefix of suffixes starting at given two positions. We show that the signature encoding

\mathcal{G}

of size

w = O(\min(z \log N \log^* M, N))

[Mehlhorn et al., Algorithmica 17(2):183-198, 1997] of

T

, which can be seen as a compressed representation of

T

, has a capability to support LCE queries in

O(\log N + \log \ell \log^* M)

time, where

\ell

is the answer to the query,

z

is the size of the Lempel-Ziv77 (LZ77) factorization of

T

, and

M \geq 4N

is an integer that can be handled in constant time under word RAM model. In compressed space, this is the fastest deterministic LCE data structure in many cases. Moreover,

\mathcal{G}

can be enhanced to support efficient update operations: After processing

\mathcal{G}

O(w f_{\mathcal{A}})

time, we can insert/delete any (sub)string of length

y

into/from an arbitrary position of

T

O((y+ \log N\log^* M) f_{\mathcal{A}})

time, where

f_{\mathcal{A}} = O(\min \{ \frac{\log\log M \log\log w}{\log\log\log M}, \sqrt{\frac{\log w}{\log\log w}} \})

. This yields the first fully dynamic LCE data structure. We also present efficient construction algorithms from various types of inputs: We can construct

\mathcal{G}

O(N f_{\mathcal{A}})

time from uncompressed string

T

; in

O(n \log\log n \log N \log^* M)

time from grammar-compressed string

T

represented by a straight-line program of size

n

; and in

O(z f_{\mathcal{A}} \log N \log^* M)

time from LZ77-compressed string

T

with

z

factors. On top of the above contributions, we show several applications of our data structures which improve previous best known results on grammar-compressed string processing.Comment: arXiv admin note: text overlap with arXiv:1504.0695

arXiv.org e-Print Archive

Dagstuhl Research Online Publication Server

The k-mismatch problem revisited

Author: Clifford Raphaël
Fontaine Allyx
Porat Ely
Sach Benjamin
Starikovskaya Tatiana
Publication venue
Publication date: 27/08/2015
Field of study

We revisit the complexity of one of the most basic problems in pattern matching. In the k-mismatch problem we must compute the Hamming distance between a pattern of length m and every m-length substring of a text of length n, as long as that Hamming distance is at most k. Where the Hamming distance is greater than k at some alignment of the pattern and text, we simply output "No". We study this problem in both the standard offline setting and also as a streaming problem. In the streaming k-mismatch problem the text arrives one symbol at a time and we must give an output before processing any future symbols. Our main results are as follows: 1) Our first result is a deterministic

O(n k^2\log{k} / m+n \text{polylog} m)

time offline algorithm for k-mismatch on a text of length n. This is a factor of k improvement over the fastest previous result of this form from SODA 2000 by Amihood Amir et al. 2) We then give a randomised and online algorithm which runs in the same time complexity but requires only

O(k^2\text{polylog} {m})

space in total. 3) Next we give a randomised

(1+\epsilon)

-approximation algorithm for the streaming k-mismatch problem which uses

O(k^2\text{polylog} m / \epsilon^2)

space and runs in

O(\text{polylog} m / \epsilon^2)

worst-case time per arriving symbol. 4) Finally we combine our new results to derive a randomised

O(k^2\text{polylog} {m})

space algorithm for the streaming k-mismatch problem which runs in

O(\sqrt{k}\log{k} + \text{polylog} {m})

worst-case time per arriving symbol. This improves the best previous space complexity for streaming k-mismatch from FOCS 2009 by Benny Porat and Ely Porat by a factor of k. We also improve the time complexity of this previous result by an even greater factor to match the fastest known offline algorithm (up to logarithmic factors)

arXiv.org e-Print Archive

Compressed Membership for NFA (DFA) with Compressed Labels is in NP (P)

Author: A. Amir
A. Jeż
A. Jeż
A. Jeż
A. Jeż
Artur Jeż
B. Genest
G. Navarro
J. MacDonald
K. Mehlhorn
L. Gąsieniec
L. Gąsieniec
L. Gąsieniec
M. Beaudry
M. Charikar
M. Farach
M. Lohrey
M. Lohrey
M. Lohrey
M. Lohrey
M. Lohrey
N. Markey
P. Bille
P. Ferragina
P. Gawrychowski
P. Gawrychowski
P. Gawrychowski
P. Gawrychowski
S. Alstrup
S. Lasota
S.R. Kosaraju
T. Kida
W. Czerwiński
W. Plandowski
W. Plandowski
W. Plandowski
W. Plandowski
W. Rytter
Y. Lifshits
Y. Lifshits
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 11/10/2011
Field of study

In this paper, a compressed membership problem for finite automata, both deterministic and non-deterministic, with compressed transition labels is studied. The compression is represented by straight-line programs (SLPs), i.e. context-free grammars generating exactly one string. A novel technique of dealing with SLPs is introduced: the SLPs are recompressed, so that substrings of the input text are encoded in SLPs labelling the transitions of the NFA (DFA) in the same way, as in the SLP representing the input text. To this end, the SLPs are locally decompressed and then recompressed in a uniform way. Furthermore, such recompression induces only small changes in the automaton, in particular, the size of the automaton remains polynomial. Using this technique it is shown that the compressed membership for NFA with compressed labels is in NP, thus confirming the conjecture of Plandowski and Rytter and extending the partial result of Lohrey and Mathissen; as it is already known, that this problem is NP-hard, we settle its exact computational complexity. Moreover, the same technique applied to the compressed membership for DFA with compressed labels yields that this problem is in P; for this problem, only trivial upper-bound PSPACE was known

arXiv.org e-Print Archive

CiteSeerX

Springer - Publisher Connector

Dagstuhl Research Online Publication Server