Fast, Parallel, and Cache-Friendly Suffix Array Construction

Dhulipala, Laxman; Khan, Jamshed; Molloy, Erin; Patro, Rob; Rubel, Tobias

Fast, Parallel, and Cache-Friendly Suffix Array Construction

Authors: Laxman Dhulipala
Jamshed Khan
Erin Molloy
Rob Patro
Tobias Rubel
Publication date: 1 January 2023
Publisher: LIPIcs - Leibniz International Proceedings in Informatics. 23rd International Workshop on Algorithms in Bioinformatics (WABI 2023)
Doi

Abstract

String indexes such as the suffix array (SA) and the closely related longest common prefix (LCP) array are fundamental objects in bioinformatics and have a wide variety of applications. Despite their importance in practice, few scalable parallel algorithms for constructing these are known, and the existing algorithms can be highly non-trivial to implement and parallelize. In this paper we present CaPS-SA, a simple and scalable parallel algorithm for constructing these string indexes inspired by samplesort. Due to its design, CaPS-SA has excellent memory-locality and thus incurs fewer cache misses and achieves strong performance on modern multicore systems with deep cache hierarchies. We show that despite its simple design, CaPS-SA outperforms existing state-of-the-art parallel SA and LCP-array construction algorithms on modern hardware. Finally, motivated by applications in modern aligners where the query strings have bounded lengths, we introduce the notion of a bounded-context SA and show that CaPS-SA can easily be extended to exploit this structure to obtain further speedups

Similar works

Full text

Open in the Core reader

Download PDF

Available Versions

DROPS Dagstuhl Research Online Publication Server

oai:drops-oai.dagstuhl.de:1864...

Last time updated on 12/09/2023