Search CORE

8 research outputs found

A General SIMD-based Approach to Accelerating Compression Algorithms

Author: Lemire Daniel
Nie Jian-Yun
Shan Dongdong
Wen Ji-Rong
Yan Hongfei
Zhang Xudong
Zhao Wayne Xin
Publication venue: 'Association for Computing Machinery (ACM)'
Publication date: 01/01/2015
Field of study

Compression algorithms are important for data oriented tasks, especially in the era of Big Data. Modern processors equipped with powerful SIMD instruction sets, provide us an opportunity for achieving better compression performance. Previous research has shown that SIMD-based optimizations can multiply decoding speeds. Following these pioneering studies, we propose a general approach to accelerate compression algorithms. By instantiating the approach, we have developed several novel integer compression algorithms, called Group-Simple, Group-Scheme, Group-AFOR, and Group-PFD, and implemented their corresponding vectorized versions. We evaluate the proposed algorithms on two public TREC datasets, a Wikipedia dataset and a Twitter dataset. With competitive compression ratios and encoding speeds, our SIMD-based algorithms outperform state-of-the-art non-vectorized algorithms with respect to decoding speeds

arXiv.org e-Print Archive

R-libre

Crossref

Make the most out of your SIMD investments: Counter control flow divergence in compiled query pipelines

Author: Boncz P.A. (Peter)
Kemper A. (Alfons)
Kipf A. (Andreas)
Lang H. (Harald)
Neumann T. (Thomas)
Passing L.K. (Linnea)
Publication venue: 'Association for Computing Machinery (ACM)'
Publication date: 11/06/2018
Field of study

Increasing single instruction multiple data (SIMD) capabilities in modern hardware allows for compiling efficient data-parallel query pipelines. This means GPU-alike challenges arise: control flow divergence causes underutilization of vector-processing units. In this paper, we present efficient algorithms for the AVX-512 architecture to address this issue. These algorithms allow for fine-grained assignment of new tuples to idle SIMD lanes. Furthermore, we present strategies for their integration with compiled query pipelines without introducing inefficient memory materializations. We evaluate our approach with a high-performance geospatial join query, which shows performance improvements of up to 35%

Crossref

CWI's Institutional Repository

Scipedia

A Learning Based Framework for Improving Querying on Web Interfaces of Curated Knowledge Bases

Author: Ali Shemshadi
Altman Naomi S.
Elbassuoni Shady
Han Jiawei
Hasan Rakebul
Kerry Taylor
Lehmann Jens
Lina Yao
Lorey Johannes
Morsey Mohamed
Quan Z. Sheng
Shu Yanfeng
Wei Emma Zhang
Yongrui Qin
Zhang Wei Emma
Publication venue: 'Association for Computing Machinery (ACM)'
Publication date: 01/01/2018
Field of study

Knowledge Bases (KBs) are widely used as one of the fundamental components in Semantic Web applications as they provide facts and relationships that can be automatically understood by machines. Curated knowledge bases usually use Resource Description Framework (RDF) as the data representation model. To query the RDF-presented knowledge in curated KBs, Web interfaces are built via SPARQL Endpoints. Currently, querying SPARQL Endpoints has problems like network instability and latency, which affect the query efficiency. To address these issues, we propose a client-side caching framework, SPARQL Endpoint Caching Framework (SECF), aiming at accelerating the overall querying speed over SPARQL Endpoints. SECF identifies the potential issued queries by leveraging the querying patterns learned from clients’ historical queries and prefecthes/caches these queries. In particular, we develop a distance function based on graph edit distance to measure the similarity of SPARQL queries. We propose a feature modelling method to transform SPARQL queries to vector representation that are fed into machine-learning algorithms. A time-aware smoothing-based method, Modified Simple Exponential Smoothing (MSES), is developed for cache replacement. Extensive experiments performed on real-world queries showcase the effectiveness of our approach, which outperforms the state-of-the-art work in terms of the overall querying speed

Crossref

Adelaide Research & Scholarship

University of Huddersfield Repository

Huddersfield Research Portal

From a Comprehensive Experimental Survey to a Cost-based Selection Strategy for Lightweight Integer Compression Algorithms

Author: Damme Patrick
Habich Dirk
Hildebrandt Juliana
Lehner Wolfgang
Ungethüm Annett
Publication venue: 'Association for Computing Machinery (ACM)'
Publication date: 11/01/2023
Field of study

Lightweight integer compression algorithms are frequently applied in in-memory database systems to tackle the growing gap between processor speed and main memory bandwidth. In recent years, the vectorization of basic techniques such as delta coding and null suppression has considerably enlarged the corpus of available algorithms. As a result, today there is a large number of algorithms to choose from, while different algorithms are tailored to different data characteristics. However, a comparative evaluation of these algorithms with different data and hardware characteristics has never been sufficiently conducted in the literature. To close this gap, we conducted an exhaustive experimental survey by evaluating several state-of-the-art lightweight integer compression algorithms as well as cascades of basic techniques. We systematically investigated the influence of data as well as hardware properties on the performance and the compression rates. The evaluated algorithms are based on publicly available implementations as well as our own vectorized reimplementations. We summarize our experimental findings leading to several new insights and to the conclusion that there is no single-best algorithm. Moreover, in this article, we also introduce and evaluate a novel cost model for the selection of a suitable lightweight integer compression algorithm for a given dataset

Qucosa

HSSS - Hochschulschriftenserver der SLUB

Technische Universität Dresden: Qucosa

A General SIMD-Based Approach to Accelerating Compression Algorithms

Author: Büttcher S.
Daniel Lemire
Dongdong Shan
Hongfei Yan
Ji-Rong Wen
Jian-Yun Nie
Jonassen S.
Lemke C.
Liu K.
Lomont C.
Ma W.
Wayne Xin Zhao
Witten I. H.
Xudong Zhang
Publication venue: 'Association for Computing Machinery (ACM)'
Publication date
Field of study

Crossref