Search CORE

25 research outputs found

Performance-Optimal Filtering:Bloom Overtakes Cuckoo at High Throughput

Author: Boncz P.A. (Peter)
Kemper A. (Alfons)
Lang H. (Harald)
Neumann T. (Thomas)
Publication venue: 'VLDB Endowment'
Publication date: 01/01/2019
Field of study

How Good Are Query Optimizers, Really?

Author: Boncz P.A. (Peter)
Gubichev A. (Andrey)
Kemper A. (Alfons)
Leis V. (Viktor)
Mirchev A. (Atanas)
Neumann T. (Thomas)
Publication venue
Publication date: 01/11/2015
Field of study

Finding a good join order is crucial for query performance. In this paper, we introduce the Join Order Benchmark (JOB) and experimentally revisi

CWI's Institutional Repository

Learned cardinalities: Estimating correlated joins with deep learning

Author: Boncz P.A. (Peter)
Kemper A. (Alfons)
Kipf A. (Andreas)
Kipf T. (Thomas)
Leis V. (Viktor)
Radke B. (Bernhard)
Publication venue
Publication date: 01/01/2019
Field of study

We describe a new deep learning approach to cardinality estimation. MSCN is a multi-set convolutional network, tailored to representing relational query plans, that employs set semantics to capture query features and true cardinalities. MSCN builds on sampling-based estimation, addressing its weaknesses when no sampled tuples qualify a predicate, and in capturing join-crossing correlations. Our evaluation of MSCN using a real-world dataset shows that deep learning signiicantly enhances the quality of cardinality estimation, which is the core problem in query optimization

CWI's Institutional Repository

Learned Cardinalities: Estimating Correlated Joins with Deep Learning

Author: Boncz P.A. (Peter)
Kemper A. (Alfons)
Kipf A. (Andreas)
Kipf T. (Thomas)
Leis V. (Viktor)
Radke B. (Bernhard)
Publication venue
Publication date: 18/12/2018
Field of study

CWI's Institutional Repository

Make the most out of your SIMD investments: Counter control flow divergence in compiled query pipelines

Author: Boncz P.A. (Peter)
Kemper A. (Alfons)
Kipf A. (Andreas)
Lang H. (Harald)
Neumann T. (Thomas)
Passing L.K. (Linnea)
Publication venue: 'Association for Computing Machinery (ACM)'
Publication date: 11/06/2018
Field of study

Increasing single instruction multiple data (SIMD) capabilities in modern hardware allows for compiling efficient data-parallel query pipelines. This means GPU-alike challenges arise: control flow divergence causes underutilization of vector-processing units. In this paper, we present efficient algorithms for the AVX-512 architecture to address this issue. These algorithms allow for fine-grained assignment of new tuples to idle SIMD lanes. Furthermore, we present strategies for their integration with compiled query pipelines without introducing inefficient memory materializations. We evaluate our approach with a high-performance geospatial join query, which shows performance improvements of up to 35%

Crossref

CWI's Institutional Repository

Scipedia

Everything You Always Wanted to Know About Compiled and Vectorized Queries But Were Afraid to Ask

Author: Boncz P.A. (Peter)
Kemper A. (Alfons)
Kersten T. (Timo)
Leis V. (Viktor)
Neumann T. (Thomas)
Pavlo A. (Andrew)
Publication venue
Publication date: 01/01/2018
Field of study

CWI's Institutional Repository

Make the most out of your SIMD investments: counter control flow divergence in compiled query pipelines

Author: Boncz P.A. (Peter)
Kemper A. (Alfons)
Kipf A. (Andreas)
Lang H. (Harald)
Neumann T. (Thomas)
Passing L.K. (Linnea)
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 16/07/2019
Field of study

Increasing single instruction multiple data (SIMD) capabilities in modern hardware allows for the compilation of data-parallel query pipelines. This means GPU-alike challenges arise: control flow divergence causes the underutilization of vector-processing units. In this paper, we present efficient algorithms for the AVX-512 architecture to address this issue. These algorithms allow for the fine-grained assignment of new tuples to idle SIMD lanes. Furthermore, we present strategies for their integration with compiled query pipelines so that tuples are never evicted from registers. We evaluate our approach with three query types: (i) a table scan query based on TPC-H Query 1, that performs up to 34% faster when addressing underutilization, (ii) a hashjoin query, where we observe up to 25% higher performance, and (iii) an approximate geospatial join query, which shows performance improvements of up to 30%

CWI's Institutional Repository

Adaptive geospatial joins for modern hardware

Author: Boncz P.A. (Peter)
Kemper A. (Alfons)
Kipf A. (Andreas)
Lang H. (Harald)
Neumann T. (Thomas)
Pandey V.N. (Varun)
Persa R.A. (Raul Alexandru)
Publication venue
Publication date: 26/02/2018
Field of study

Geospatial joins are a core building block of connected mobility applications. An especially challenging problem are joins between streaming points and static polygons. Since points are not known beforehand, they cannot be indexed. Nevertheless, points need to be mapped to polygons with low latencies to enable real-time feedback. We present an adaptive geospatial join that uses true hit filtering to avoid expensive geometric computations in most cases. Our technique uses a quadtree-based hierarchical grid to approximate polygons and stores these approximations in a specialized radix tree. We emphasize on an approximate version of our algorithm that guarantees a user-defined precision. The exact version of our algorithm can adapt to the expected point distribution by refining the index. We optimized our implementation for modern hardware architectures with wide SIMD vector processing units, including Intel’s brand new Knights Landing. Overall, our approach can perform up to two orders of magnitude faster than existing techniques

CWI's Institutional Repository

Approximate geospatial joins with precision guarantees

Author: Boncz P.A. (Peter)
Kemper A. (Alfons)
Kipf A. (Andreas)
Lang H. (Harald)
Neumann T. (Thomas)
Pandey V.N. (Varun)
Persa R.A. (Raul Alexandru)
Publication venue
Publication date: 16/04/2018
Field of study

Geospatial joins are a core building block of con- nected mobility applications. An especially challenging problem are joins between streaming points and static polygons. Since points are not known beforehand, they cannot be indexed. Nevertheless, points need to be mapped to polygons with low latencies to enable real-time feedback. We present an approximate geospatial join that guarantees a user-defined precision. Our technique uses a quadtree-based hierarchical grid to approximate polygons and stores these approximations in a specialized radix tree. Our approach can perform up to several orders of magnitude faster than existing techniques while providing sufficiently precise results for many applications

Crossref

CWI's Institutional Repository

Estimating cardinalities with deep sketches

Author: Boncz P.A. (Peter)
Kemper A. (Alfons)
Kipf A. (Andreas)
Kipf T. (Thomas)
Leis V. (Viktor)
Müller J. (Jonas)
Neumann T. (Thomas)
Radke B. (Bernhard)
Vorona (Dimitri)
Publication venue: 'Association for Computing Machinery (ACM)'
Publication date: 30/06/2019
Field of study

We introduce Deep Sketches, which are compact models of databases that allow us to estimate the result sizes of SQL queries. Deep Sketches are powered by a new deep learning approach to cardinality estimation that can capture correlations between columns, even across tables. Our demonstration allows users to define such sketches on the TPC-H and IMDb datasets, monitor the training process, and run ad-hoc queries against trained sketches. We also estimate query cardinalities with HyPer and PostgreSQL to visualize the gains over traditional cardinality estimators

CWI's Institutional Repository