Search CORE

3 research outputs found

Business Analytics in (a) Blink

Author: Barber R.
Bendel P.
Czech M.
Draese O.
Ho F.
Hrle N.
Idreos S. (Stratos)
Kim M
Koeth O
Lee J. (Jae-Gil)
Lohman G.
Morfonios K.
Mueller R. (René)
Murthy K
Pandis I.
Qiao L.
Raman V. (Vijayshankar)
Sidle R.
Stolze K
Szabo S.
Tim Li T. (Tianchao)
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 01/03/2012
Field of study

The Blink project’s ambitious goal is to answer all Business Intelligence (BI) queries in mere seconds, regardless of the database size, with an extremely low total cost of ownership. Blink is a new DBMS aimed primarily at read-mostly BI query processing that exploits scale-out of commodity multi-core processors and cheap DRAM to retain a (copy of a) data mart completely in main memory. Additionally, it exploits proprietary compression technology and cache-conscious algorithms that reduce memory bandwidth consumption and allow most SQL query processing to be performed on the compressed data. Blink always scans (portions of) the data mart in parallel on all nodes, without using any indexes or materialized views, and without any query optimizer to choose among them. The Blink technology has thus far been incorp

CWI's Institutional Repository

Composite Group-Keys

Author: F Färber
J Krüger
J Rao
M Böhm
Martin Grund
V Sikka
Vijayshankar Raman
Publication venue: 'Springer Science and Business Media LLC'
Publication date
Field of study

Crossref

Joins on encoded and partitioned data

Author: Attaluri Gopi K.
Barber Ronald
Chainani Naresh
Draese Oliver
Ho Frederick
Idreos Stratos
Kim Min Soo
Lee Jae Gil
Lightstone Sam S.
Lohman Guy M.
Morfonios Konstantinos
Murthy Keshava V Sreerama
Pandis Ippokratis
Qiao Lin
Raman Vijayshankar
Samy Vincent Kulandai
Sidle Richard
Stolze Knut
Zhang Liping
Publication venue: ACM SIGMOD and VLDB Endowment
Publication date: 01/09/2014
Field of study

Compression has historically been used to reduce the cost of storage, I/Os from that storage, and buffer pool utilization, at the expense of the CPU required to decompress data every time it is queried. However, significant additional CPU efficiencies can be achieved by deferring decompression as late in query processing as possible and performing query processing operations directly on the still-compressed data. In this paper, we investigate the benefits and challenges of performing joins on compressed (or encoded) data. We demonstrate the benefit of independently optimizing the compression scheme of each join column, even though join predicates relating values from multiple columns may require translation of the encoding of one join column into the encoding of the other. We also show the benefit of compressing "payload" data other than the join columns "on the fly," to minimize the size of hash tables used in the join. By partitioning the domain of each column and defining separate dictionaries for each partition, we can achieve even better overall compression as well as increased flexibility in dealing with new values introduced by updates. Instead of decompressing both join columns participating in a join to resolve their different compression schemes, our system performs a light-weight mapping of only qualifying rows from one of the join columns to the encoding space of the other at run time. Consequently, join predicates can be applied directly on the compressed data. We call this procedure encoding translation. Two alternatives of encoding translation are developed and compared in the paper. We provide a comprehensive evaluation of these alternatives using product implementations of each on the TPC-H data set, and demonstrate that performing joins on encoded and partitioned data achieves both superior performance and excellent compression. © 2014 VLDB Endowment 2150-8097/14/08

DGIST Library Institutional Repository