Search CORE

6,629 research outputs found

Learning Binary Code for Fast Nearest Subspace Search

Author: Andoni
Andoni
Arya
Bai
Baraniuk
Basri
Basri
Basri
Bauml
Beveridge
Blank
Blanz
Broomhead
Broomhead
Chang
Datar
Dong
Dong
Edelman
Edwin R. Hancock
Fitzgibbon
Ghasedi Dizaji
Gionis
Goldstein
Golub
Gong
Gross
Hamm
Hotelling
Ji
Ji
Jun Zhou
Kim
Kushilevitz
Lei Zhou
Li
Lin
Lin
Liong
Liu
Liu
Liu
Luo
Marrinan
Muja
O’Hara
Pirsiavash
Shen
Song
Soomro
Sun
Vidal
Wang
Wang
Wang
Wang
Wang
Weiss
Wolf
Wright
Xianglong Liu
Xiao Bai
Xu
Zhang
Zhang
Zhang
Zhou
Publication venue: 'Elsevier BV'
Publication date: 01/02/2020
Field of study

Crossref

White Rose Research Online

Optimized Cartesian $K$ -Means

Author: Li Shipeng
Shen Heng Tao
Song Jingkuan
Wang Jianfeng
Wang Jingdong
Xu Xin-Shun
Publication venue
Publication date: 15/05/2014
Field of study

Product quantization-based approaches are effective to encode high-dimensional data points for approximate nearest neighbor search. The space is decomposed into a Cartesian product of low-dimensional subspaces, each of which generates a sub codebook. Data points are encoded as compact binary codes using these sub codebooks, and the distance between two data points can be approximated efficiently from their codes by the precomputed lookup tables. Traditionally, to encode a subvector of a data point in a subspace, only one sub codeword in the corresponding sub codebook is selected, which may impose strict restrictions on the search accuracy. In this paper, we propose a novel approach, named Optimized Cartesian

K

-Means (OCKM), to better encode the data points for more accurate approximate nearest neighbor search. In OCKM, multiple sub codewords are used to encode the subvector of a data point in a subspace. Each sub codeword stems from different sub codebooks in each subspace, which are optimally generated with regards to the minimization of the distortion errors. The high-dimensional data point is then encoded as the concatenation of the indices of multiple sub codewords from all the subspaces. This can provide more flexibility and lower distortion errors than traditional methods. Experimental results on the standard real-life datasets demonstrate the superiority over state-of-the-art approaches for approximate nearest neighbor search.Comment: to appear in IEEE TKDE, accepted in Apr. 201

arXiv.org e-Print Archive

University of Queensland eSpace

Streaming Binary Sketching based on Subspace Tracking and Diagonal Uniformization

Author: Atif Jamal
Gouy-Pailler Cédric
Morvan Anne
Souloumiac Antoine
Publication venue
Publication date: 08/02/2018
Field of study

In this paper, we address the problem of learning compact similarity-preserving embeddings for massive high-dimensional streams of data in order to perform efficient similarity search. We present a new online method for computing binary compressed representations -sketches- of high-dimensional real feature vectors. Given an expected code length

c

and high-dimensional input data points, our algorithm provides a

c

-bits binary code for preserving the distance between the points from the original high-dimensional space. Our algorithm does not require neither the storage of the whole dataset nor a chunk, thus it is fully adaptable to the streaming setting. It also provides low time complexity and convergence guarantees. We demonstrate the quality of our binary sketches through experiments on real data for the nearest neighbors search task in the online setting

arXiv.org e-Print Archive

Crossref

Scalable Image Retrieval by Sparse Product Quantization

Author: Chen Chun
Hoi Steven C. H.
Ning Qingqun
Zhong Zhiyuan
Zhu Jianke
Publication venue
Publication date: 15/03/2016
Field of study

Fast Approximate Nearest Neighbor (ANN) search technique for high-dimensional feature indexing and retrieval is the crux of large-scale image retrieval. A recent promising technique is Product Quantization, which attempts to index high-dimensional image features by decomposing the feature space into a Cartesian product of low dimensional subspaces and quantizing each of them separately. Despite the promising results reported, their quantization approach follows the typical hard assignment of traditional quantization methods, which may result in large quantization errors and thus inferior search performance. Unlike the existing approaches, in this paper, we propose a novel approach called Sparse Product Quantization (SPQ) to encoding the high-dimensional feature vectors into sparse representation. We optimize the sparse representations of the feature vectors by minimizing their quantization errors, making the resulting representation is essentially close to the original data in practice. Experiments show that the proposed SPQ technique is not only able to compress data, but also an effective encoding technique. We obtain state-of-the-art results for ANN search on four public image datasets and the promising results of content-based image retrieval further validate the efficacy of our proposed method.Comment: 12 page

arXiv.org e-Print Archive

Institutional Knowledge at Singapore Management University

Online Product Quantization

Author: Tsang Ivor W.
Xu Donna
Zhang Ying
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 24/03/2018
Field of study

Approximate nearest neighbor (ANN) search has achieved great success in many tasks. However, existing popular methods for ANN search, such as hashing and quantization methods, are designed for static databases only. They cannot handle well the database with data distribution evolving dynamically, due to the high computational effort for retraining the model based on the new database. In this paper, we address the problem by developing an online product quantization (online PQ) model and incrementally updating the quantization codebook that accommodates to the incoming streaming data. Moreover, to further alleviate the issue of large scale computation for the online PQ update, we design two budget constraints for the model to update partial PQ codebook instead of all. We derive a loss bound which guarantees the performance of our online PQ model. Furthermore, we develop an online PQ model over a sliding window with both data insertion and deletion supported, to reflect the real-time behaviour of the data. The experiments demonstrate that our online PQ model is both time-efficient and effective for ANN search in dynamic large scale databases compared with baseline methods and the idea of partial PQ codebook update further reduces the update cost.Comment: To appear in IEEE Transactions on Knowledge and Data Engineering (DOI: 10.1109/TKDE.2018.2817526

arXiv.org e-Print Archive

OPUS - University of Technology Sydney