Search CORE

302 research outputs found

GraphH: High Performance Big Graph Analytics in Small Clusters

Author: Duong Ta Nguyen Binh
Sun Peng
Wen Yonggang
Xiao Xiaokui
Publication venue
Publication date: 07/08/2017
Field of study

It is common for real-world applications to analyze big graphs using distributed graph processing systems. Popular in-memory systems require an enormous amount of resources to handle big graphs. While several out-of-core approaches have been proposed for processing big graphs on disk, the high disk I/O overhead could significantly reduce performance. In this paper, we propose GraphH to enable high-performance big graph analytics in small clusters. Specifically, we design a two-stage graph partition scheme to evenly divide the input graph into partitions, and propose a GAB (Gather-Apply-Broadcast) computation model to make each worker process a partition in memory at a time. We use an edge cache mechanism to reduce the disk I/O overhead, and design a hybrid strategy to improve the communication performance. GraphH can efficiently process big graphs in small clusters or even a single commodity server. Extensive evaluations have shown that GraphH could be up to 7.8x faster compared to popular in-memory systems, such as Pregel+ and PowerGraph when processing generic graphs, and more than 100x faster than recently proposed out-of-core systems, such as GraphD and Chaos when processing big graphs

arXiv.org e-Print Archive

Crossref

GraphMP: An Efficient Semi-External-Memory Big Graph Processing System on a Single Machine

Author: Duong Ta Nguyen Binh
Sun Peng
Wen Yonggang
Xiao Xiaokui
Publication venue
Publication date: 09/07/2017
Field of study

Recent studies showed that single-machine graph processing systems can be as highly competitive as cluster-based approaches on large-scale problems. While several out-of-core graph processing systems and computation models have been proposed, the high disk I/O overhead could significantly reduce performance in many practical cases. In this paper, we propose GraphMP to tackle big graph analytics on a single machine. GraphMP achieves low disk I/O overhead with three techniques. First, we design a vertex-centric sliding window (VSW) computation model to avoid reading and writing vertices on disk. Second, we propose a selective scheduling method to skip loading and processing unnecessary edge shards on disk. Third, we use a compressed edge cache mechanism to fully utilize the available memory of a machine to reduce the amount of disk accesses for edges. Extensive evaluations have shown that GraphMP could outperform state-of-the-art systems such as GraphChi, X-Stream and GridGraph by 31.6x, 54.5x and 23.1x respectively, when running popular graph applications on a billion-vertex graph

arXiv.org e-Print Archive

Crossref

Towards distributed machine learning in shared clusters: A dynamically-partitioned approach

Author: SUN Peng
TA Nguyen Binh Duong
WEN Yonggang
YAN Shengen
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 01/05/2017
Field of study

Crossref

Institutional Knowledge at Singapore Management University

GraphH: High performance big graph analytics in small clusters

Author: SUN Peng
TA Nguyen Binh Duong
WEN Yonggang
XIAO Xiaokui
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 08/09/2017
Field of study

Institutional Knowledge at Singapore Management University

Metaflow: a scalable metadata lookup service for distributed file systems in data centers

Author: SUN Peng
TA Nguyen Binh Duong
WEN Yonggang
XIE Haiyong
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 01/09/2016
Field of study

Institutional Knowledge at Singapore Management University

ConjNorm: Tractable Density Estimation for Out-of-Distribution Detection

Author: Fang Zhen
Li Yixuan
Luo Yadan
Peng Bo
Zhang Yonggang
Publication venue
Publication date: 27/02/2024
Field of study

Post-hoc out-of-distribution (OOD) detection has garnered intensive attention in reliable machine learning. Many efforts have been dedicated to deriving score functions based on logits, distances, or rigorous data distribution assumptions to identify low-scoring OOD samples. Nevertheless, these estimate scores may fail to accurately reflect the true data density or impose impractical constraints. To provide a unified perspective on density-based score design, we propose a novel theoretical framework grounded in Bregman divergence, which extends distribution considerations to encompass an exponential family of distributions. Leveraging the conjugation constraint revealed in our theorem, we introduce a \textsc{ConjNorm} method, reframing density function design as a search for the optimal norm coefficient

p

against the given dataset. In light of the computational challenges of normalization, we devise an unbiased and analytically tractable estimator of the partition function using the Monte Carlo-based importance sampling technique. Extensive experiments across OOD detection benchmarks empirically demonstrate that our proposed \textsc{ConjNorm} has established a new state-of-the-art in a variety of OOD detection setups, outperforming the current best method by up to 13.25

\%

and 28.19

\%

(FPR95) on CIFAR-100 and ImageNet-1K, respectively.Comment: ICLR24 poste

arXiv.org e-Print Archive

GraphMP: I/O-Efficient big graph analytics on a single commodity machine

Author: SUN Peng
TA Nguyen Binh Duong
WEN Yonggang
XIAO Xiaokui
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 01/12/2020
Field of study

Institutional Knowledge at Singapore Management University