Search CORE

1,533 research outputs found

Generic Techniques in General Purpose GPU Programming with Applications to Ant Colony and Image Processing Algorithms

Author: DAWSON LAURENCE,JAMES
Publication venue
Publication date: 01/01/2015
Field of study

In 2006 NVIDIA introduced a new unified GPU architecture facilitating general-purpose computation on the GPU. The following year NVIDIA introduced CUDA, a parallel programming architecture for developing general purpose applications for direct execution on the new unified GPU. CUDA exposes the GPU's massively parallel architecture of the GPU so that parallel code can be written to execute much faster than its sequential counterpart. Although CUDA abstracts the underlying architecture, fully utilising and scheduling the GPU is non-trivial and has given rise to a new active area of research. Due to the inherent complexities pertaining to GPU development, in this thesis we explore and find efficient parallel mappings of existing and new parallel algorithms on the GPU using NVIDIA CUDA. We place particular emphasis on metaheuristics, image processing and designing reusable techniques and mappings that can be applied to other problems and domains. We begin by focusing on Ant Colony Optimisation (ACO), a nature inspired heuristic approach for solving optimisation problems. We present a versatile improved data-parallel approach for solving the Travelling Salesman Problem using ACO resulting in significant speedups. By extending our initial work, we show how existing mappings of ACO on the GPU are unable to compete against their sequential counterpart when common CPU optimisation strategies are employed and detail three distinct candidate set parallelisation strategies for execution on the GPU. By further extending our data-parallel approach we present the first implementation of an ACO-based edge detection algorithm on the GPU to reduce the execution time and improve the viability of ACO-based edge detection. We finish by presenting a new color edge detection technique using the volume of a pixel in the HSI color space along with a parallel GPU implementation that is able to withstand greater levels of noise than existing algorithms

Durham e-Theses

Parallelization of Ant System for GPU under the PRAM Model

Author: Brodnik Andrej
Grgurovič Marko
Publication venue: Institute of Informatics, Slovak Academy of Sciences
Publication date: 03/05/2018
Field of study

We study the parallelized ant system algorithm solving the traveling salesman problem on n cities. First, following the series of recent results for the graphics processing unit, we show that they translate to the PRAM (parallel random access machine) model. In addition, we develop a novel pheromone matrix update method under the PRAM CREW (concurrent-read exclusive-write) model and translate it to the graphics processing unit without atomic instructions. As a consequence, we give new asymptotic bounds for the parallel ant system, resulting in step complexities O(n łg łg n) on CRCW (concurrent-read concurrent-write) and O(n łg n) on CREW variants of PRAM using n2 processors in both cases. Finally, we present an experimental comparison with the currently known pheromone matrix update methods on the graphics processing unit and obtain encouraging results

Computing and Informatics (E-Journal - Institute of Informatics, SAS, Bratislava)

Accelerating ant colony optimization-based edge detection on the GPU using CUDA

Author: Dawson L.
Stewart I.A.
Publication venue
Publication date: 01/07/2014
Field of study

Ant Colony Optimization (ACO) is a nature-inspired metaheuristic that can be applied to a wide range of optimization problems. In this paper we present the first parallel implementation of an ACO-based (image processing) edge detection algorithm on the Graphics Processing Unit (GPU) using NVIDIA CUDA. We extend recent work so that we are able to implement a novel data-parallel approach that maps individual ants to thread warps. By exploiting the massively parallel nature of the GPU, we are able to execute significantly more ants per ACO-iteration allowing us to reduce the total number of iterations required to create an edge map. We hope that reducing the execution time of an ACO-based implementation of edge detection will increase its viability in image processing and computer vision

Durham Research Online

Crossref

Parallelization Strategies for Ant Colony Optimisation on GPUs

Author: Amos Martyn
Cecilia Jose M.
Garcia Jose M.
Nisbet Andy
Ujaldon Manuel
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 01/01/2011
Field of study

Ant Colony Optimisation (ACO) is an effective population-based meta-heuristic for the solution of a wide variety of problems. As a population-based algorithm, its computation is intrinsically massively parallel, and it is there- fore theoretically well-suited for implementation on Graphics Processing Units (GPUs). The ACO algorithm comprises two main stages: Tour construction and Pheromone update. The former has been previously implemented on the GPU, using a task-based parallelism approach. However, up until now, the latter has always been implemented on the CPU. In this paper, we discuss several parallelisation strategies for both stages of the ACO algorithm on the GPU. We propose an alternative data-based parallelism scheme for Tour construction, which fits better on the GPU architecture. We also describe novel GPU programming strategies for the Pheromone update stage. Our results show a total speed-up exceeding 28x for the Tour construction stage, and 20x for Pheromone update, and suggest that ACO is a potentially fruitful area for future research in the GPU domain.Comment: Accepted by 14th International Workshop on Nature Inspired Distributed Computing (NIDISC 2011), held in conjunction with the 25th IEEE/ACM International Parallel and Distributed Processing Symposium (IPDPS 2011

arXiv.org e-Print Archive

CiteSeerX

Northumbria Research Link

Crossref

Ant Colony Optimization

Author
Publication venue: 'IntechOpen'
Publication date: 20/04/2021
Field of study

Ant Colony Optimization (ACO) is the best example of how studies aimed at understanding and modeling the behavior of ants and other social insects can provide inspiration for the development of computational algorithms for the solution of difficult mathematical problems. Introduced by Marco Dorigo in his PhD thesis (1992) and initially applied to the travelling salesman problem, the ACO field has experienced a tremendous growth, standing today as an important nature-inspired stochastic metaheuristic for hard optimization problems. This book presents state-of-the-art ACO methods and is divided into two parts: (I) Techniques, which includes parallel implementations, and (II) Applications, where recent contributions of ACO to diverse fields, such as traffic congestion and control, structural optimization, manufacturing, and genomics are presented

Directory of Open Access Books (DOAB)

Distributed evolutionary algorithms and their models: A survey of the state-of-the-art

Author: Alba
Alba
Alba
Alba
Alba
Anglano
Apolloni
Bai
Bollini
Bouvry
Branke
Burczynski
Burczyński
Cahon
Cahon
Cantu-Paz
Cantu-Paz
Cantú-Paz
Chatzimilioudis
Chen
Creput
Danoy
Davis
de Toro Negro
Dean
Deb
Decraene
Desell
Dorronsoro
Du
Dubreuil
Durillo
Durillo
Durillo
Durillo
Durillo
Epitropakis
Escuela
Ewald
Fok
Folino
Folino
Gagné
Garcia-Arenas
García-Arenas
Giacobini
Giacobini
Giacobini
Giacobini
Giacobini
Giacobini
Goh
Goldberg
Gong
Gonzalez
Herrera
Herrera
Hidalgo
Hidalgo
Hosseini
Iimura
Ishimizu
Ismail
Jin
Jing-Jing Li
Johar
Jun Zhang
Kattan
Kattan
Kirley
Kirley
Kwok
Laredo
Li
Liang
Lienig
Lim
Lim
Liu
Llora
Lorion
Manfrin
McNabb
Melab
Melab
Mendiburu
Merelo
Merelo-Guervos
Merelo-Guervós
Merelo-Guervós
Michel
Mostaghim
Mussi
Nebro
Nebro
Nesmachnow
Nojima
Ordeshook
Pedemonte
Pendharkar
Pierreval
Piriyakumar
Potter
Qingfu Zhang
Ray
Robilliard
Roy
Ruiz-Andino
Said
Said
Sarma
Schutte
Schönfisch
Scriven
Sefrioui
Seredynski
Seredynski
Seredynski
Sherry
Soca
Starzynski
Stützle
Su
Subbu
Subbu
Suganthan
Tagawa
Tan
Tan
Tan
Tan
Tan
Tasoulis
Tomassini
Tomassini
Umbarkar
Van Veldhuizen
Verma
Veronese
Vlachogiannis
Weber
Weber
Wei-Neng Chen
Whitley
Wickramasinghe
Wu
Xiong
Xu
Yang
Yu
Yu
Yue-Jiao Gong
Yun Li
Zhang
Zhang
Zhang
Zhao
Zhao
Zhi-Hui Zhan
Zhou
Zhou
Zhu
Publication venue: 'Elsevier BV'
Publication date: 11/05/2015
Field of study

The increasing complexity of real-world optimization problems raises new challenges to evolutionary computation. Responding to these challenges, distributed evolutionary computation has received considerable attention over the past decade. This article provides a comprehensive survey of the state-of-the-art distributed evolutionary algorithms and models, which have been classified into two groups according to their task division mechanism. Population-distributed models are presented with master-slave, island, cellular, hierarchical, and pool architectures, which parallelize an evolution task at population, individual, or operation levels. Dimension-distributed models include coevolution and multi-agent models, which focus on dimension reduction. Insights into the models, such as synchronization, homogeneity, communication, topology, speedup, advantages and disadvantages are also presented and discussed. The study of these models helps guide future development of different and/or improved algorithms. Also highlighted are recent hotspots in this area, including the cloud and MapReduce-based implementations, GPU and CUDA-based implementations, distributed evolutionary multiobjective optimization, and real-world applications. Further, a number of future research directions have been discussed, with a conclusion that the development of distributed evolutionary computation will continue to flourish

University of Essex Research Repository

Crossref

Enlighten