Search CORE

2,508 research outputs found

Assessment, Design and Implementation of a Private Cloud for MapReduce Applications

Author: Cabaleiro J.C.
Fernández Pena Tomás
González Patricia
Salgueiro M.
Publication venue: 'Scientific Research Publishing, Inc.'
Publication date: 01/01/2014
Field of study

[Abstract] Scientific computation and data intensive analyses are ever more frequent. On the one hand, the MapReduce programming model has gained a lot of attention for its applicability in large parallel data analyses and Big Data applications. On the other hand, Cloud computing seems to be increasingly attractive in solving these computing problems that demand a lot of resources. This paper explores the potential symbiosis between MapReduce and Cloud Computing, in order to create a robust and scalable environment to execute MapReduce workflows regardless of the underlaying infrastructure. The main goal of this work is to provide an easy-to-install interface, so as non-expert scientists can deploy a suitable testbed for their MapReduce experiments on local resources of their institution. Testing cases were performed in order to evaluate the required time for the whole executing process on a real cluster

Repositorio da Universidade da Coruña

LAReferencia - Red Federada de Repositorios Institucionales de Publicaciones Científicas Latinoamericanas

Crossref

Assessment, Design and Implementation of a Private Cloud for MapReduce Applications

Author: Cabaleiro Domínguez José Carlos
Fernández Pena Anselmo Tomás
González Gómez Patricia
Salgueiro M.
Publication venue: 'Scientific Research Publishing, Inc.'
Publication date: 01/01/2014
Field of study

Scientific computation and data intensive analyses are ever more frequent. On the one hand, the MapReduce programming model has gained a lot of attention for its applicability in large parallel data analyses and Big Data applications. On the other hand, Cloud computing seems to be increasingly attractive in solving these computing problems that demand a lot of resources. This paper explores the potential symbiosis between MapReduce and Cloud Computing, in order to create a robust and scalable environment to execute MapReduce workflows regardless of the underlaying infrastructure. The main goal of this work is to provide an easy-to-install interface, so as non-expert scientists can deploy a suitable testbed for their MapReduce experiments on local resources of their institution. Testing cases were performed in order to evaluate the required time for the whole executing process on a real clusterS

LAReferencia - Red Federada de Repositorios Institucionales de Publicaciones Científicas Latinoamericanas

Repositorio Institucional da Universidade de Santiago de Compostela

A STUDY FOR HANDELLING OF HIGH-PERFORMANCE CLIMATE DATA USING HADOOP

Author: . Ajay K P
. Dr Nagesh H R
. Dr. K C Gouda
Publication venue: International Journal of Innovative Technology and Research
Publication date: 16/04/2015
Field of study

Introduction of Hadoop has become the de factor for large-scale data analysis in commercial applications, and nowadays finding its prominence in scientific applications also. In climate research, where there is a need for high-performance analytics, the Hadoop MapReduce may be useful in providing solution to data intensive problems. It makes use of parallel computation for analysis paradigm that uses clusters of computers and combines distributed storage of large data sets. This paper presents the potential of MapReduce for scientific data sets which is in the NetCDF format, and performs basic operations common to a wide range of analyses. This provides a prototype for series of canonical MapReduce operations for number of observational and climate simulation datasets. Our work provides solutions on how to tackle with arbitrary spatial and temporal global climate data which is in NetCDF form. This approach can improve efficiencies within data intensive analytic workflows

International Journal of Innovative Technology and Research (IJITR)

MapReduce in the Clouds for Science

Author: Geoffrey Fox
Judy Qiu
Tak-lon Wu
Thilina Gunarathne
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 01/01/2010
Field of study

Abstract — The utility computing model introduced by cloud computing combined with the rich set of cloud infrastructure services offers a very viable alternative to traditional servers and computing clusters. MapReduce distributed data processing architecture has become the weapon of choice for data-intensive analyses in the clouds and in commodity clusters due to its excellent fault tolerance features, scalability and the ease of use. Currently, there are several options for using MapReduce in cloud environments, such as using MapReduce as a service, setting up one’s own MapReduce cluster on cloud instances, or using specialized cloud MapReduce runtimes that take advantage of cloud infrastructure services. In this paper, we introduce AzureMapReduce, a novel MapReduce runtime built using the Microsoft Azure cloud infrastructure services. AzureMapReduce architecture successfully leverages the high latency, eventually consistent, yet highly scalable Azure infrastructure services to provide an efficient, on demand alternative to traditional MapReduce clusters. Further we evaluate the use and performance of MapReduce frameworks, including AzureMapReduce, in cloud environments for scientific applications using sequence assembly and sequence alignment as use cases

CiteSeerX

Crossref

Scientific Computing Meets Big Data Technology: An Astronomy Use Case

Author: Barbary Kyle
Franklin Michael J.
Nothaft Frank Austin
Patterson David A.
Perlmutter Saul
Sparks Evan
Zahn Oliver
Zhang Zhao
Publication venue
Publication date: 22/12/2015
Field of study

Scientific analyses commonly compose multiple single-process programs into a dataflow. An end-to-end dataflow of single-process programs is known as a many-task application. Typically, tools from the HPC software stack are used to parallelize these analyses. In this work, we investigate an alternate approach that uses Apache Spark -- a modern big data platform -- to parallelize many-task applications. We present Kira, a flexible and distributed astronomy image processing toolkit using Apache Spark. We then use the Kira toolkit to implement a Source Extractor application for astronomy images, called Kira SE. With Kira SE as the use case, we study the programming flexibility, dataflow richness, scheduling capacity and performance of Apache Spark running on the EC2 cloud. By exploiting data locality, Kira SE achieves a 2.5x speedup over an equivalent C program when analyzing a 1TB dataset using 512 cores on the Amazon EC2 cloud. Furthermore, we show that by leveraging software originally designed for big data infrastructure, Kira SE achieves competitive performance to the C implementation running on the NERSC Edison supercomputer. Our experience with Kira indicates that emerging Big Data platforms such as Apache Spark are a performant alternative for many-task scientific applications

arXiv.org e-Print Archive

Crossref

eScholarship - University of California