Search CORE

48 research outputs found

EPCC's Exascale journey: a retrospective of the past 10 years and a vision of the future

Author: Parsons Mark
Weiland Michele
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 11/10/2021
Field of study

Edinburgh Research Explorer

Progressive Load Balancing in Distributed Memory

Author: Weiland Michele
Zarins Justs
Publication venue: 'IOS Press'
Publication date: 01/04/2020
Field of study

Edinburgh Research Explorer

Exploiting the Performance Benefits of Storage Class Memory for HPC and HPDA Workflows

Author: Jackson Adrian
Johnson Nick
Parsons Mark
Weiland Michele
Publication venue: 'FSAEIHE South Ural State University (National Research University)'
Publication date: 30/04/2018
Field of study

Crossref

Edinburgh Research Explorer

Vectorizing and distributing number-theoretic transform to count Goldbach partitions on Arm-based supercomputers

Author: Jesus Ricardo
Oliveira e Silva Tomas
Weiland Michele
Publication venue
Publication date: 14/08/2023
Field of study

In this article, we explore the usage of scalable vector extension (SVE) to vectorize number-theoretic transforms (NTTs). In particular, we show that 64-bit modular arithmetic operations, including modular multiplication, can be efficiently implemented with SVE instructions. The vectorization of NTT loops and kernels involving 64-bit modular operations was not possible in previous Arm-based single instruction multiple data architectures since these architectures lacked crucial instructions to efficiently implement modular multiplication. We test and evaluate our SVE implementation on the A64FX processor in an HPE Apollo 80 system. Furthermore, we implement a distributed NTT for the computation of large-scale exact integer convolutions. We evaluate this transform on HPE Apollo 70, Cray XC50, HPE Apollo 80, and HPE Cray EX systems, where we demonstrate good scalability to thousands of cores. Finally, we describe how these methods can be utilized to count the number of Goldbach partitions of all even numbers to large limits. We present some preliminary results concerning this problem, in particular a histogram of the number of Goldbach partitions of the even numbers up to 2 40.</p

Edinburgh Research Explorer

Evaluation Methodology of an NVRAM-based Platform for the Exascale

Author: Herrera Juan F. R.
Parsons Mark
Prabhakaran Suraj
Weiland Michele
Publication venue
Publication date: 01/01/2018
Field of study

Edinburgh Research Explorer

Morpheus unleashed: Fast cross-platform SpMV on emerging architectures

Author: Brown Nick
Jesus Ricardo
Klaisoongnoen Mark
Stylianou Christodoulos
Weiland Michele
Publication venue
Publication date: 11/05/2023
Field of study

Edinburgh Research Explorer

Detecting scale-induced overflow bugs in production HPC codes

Author: Bartholomew Paul
Lapworth Leigh
Parsons Mark
Weiland Michele
Zarins Justs
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 04/01/2023
Field of study

Edinburgh Research Explorer

Investigating applications on the A64FX

Author: Brown Nick
Jackson Adrian
Parsons Mark
Turner Andrew
Weiland Michele
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 24/09/2020
Field of study

The A64FX processor from Fujitsu, being designed for computational simulation and machine learning applications, has the potential for unprecedented performance in HPC systems. In this paper, we evaluate the A64FX by benchmarking against a range of production HPC platforms that cover a number of processor technologies. We investigate the performance of complex scientific applications across multiple nodes, as well as single node and mini-kernel benchmarks. This paper finds that the performance of the A64FX processor across our chosen benchmarks often significantly exceeds other platforms, even without specific application optimisations for the processor instruction set or hardware. However, this is not true for all the benchmarks we have undertaken. Furthermore, the specific configuration of applications can have an impact on the runtime and performance experienced

arXiv.org e-Print Archive

Crossref

Edinburgh Research Explorer