Anomaly detection in cloud-native systems

Lenarduzzi, V. (Valentina); Li, X. (Xiaozhou); Lomio, F. (Francesco); Moreschini, S. (Sergio)

Anomaly detection in cloud-native systems

Authors: V. (Valentina) Lenarduzzi
X. (Xiaozhou) Li
F. (Francesco) Lomio
S. (Sergio) Moreschini
Publication date: 1 January 2022
Publisher

Abstract

Abstract Companies develop cloud-native systems deployed on public and private clouds. Since private clouds have limited resources, the systems should run efficiently by keeping performance related anomalies under control. The goal of this work is to understand whether a set of five performance-related KPIs depends on the metrics collected at runtime by Kafka, Zookeeper, and other tools (168 different metrics). We considered four weeks worth of runtime data collected from a system running in production. We trained eight Machine Learning algorithms on three weeks worth of data and tested them on one week’s worth of data to compare their prediction accuracy and their training and testing time. It is possible to detect performance-related anomalies with a very high level of accuracy (higher than 95% AUC) and with very limited training time (between 8 and 17 minutes). Machine Learning algorithms can help to identify runtime anomalies and to detect them efficiently. Future work will include the identification of a proactive approach to recognize the root cause of the anomalies and to prevent them as early as possible

Similar works

Full text

Available Versions

University of Oulu Repository - Jultika

oai:oulu.fi:nbnfi-fe2023032333...

Last time updated on 15/04/2023