Search CORE

8,487 research outputs found

The Case for Learned Index Structures

Author: Abadi M.
Armbrust M.
Böhm M.
Chang F.
Goodfellow I.
Grossi R.
Lehman T. J.
Litwin W.
Magdon-Ismail M.
Miller D. J.
Moerkotte G.
Sutskever I.
You S.
Publication venue
Publication date: 30/04/2018
Field of study

Indexes are models: a B-Tree-Index can be seen as a model to map a key to the position of a record within a sorted array, a Hash-Index as a model to map a key to a position of a record within an unsorted array, and a BitMap-Index as a model to indicate if a data record exists or not. In this exploratory research paper, we start from this premise and posit that all existing index structures can be replaced with other types of models, including deep-learning models, which we term learned indexes. The key idea is that a model can learn the sort order or structure of lookup keys and use this signal to effectively predict the position or existence of records. We theoretically analyze under which conditions learned indexes outperform traditional index structures and describe the main challenges in designing learned index structures. Our initial results show, that by using neural nets we are able to outperform cache-optimized B-Trees by up to 70% in speed while saving an order-of-magnitude in memory over several real-world data sets. More importantly though, we believe that the idea of replacing core components of a data management system through learned models has far reaching implications for future systems designs and that this work just provides a glimpse of what might be possible

arXiv.org e-Print Archive

Crossref

Modelling and trading the Greek stock market with gene expression and genetic programing algorithms

Author: Dunis Christian
Karathanasopoulos Andreas
Laws Jason
Sermpinis Georgios
Publication venue: 'Wiley'
Publication date: 25/09/2014
Field of study

This paper presents an application of the gene expression programming (GEP) and integrated genetic programming (GP) algorithms to the modelling of ASE 20 Greek index. GEP and GP are robust evolutionary algorithms that evolve computer programs in the form of mathematical expressions, decision trees or logical expressions. The results indicate that GEP and GP produce significant trading performance when applied to ASE 20 and outperform the well-known existing methods. The trading performance of the derived models is further enhanced by applying a leverage filter

Enlighten

Bridging the gap between algorithmic and learned index structures

Author: Hadian Ali
Publication venue: Computing, Imperial College London
Publication date: 01/07/2022
Field of study

Index structures such as B-trees and bloom filters are the well-established petrol engines of database systems. However, these structures do not fully exploit patterns in data distribution. To address this, researchers have suggested using machine learning models as electric engines that can entirely replace index structures. Such a paradigm shift in data system design, however, opens many unsolved design challenges. More research is needed to understand the theoretical guarantees and design efficient support for insertion and deletion. In this thesis, we adopt a different position: index algorithms are good enough, and instead of going back to the drawing board to fit data systems with learned models, we should develop lightweight hybrid engines that build on the benefits of both algorithmic and learned index structures. The indexes that we suggest provide the theoretical performance guarantees and updatability of algorithmic indexes while using position prediction models to leverage the data distributions and thereby improve the performance of the index structure. We investigate the potential for minimal modifications to algorithmic indexes such that they can leverage data distribution similar to how learned indexes work. In this regard, we propose and explore the use of helping models that boost classical index performance using techniques from machine learning. Our suggested approach inherits performance guarantees from its algorithmic baseline index, but at the same time it considers the data distribution to improve performance considerably. We study single-dimensional range indexes, spatial indexes, and stream indexing, and show that the suggested approach results in range indexes that outperform the algorithmic indexes and have comparable performance to the read-only, fully learned indexes and hence can be reliably used as a default index structure in a database engine. Besides, we consider the updatability of the indexes and suggest solutions for updating the index, notably when the data distribution drastically changes over time (e.g., for indexing data streams). In particular, we propose a specific learning-augmented index for indexing a sliding window with timestamps in a data stream. Additionally, we highlight the limitations of learned indexes for low-latency lookup on real- world data distributions. To tackle this issue, we suggest adding an algorithmic enhancement layer to a learned model to correct the prediction error with a small memory latency. This approach enables efficient modelling of the data distribution and resolves the local biases of a learned model at the cost of roughly one memory lookup.Open Acces

Spiral - Imperial College Digital Repository

Benchmarking Learned Indexes

Author: Kemper Alfons
Kipf Andreas
Kraska Tim
Marcus Ryan
Misra Sanchit
Neumann Thomas
Stoian Mihail
van Renen Alexander
Publication venue
Publication date: 29/06/2020
Field of study

Recent advancements in learned index structures propose replacing existing index structures, like B-Trees, with approximate learned models. In this work, we present a unified benchmark that compares well-tuned implementations of three learned index structures against several state-of-the-art "traditional" baselines. Using four real-world datasets, we demonstrate that learned index structures can indeed outperform non-learned indexes in read-only in-memory workloads over a dense array. We also investigate the impact of caching, pipelining, dataset size, and key size. We study the performance profile of learned index structures, and build an explanation for why learned models achieve such good performance. Finally, we investigate other important properties of learned index structures, such as their performance in multi-threaded systems and their build times

arXiv.org e-Print Archive

DSpace@MIT

A Survey of Symbolic Execution Techniques

Author: Baldoni Roberto
Coppa Emilio
D'Elia Daniele Cono
Demetrescu Camil
Finocchi Irene
Publication venue
Publication date: 01/01/2018
Field of study

Many security and software testing applications require checking whether certain properties of a program hold for any possible usage scenario. For instance, a tool for identifying software vulnerabilities may need to rule out the existence of any backdoor to bypass a program's authentication. One approach would be to test the program using different, possibly random inputs. As the backdoor may only be hit for very specific program workloads, automated exploration of the space of possible inputs is of the essence. Symbolic execution provides an elegant solution to the problem, by systematically exploring many possible execution paths at the same time without necessarily requiring concrete inputs. Rather than taking on fully specified input values, the technique abstractly represents them as symbols, resorting to constraint solvers to construct actual instances that would cause property violations. Symbolic execution has been incubated in dozens of tools developed over the last four decades, leading to major practical breakthroughs in a number of prominent software reliability applications. The goal of this survey is to provide an overview of the main ideas, challenges, and solutions developed in the area, distilling them for a broad audience. The present survey has been accepted for publication at ACM Computing Surveys. If you are considering citing this survey, we would appreciate if you could use the following BibTeX entry: http://goo.gl/Hf5FvcComment: This is the authors pre-print copy. If you are considering citing this survey, we would appreciate if you could use the following BibTeX entry: http://goo.gl/Hf5Fv

arXiv.org e-Print Archive

Archivio della ricerca- LUISS Libera Università Internazionale degli Studi Sociali Guido Carli di Roma

Archivio della ricerca- Università di Roma La Sapienza

Machine learning-guided synthesis of advanced inorganic materials

Author: Chouhan Tushar
Golani Prafful
Guan Cuntai
Liu Zheng
Lu Yuhao
Tang Bijun
Wang Han
Xu Manzhang
Xu Quan
Zhou Jiadong
Publication venue: 'Elsevier BV'
Publication date: 10/05/2019
Field of study

Synthesis of advanced inorganic materials with minimum number of trials is of paramount importance towards the acceleration of inorganic materials development. The enormous complexity involved in existing multi-variable synthesis methods leads to high uncertainty, numerous trials and exorbitant cost. Recently, machine learning (ML) has demonstrated tremendous potential for material research. Here, we report the application of ML to optimize and accelerate material synthesis process in two representative multi-variable systems. A classification ML model on chemical vapor deposition-grown MoS2 is established, capable of optimizing the synthesis conditions to achieve higher success rate. While a regression model is constructed on the hydrothermal-synthesized carbon quantum dots, to enhance the process-related properties such as the photoluminescence quantum yield. Progressive adaptive model is further developed, aiming to involve ML at the beginning stage of new material synthesis. Optimization of the experimental outcome with minimized number of trials can be achieved with the effective feedback loops. This work serves as proof of concept revealing the feasibility and remarkable capability of ML to facilitate the synthesis of inorganic materials, and opens up a new window for accelerating material development

arXiv.org e-Print Archive

DR-NTU (Digital Repository of NTU)

Air quality and urban sustainable development: the application of machine learning tools

Author: A Al-Dabbous
A Kadiyala
A Kadiyala
A Sayegh
A Suárez
B Paas
B Sierra
B Wang
B-Ch Liu
C Cruz
C Madu
Corani
D Antanasijević
D Gounaridis
D Ma
F Franceschi
F Tzima
G Cervone
G Pandey
H Karimian
H Peng
J Holloway
K de Hoogh
K Gibert
K Lässig
K Mellos
K Shaban
K Singh
L Zeng
M Krzyzanowski
M Lubell
M Oprea
M Pérez-Ortíz
N García
N Yeganeh
O Toumi
P Ifaei
R Souza
R Zalakeviciute
S Chen
S Saeed
W Tamas
W Wang
WCED
X Wang
XY Ni
Y Phillis
Y Zhan
Y Zhou
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/04/2021
Field of study

[EN] Air quality has an efect on a population¿s quality of life. As a dimension of sustainable urban development, governments have been concerned about this indicator. This is refected in the references consulted that have demonstrated progress in forecasting pollution events to issue early warnings using conventional tools which, as a result of the new era of big data, are becoming obsolete. There are a limited number of studies with applications of machine learning tools to characterize and forecast behavior of the environmental, social and economic dimensions of sustainable development as they pertain to air quality. This article presents an analysis of studies that developed machine learning models to forecast sustainable development and air quality. Additionally, this paper sets out to present research that studied the relationship between air quality and urban sustainable development to identify the reliability and possible applications in diferent urban contexts of these machine learning tools. To that end, a systematic review was carried out, revealing that machine learning tools have been primarily used for clustering and classifying variables and indicators according to the problem analyzed, while tools such as artifcial neural networks and support vector machines are the most widely used to predict diferent types of events. The nonlinear nature and synergy of the dimensions of sustainable development are of great interest for the application of machine learning tools.Molina-Gómez, NI.; Díaz-Arévalo, JL.; López Jiménez, PA. (2021). Air quality and urban sustainable development: the application of machine learning tools. International Journal of Environmental Science and Technology. 18(4):1-18. https://doi.org/10.1007/s13762-020-02896-6S118184Al-Dabbous A, Kumar P, Khan A (2017) Prediction of airborne nanoparticles at roadside location using a feed–forward artificial neural network. Atmos Pollut Res 8:446–454. https://doi.org/10.1016/j.apr.2016.11.004Antanasijević D, Pocajt V, Povrenović D, Ristić M, Perić-Grujić A (2013) PM10 emission forecasting using artificial neural networks and genetic algorithm input variable optimization. Sci Total Environ 443:511–519. https://doi.org/10.1016/j.scitotenv.2012.10.110Brink H, Richards JW, Fetherolf M (2016) Real-world machine learning. Richards JW, Fetherolf M (eds) Manning Publications Co. Berkeley, CA. https://www.manning.com/books/real-world-machine-learning. Accessed 26 Apr 2020Cervone G, Franzese P, Ezber Y, Boybeyi Z (2008) Risk assessment of atmospheric emissions using machine learning. Nat Hazard Earth Syst 8:991–1000. https://doi.org/10.5194/nhess-8-991-2008Chen S, Kan G, Li J, Liang K, Hong Y (2018) Investigating China’s urban air quality using big data, information theory, and machine learning. Pol J Environ Stud 27:565–578. https://doi.org/10.15244/pjoes/75159Corani (2005) Air quality prediction in Milan: feed-forward neural networks, pruned neural networks and lazy learning. Ecol Model 185:513–529. https://doi.org/10.1016/j.ecolmodel.2005.01.008Cruz C, Gómez A, Ramírez L, Villalva A, Monge O, Varela J, Quiroz J, Duarte H (2017) Calidad del aire respecto de metales (Pb, Cd, Ni, Cu, Cr) y relación con salud respiratoria: caso Sonora, México. Rev Int Contam Ambient 33:23–34. https://doi.org/10.20937/RICA.2017.33.esp02.02de Hoogh K, Héritier H, Stafoggia M, Künzli N, Kloog I (2018) Modelling daily PM2.5 concentrations at high spatio-temporal resolution across Switzerland. Environ Pollut 233:1147–1154. https://doi.org/10.1016/j.envpol.2017.10.025Franceschi F, Cobo M, Figueredo M (2018) Discovering relationships and forecasting PM10 and PM2.5 concentrations in Bogotá, Colombia, using Artificial Neural Networks, Principal Component Analysis, and k-means clustering. Atmos Pollut Res 9:912–922. https://doi.org/10.1016/j.apr.2018.02.006García N, Combarro E, del Coz J, Montañes E (2013) A SVM-based regression model to study the air quality at local scale in Oviedo urban area (Northern Spain): a case study. Appl Math Comput 219:8923–8937. https://doi.org/10.1016/j.amc.2013.03.018Gibert K, Sànchez-Màrre M, Sevilla B (2012) Tools for environmental data mining and intelligent decision support. In iEMSs. Leipzig, Germany. http://www.iemss.org/society/index.php/iemss-2012-proceedings. Accessed 26 Nov 2018Gibert K, Sànchez-Marrè M, Izquierdo J (2016) A survey on pre-processing techniques: relevant issues in the context of environmental data mining. Ai Commun 29:627–663. https://doi.org/10.3233/AIC-160710Gounaridis D, Chorianopoulos I, Koukoulas S (2018) Exploring prospective urban growth trends under different economic outlooks and land-use planning scenarios: the case of Athens. Appl Geogr 90:134–144. https://doi.org/10.1016/j.apgeog.2017.12.001Holloway J, Mengersen K (2018) Statistical machine learning methods and remote sensing for sustainable development goals: a review. Remote Sens 10:1–21. https://doi.org/10.3390/rs10091365Ifaei P, Karbassi A, Lee S, Yoo Ch (2017) A renewable energies-assisted sustainable development plan for Iran using techno-econo-socio-environmental multivariate analysis and big data. Energy Convers Manag 153:257–277. https://doi.org/10.1016/j.enconman.2017.10.014Kadiyala A, Kumar A (2017a) Applications of R to evaluate environmental data science problems. Environ Prog Sustain 36:1358–1364. https://doi.org/10.1002/ep.12676Kadiyala A, Kumar A (2017b) Vector time series-based radial basis function neural network modeling of air quality inside a public transportation bus using available software. Environ Prog Sustain 36:4–10. https://doi.org/10.1002/ep.12523Karimian H, Li Q, Wu Ch, Qi Y, Mo Y, Chen G, Zhang X, Sachdeva S (2019) Evaluation of different machine learning approaches to forecasting PM2.5 mass concentrations. Aerosol Air Qual Res 19:1400–1410. https://doi.org/10.4209/aaqr.2018.12.0450Krzyzanowski M, Apte J, Bonjour S, Brauer M, Cohen A, Prüss-Ustun A (2014) Air pollution in the mega-cities. Curr Environ Health Rep 1:185–191. https://doi.org/10.1007/s40572-014-0019-7Lässig K, Morik (2016) Computat sustainability. Springer, Berlin. https://doi.org/10.1007/978-3-319-31858-5Li Y, Wu Y-X, Zeng Z-X, Guo L (2006) Research on forecast model for sustainable development of economy-environment system based on PCA and SVM. In: Proceedings of the 2006 international conference on machine learning and cybernetics, vol 2006. IEEE, Dalian, China, pp 3590–3593. https://doi.org/10.1109/ICMLC.2006.258576Liu B-Ch, Binaykia A, Chang P-Ch, Tiwari M, Tsao Ch-Ch (2017) Urban air quality forecasting based on multi- dimensional collaborative support vector regression (SVR): a case study of Beijing-Tianjin-Shijiazhuang. PLoS ONE 12:1–17. https://doi.org/10.1371/journal.pone.0179763Lubell M, Feiock R, Handy S (2009) City adoption of environmentally sustainable policies in California’s Central Valley. J Am Plan Assoc 75:293–308. https://doi.org/10.1080/01944360902952295Ma D, Zhang Z (2016) Contaminant dispersion prediction and source estimation with integrated Gaussian-machine learning network model for point source emission in atmosphere. J Hazard Mater 311:237–245. https://doi.org/10.1016/j.jhazmat.2016.03.022Madu C, Kuei N, Lee P (2017) Urban sustainability management: a deep learning perspective. Sustain Cities Soc 30:1–17. https://doi.org/10.1016/j.scs.2016.12.012Mellos K (1988) Theory of eco-development. In: Perspectives on ecology. Palgrave Macmillan, London. https://doi.org/10.1007/978-1-349-19598-5_4Ni XY, Huang H, Du WP (2017) Relevance analysis and short-term prediction of PM2.5 concentrations in Beijing based on multi-source data. Atmos Environ 150:146–161. https://doi.org/10.1016/j.atmosenv.2016.11.054Oprea M, Dragomir E, Popescu M, Mihalache S (2016) Particulate matter air pollutants forecasting using inductive learning approach. Rev Chim 67:2075–2081Paas B, Stienen J, Vorländer M, Schneider Ch (2017) Modelling of urban near-road atmospheric PM concentrations using an artificial neural network approach with acoustic data input. Environments 4:1–25. https://doi.org/10.3390/environments4020026Pandey G, Zhang B, Jian L (2013) Predicting submicron air pollution indicators: a machine learning approach. Environ Sci Proc Impacts 15:996–1005. https://doi.org/10.1039/c3em30890aPeng H, Lima A, Teakles A, Jin J, Cannon A, Hsieh W (2017) Evaluating hourly air quality forecasting in Canada with nonlinear updatable machine learning methods. Air Qual Atmos Health 10:195–211. https://doi.org/10.1007/s11869-016-0414-3Pérez-Ortíz M, de La Paz-Marín M, Gutiérrez PA, Hervás-Martínez C (2014) Classification of EU countries’ progress towards sustainable development based on ordinal regression techniques. Knowl Based Syst 66:178–189. https://doi.org/10.1016/j.knosys.2014.04.041Phillis Y, Kouikoglou V, Verdugo C (2017) Urban sustainability assessment and ranking of cities. Comput Environ Urban 64:254–265. https://doi.org/10.1016/j.compenvurbsys.2017.03.002Saeed S, Hussain L, Awan I, Idris A (2017) Comparative analysis of different statistical methods for prediction of PM2.5 and PM10 concentrations in advance for several hours. Int J Comput Sci Netw Secur 17:45–52Sayegh A, Munir S, Habeebullah T (2014) Comparing the performance of statistical models for predicting PM10 concentrations. Aerosol Air Qual Res 14:653–665. https://doi.org/10.4209/aaqr.2013.07.0259Shaban K, Kadri A, Rezk E (2016) Urban air pollution monitoring system with forecasting models. IEEE Sens J 16:2598–2606. https://doi.org/10.1109/JSEN.2016.2514378Sierra B (2006) Aprendizaje automático conceptos básicos y avanzados Aspectos prácticos utilizando el software Weka. Madrid Pearson Prentice Hall, MadridSingh K, Gupta S, Rai P (2013) Identifying pollution sources and predicting urban air quality using ensemble learning methods. Atmos Environ 80:426–437. https://doi.org/10.1016/j.atmosenv.2013.08.023Song L, Pang S, Longley I, Olivares G, Sarrafzadeh A (2014) Spatio-temporal PM2.5 prediction by spatial data aided incremental support vector regression. In: International joint conference on neural networks. IEEE, Beijing, pp 623–630. https://doi.org/10.1109/IJCNN.2014.6889521Souza R, Coelho G, da Silva A, Pozza S (2015) Using ensembles of artificial neural networks to improve PM10 forecasts. Chem Eng Trans 43:2161–2166. https://doi.org/10.3303/CET1543361Suárez A, García PJ, Riesgo P, del Coz JJ, Iglesias-Rodríguez FJ (2011) Application of an SVM-based regression model to the air quality study at local scale in the Avilés urban area (Spain). Math Comput Model 54:453–1466. https://doi.org/10.1016/j.mcm.2011.04.017Tamas W, Notton G, Paoli C, Nivet M, Voyant C (2016) Hybridization of air quality forecasting models using machine learning and clustering: an original approach to detect pollutant peaks. Aerosol Air Qual Res 16:405–416. https://doi.org/10.4209/aaqr.2015.03.0193Toumi O, Le Gallo J, Ben Rejeb J (2017) Assessment of Latin American sustainability. Renew Sustain Energy Rev 78:878–885. https://doi.org/10.1016/j.rser.2017.05.013Tzima F, Mitkas P, Voukantsis D, Karatzas K (2011) Sparse episode identification in environmental datasets: the case of air quality assessment. Expert Syst Appl 38:5019–5027. https://doi.org/10.1016/j.eswa.2010.09.148United Nations, Department of Economic and Social Affairs (2019) World urbanization prospects The 2018 Revision. New York. https://doi.org/10.18356/b9e995fe-enWang B (2019) Applying machine-learning methods based on causality analysis to determine air quality in China. Pol J Environ Stud 28:3877–3885. https://doi.org/10.15244/pjoes/99639Wang X, Xiao Z (2017) Regional eco-efficiency prediction with support vector spatial dynamic MIDAS. J Clean Prod 161:165–177. https://doi.org/10.1016/j.jclepro.2017.05.077Wang W, Men C, Lu W (2008) Online prediction model based on support vector machine. Neurocomputing 71:550–558. https://doi.org/10.1016/j.neucom.2007.07.020WCED (1987) Report of the world commission on environment and development: our common future: report of the world commission on environment and development. WCED, Oslo. https://doi.org/10.1080/07488008808408783Weizhen H, Zhengqiang L, Yuhuan Z, Hua X, Ying Z, Kaitao L, Donghui L, Peng W, Yan M (2014) Using support vector regression to predict PM10 and PM2.5. In: IOP conference series: earth and environmental science, vol 17. IOP. https://doi.org/10.1088/1755-1315/17/1/012268WHO (2016) OMS | La OMS publica estimaciones nacionales sobre la exposición a la contaminación del aire y sus repercusiones para la salud. WHO. http://www.who.int/mediacentre/news/releases/2016/air-pollution-estimates/es/. Accesed 26 Nov 2018Yeganeh N, Shafie MP, Rashidi Y, Kamalan H (2012) Prediction of CO concentrations based on a hybrid partial least square and support vector machine model. Atmos Environ 55:357–365. https://doi.org/10.1016/j.atmosenv.2012.02.092Zalakeviciute R, Bastidas M, Buenaño A, Rybarczyk Y (2020) A traffic-based method to predict and map urban air quality. Appl Sci. https://doi.org/10.3390/app10062035Zeng L, Guo J, Wang B, Lv J, Wang Q (2019) Analyzing sustainability of Chinese coal cities using a decision tree modeling approach. Resour Policy 64:101501. https://doi.org/10.1016/j.resourpol.2019.101501Zhan Y, Luo Y, Deng X, Grieneisen M, Zhang M, Di B (2018) Spatiotemporal prediction of daily ambient ozone levels across China using random forest for human exposure assessment. Environ Pollut 233:464–473. https://doi.org/10.1016/j.envpol.2017.10.029Zhang Y, Huan Q (2006) Research on the evaluation of sustainable development in Cangzhou city based on neural-network-AHP. In: Proceedings of the fifth international conference on machine learning and cybernetics, vol 2006. pp 3144–3147. https://doi.org/10.1109/ICMLC.2006.258407Zhang Y, Shang W, Wu Y (2009) Research on sustainable development based on neural network. In: 2009 Chinese control and decision conference. IEEE, pp 3273–3276. https://doi.org/10.1109/CCDC.2009.5192476Zhou Y, Chang F-J, Chang L-Ch, Kao I-F, Wang YS (2019) Explore a deep learning multi-output neural network for regional multi-step-ahead air quality forecasts. J Clean Prod 209:134–145. https://doi.org/10.1016/j.jclepro.2018.10.24

Crossref

RiuNet