Search CORE

10 research outputs found

Modeling Semi-Bounded Support Data using Non-Gaussian Hidden Markov Models with Applications

Author: Nasfi Rim
Publication venue
Publication date: 01/04/2022
Field of study

With the exponential growth of data in all formats, and data categorization rapidly becoming one of the most essential components of data analysis, it is crucial to research and identify hidden patterns in order to extract valuable information that promotes accurate and solid decision making. Because data modeling is the first stage in accomplishing any of these tasks, its accuracy and consistency are critical for later development of a complete data processing framework. Furthermore, an appropriate distribution selection that corresponds to the nature of the data is a particularly interesting subject of research. Hidden Markov Models (HMMs) are some of the most impressively powerful probabilistic models, which have recently made a big resurgence in the machine learning industry, despite having been recognized for decades. Their ever-increasing application in a variety of critical practical settings to model varied and heterogeneous data (image, video, audio, time series, etc.) is the subject of countless extensions. Equally prevalent, finite mixture models are a potent tool for modeling heterogeneous data of various natures. The over-use of Gaussian mixture models for data modeling in the literature is one of the main driving forces for this thesis. This work focuses on modeling positive vectors, which naturally occur in a variety of real-life applications, by proposing novel HMMs extensions using the Inverted Dirichlet, the Generalized Inverted Dirichlet and the BetaLiouville mixture models as emission probabilities. These extensions are motivated by the proven capacity of these mixtures to deal with positive vectors and overcome mixture models’ impotence to account for any ordering or temporal limitations relative to the information. We utilize the aforementioned distributions to derive several theoretical approaches for learning and deploying Hidden Markov Modelsinreal-world settings. Further, we study online learning of parameters and explore the integration of a feature selection methodology. Extensive experimentation on highly challenging applications ranging from image categorization, video categorization, indoor occupancy estimation and Natural Language Processing, reveals scenarios in which such models are appropriate to apply, and proves their effectiveness compared to the extensively used Gaussian-based models

Concordia University Research Repository

Connecting a Digital Europe through Location and Place. Selected best short papers and posters of the AGILE 2014 Conference, 03-06 June 2014, Castellón, Spain

Author: Granell Carlos
Huerta Joaquin
Schade Sven
Publication venue: AGILE Digital Editions
Publication date: 01/01/2014
Field of study

LAReferencia - Red Federada de Repositorios Institucionales de Publicaciones Científicas Latinoamericanas

Repositori Institucional de la Universitat Jaume I

SIS 2017. Statistics and Data Science: new challenges, new generations

Author
Publication venue: 'Firenze University Press'
Publication date: 31/05/2022
Field of study

The 2017 SIS Conference aims to highlight the crucial role of the Statistics in Data Science. In this new domain of ‘meaning’ extracted from the data, the increasing amount of produced and available data in databases, nowadays, has brought new challenges. That involves different fields of statistics, machine learning, information and computer science, optimization, pattern recognition. These afford together a considerable contribute in the analysis of ‘Big data’, open data, relational and complex data, structured and no-structured. The interest is to collect the contributes which provide from the different domains of Statistics, in the high dimensional data quality validation, sampling extraction, dimensional reduction, pattern selection, data modelling, testing hypotheses and confirming conclusions drawn from the data

Directory of Open Access Books (DOAB)

Livro de Atas do III Encontro Luso-Galaico de Biometria

Author: Costa Marco
Freitas Adelaide
Monteiro Magda
Teixeira Laetitia
Publication venue: Sociedade Portuguesa de Estatística
Publication date: 01/01/2018
Field of study

A cidade de Aveiro foi escolhida, pelas Sociedade Portuguesa de Estatística (SPE) e Sociedade Galega para a Promoción da Estatística e Investigación de Operacións (SGAPEIO), para acolher o III Encontro Luso-Galaico de Biometria (EBio2018). Neste terceiro encontro sobre Biometria, realizado no Departamento de Matemática da Universidade de Aveiro, de 28 a 30 de junho de 2018, reunimos cerca de 100 participantes que desenvolvem e/ou aplicam metodologias estatísticas a dados das ciências da vida e do meio ambiente. Com o objetivo de criar sinergias entre diferentes áreas de aplicação e discutir os mais recentes desenvolvimentos metodológicos em Biometria, este encontro encoraja futuras colaborações e discussão entre os participantes. (...

Repositório Institucional da Universidade de Aveiro

Personalized information retrieval based on time-sensitive user profile

Author: Kacem Sahraoui Ameni
Publication venue
Publication date: 13/06/2017
Field of study

Les moteurs de recherche, largement utilisés dans différents domaines, sont devenus la principale source d'information pour de nombreux utilisateurs. Cependant, les Systèmes de Recherche d'Information (SRI) font face à de nouveaux défis liés à la croissance et à la diversité des données disponibles. Un SRI analyse la requête soumise par l'utilisateur et explore des collections de données de nature non structurée ou semi-structurée (par exemple : texte, image, vidéo, page Web, etc.) afin de fournir des résultats qui correspondent le mieux à son intention et ses intérêts. Afin d'atteindre cet objectif, au lieu de prendre en considération l'appariement requête-document uniquement, les SRI s'intéressent aussi au contexte de l'utilisateur. En effet, le profil utilisateur a été considéré dans la littérature comme l'élément contextuel le plus important permettant d'améliorer la pertinence de la recherche. Il est intégré dans le processus de recherche d'information afin d'améliorer l'expérience utilisateur en recherchant des informations spécifiques. Comme le facteur temps a gagné beaucoup d'importance ces dernières années, la dynamique temporelle est introduite pour étudier l'évolution du profil utilisateur qui consiste principalement à saisir les changements du comportement, des intérêts et des préférences de l'utilisateur en fonction du temps et à actualiser le profil en conséquence. Les travaux antérieurs ont distingué deux types de profils utilisateurs : les profils à court-terme et ceux à long-terme. Le premier type de profil est limité aux intérêts liés aux activités actuelles de l'utilisateur tandis que le second représente les intérêts persistants de l'utilisateur extraits de ses activités antérieures tout en excluant les intérêts récents. Toutefois, pour les utilisateurs qui ne sont pas très actifs dont les activités sont peu nombreuses et séparées dans le temps, le profil à court-terme peut éliminer des résultats pertinents qui sont davantage liés à leurs intérêts personnels. Pour les utilisateurs qui sont très actifs, l'agrégation des activités récentes sans ignorer les intérêts anciens serait très intéressante parce que ce type de profil est généralement en évolution au fil du temps. Contrairement à ces approches, nous proposons, dans cette thèse, un profil utilisateur générique et sensible au temps qui est implicitement construit comme un vecteur de termes pondérés afin de trouver un compromis en unifiant les intérêts récents et anciens. Les informations du profil utilisateur peuvent être extraites à partir de sources multiples. Parmi les méthodes les plus prometteuses, nous proposons d'utiliser, d'une part, l'historique de recherche, et d'autre part les médias sociaux. En effet, les données de l'historique de recherche peuvent être extraites implicitement sans aucun effort de l'utilisateur et comprennent les requêtes émises, les résultats correspondants, les requêtes reformulées et les données de clics qui ont un potentiel de retour de pertinence/rétroaction. Par ailleurs, la popularité des médias sociaux permet d'en faire une source inestimable de données utilisées par les utilisateurs pour exprimer, partager et marquer comme favori le contenu qui les intéresse. En premier lieu, nous avons modélisé le profil utilisateur utilisateur non seulement en fonction du contenu de ses activités mais aussi de leur fraîcheur en supposant que les termes utilisés récemment dans les activités de l'utilisateur contiennent de nouveaux intérêts, préférences et pensées et doivent être pris en considération plus que les anciens intérêts surtout que de nombreux travaux antérieurs ont prouvé que l'intérêt de l'utilisateur diminue avec le temps. Nous avons modélisé le profil utilisateur sensible au temps en fonction d'un ensemble de données collectées de Twitter (un réseau social et un service de microblogging) et nous l'avons intégré dans le processus de reclassement afin de personnaliser les résultats standards en fonction des intérêts de l'utilisateur.En second lieu, nous avons étudié la dynamique temporelle dans le cadre de la session de recherche où les requêtes récentes soumises par l'utilisateur contiennent des informations supplémentaires permettant de mieux expliquer l'intention de l'utilisateur et prouvant qu'il n'a pas trouvé les informations recherchées à partir des requêtes précédentes.Ainsi, nous avons considéré les interactions récentes et récurrentes au sein d'une session de recherche en donnant plus d'importance aux termes apparus dans les requêtes récentes et leurs résultats cliqués. Nos expérimentations sont basés sur la tâche Session TREC 2013 et la collection ClueWeb12 qui ont montré l'efficacité de notre approche par rapport à celles de l'état de l'art. Au terme de ces différentes contributions et expérimentations, nous prouvons que notre modèle générique de profil utilisateur sensible au temps assure une meilleure performance de personnalisation et aide à analyser le comportement des utilisateurs dans les contextes de session de recherche et de médias sociaux.Recently, search engines have become the main source of information for many users and have been widely used in different fields. However, Information Retrieval Systems (IRS) face new challenges due to the growth and diversity of available data. An IRS analyses the query submitted by the user and explores collections of data with unstructured or semi-structured nature (e.g. text, image, video, Web page etc.) in order to deliver items that best match his/her intent and interests. In order to achieve this goal, we have moved from considering the query-document matching to consider the user context. In fact, the user profile has been considered, in the literature, as the most important contextual element which can improve the accuracy of the search. It is integrated in the process of information retrieval in order to improve the user experience while searching for specific information. As time factor has gained increasing importance in recent years, the temporal dynamics are introduced to study the user profile evolution that consists mainly in capturing the changes of the user behavior, interests and preferences, and updating the profile accordingly. Prior work used to discern short-term and long-term profiles. The first profile type is limited to interests related to the user's current activities while the second one represents user's persisting interests extracted from his prior activities excluding the current ones. However, for users who are not very active, the short-term profile can eliminate relevant results which are more related to their personal interests. This is because their activities are few and separated over time. For users who are very active, the aggregation of recent activities without ignoring the old interests would be very interesting because this kind of profile is usually changing over time. Unlike those approaches, we propose, in this thesis, a generic time-sensitive user profile that is implicitly constructed as a vector of weighted terms in order to find a trade-off by unifying both current and recurrent interests. User profile information can be extracted from multiple sources. Among the most promising ones, we propose to use, on the one hand, searching history. Data from searching history can be extracted implicitly without any effort from the user and includes issued queries, their corresponding results, reformulated queries and click-through data that has relevance feedback potential. On the other hand, the popularity of Social Media makes it as an invaluable source of data used by users to express, share and mark as favorite the content that interests them. First, we modeled a user profile not only according to the content of his activities but also to their freshness under the assumption that terms used recently in the user's activities contain new interests, preferences and thoughts and should be considered more than old interests. In fact, many prior works have proved that the user interest is decreasing as time goes by. In order to evaluate the time-sensitive user profile, we used a set of data collected from Twitter, i.e a social networking and microblogging service. Then, we apply our re-ranking process to a Web search system in order to adapt the user's online interests to the original retrieved results. Second, we studied the temporal dynamics within session search where recent submitted queries contain additional information explaining better the user intent and prove that the user hasn't found the information sought from previous submitted ones. We integrated current and recurrent interactions within a unique session model giving more importance to terms appeared in recently submitted queries and clicked results. We conducted experiments using the 2013 TREC Session track and the ClueWeb12 collection that showed the effectiveness of our approach compared to state-of-the-art ones. Overall, in those different contributions and experiments, we prove that our time-sensitive user profile insures better performance of personalization and helps to analyze user behavior in both session search and social media contexts

Thèses en ligne de l'Université Toulouse III - Paul Sabatier

Uncertainty in Artificial Intelligence: Proceedings of the Thirty-Fourth Conference

Author
Publication venue: AUAI Press
Publication date: 01/09/2018
Field of study

UCL Discovery

Acoustic tubes with maximal and minimal resonance frequencies

Author: Bart De Boer
Publication venue: 'Acoustical Society of America (ASA)'
Publication date: 01/01/2008
Field of study

Crossref

International Migration, Integration and Social Cohesion online publications

Statistics in the 150 years from Italian Unification. SIS 2011 Statistical Conference, Bologna, 8 – 10 June 2011. Book of short paper.

Author
Publication venue: Dipartimento di Scienze Statistiche "Paolo Fortunati", Alma Mater Studiorum Università di Bologna
Publication date: 20/12/2011
Field of study

Statistics in the 150 years from Italian Unification. SIS 2011 Statistical Conference, Bologna, 8 – 10 June 2011. Book of short paper.

Author
Publication venue: Dipartimento di Scienze Statistiche "Paolo Fortunati", Alma Mater Studiorum Università di Bologna
Publication date: 20/12/2011
Field of study

AMS Acta

New Technologies and Statistics: Partners for Environmental Monitoring and City Sensing

Author: ER Tufte
ER Tufte
F Calabrese
F Calabrese
M Batty
MF Goodchild
N Maisonneuve
PA Burrough
PA Longley
Z Jirícek
Publication venue: 'Springer Science and Business Media LLC'
Publication date
Field of study

Crossref