Search CORE

1,698 research outputs found

A survey on sentiment analysis in Urdu: A resource-poor language

Author: Ahmad Shakeel
Asghar Muhammad Zubair
Asif Hassan Syed
Hameed Ibrahim A.
Khattak Asad
Saeed Anam
Publication venue: ZU Scholars
Publication date: 01/01/2020
Field of study

© 2020 Background/introduction: The dawn of the internet opened the doors to the easy and widespread sharing of information on subject matters such as products, services, events and political opinions. While the volume of studies conducted on sentiment analysis is rapidly expanding, these studies mostly address English language concerns. The primary goal of this study is to present state-of-art survey for identifying the progress and shortcomings saddling Urdu sentiment analysis and propose rectifications. Methods: We described the advancements made thus far in this area by categorising the studies along three dimensions, namely: text pre-processing lexical resources and sentiment classification. These pre-processing operations include word segmentation, text cleaning, spell checking and part-of-speech tagging. An evaluation of sophisticated lexical resources including corpuses and lexicons was carried out, and investigations were conducted on sentiment analysis constructs such as opinion words, modifiers, negations. Results and conclusions: Performance is reported for each of the reviewed study. Based on experimental results and proposals forwarded through this paper provides the groundwork for further studies on Urdu sentiment analysis

ZU Scholars (Zayed University)

A Comprehensive Review of Sentiment Analysis on Indian Regional Languages: Techniques, Challenges, and Trends

Author: Kale Sunil D.
Mahalle Parikshit N.
Mane Deepak T.
Potdar Girish P.
Prasad Rajesh
Upadhye Gopal D.
Publication venue: Auricle Global Society of Education and Research
Publication date: 31/08/2023
Field of study

Sentiment analysis (SA) is the process of understanding emotion within a text. It helps identify the opinion, attitude, and tone of a text categorizing it into positive, negative, or neutral. SA is frequently used today as more and more people get a chance to put out their thoughts due to the advent of social media. Sentiment analysis benefits industries around the globe, like finance, advertising, marketing, travel, hospitality, etc. Although the majority of work done in this field is on global languages like English, in recent years, the importance of SA in local languages has also been widely recognized. This has led to considerable research in the analysis of Indian regional languages. This paper comprehensively reviews SA in the following major Indian Regional languages: Marathi, Hindi, Tamil, Telugu, Malayalam, Bengali, Gujarati, and Urdu. Furthermore, this paper presents techniques, challenges, findings, recent research trends, and future scope for enhancing results accuracy

International Journal on Recent and Innovation Trends in Computing and Communication

Urdu Speech and Text Based Sentiment Analyzer

Author: Ahmad Waqar
Edalati Maryam
Publication venue
Publication date: 19/07/2022
Field of study

Discovering what other people think has always been a key aspect of our information-gathering strategy. People can now actively utilize information technology to seek out and comprehend the ideas of others, thanks to the increased availability and popularity of opinion-rich resources such as online review sites and personal blogs. Because of its crucial function in understanding people's opinions, sentiment analysis (SA) is a crucial task. Existing research, on the other hand, is primarily focused on the English language, with just a small amount of study devoted to low-resource languages. For sentiment analysis, this work presented a new multi-class Urdu dataset based on user evaluations. The tweeter website was used to get Urdu dataset. Our proposed dataset includes 10,000 reviews that have been carefully classified into two categories by human experts: positive, negative. The primary purpose of this research is to construct a manually annotated dataset for Urdu sentiment analysis and to establish the baseline result. Five different lexicon- and rule-based algorithms including Naivebayes, Stanza, Textblob, Vader, and Flair are employed and the experimental results show that Flair with an accuracy of 70% outperforms other tested algorithms.Comment: Sentiment Analysis, Opinion Mining, Urdu language, polarity assessment, lexicon-based metho

arXiv.org e-Print Archive

Sentiment Classification of Customer Reviews about Automobiles in Roman Urdu

Author: A Abbasi
A Jebaseel
A Rashid
AZ Syed
B Pang
K Katsiavriades
TN Khushboo
WJ Wilbur
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 30/12/2018
Field of study

Text mining is a broad field having sentiment mining as its important constituent in which we try to deduce the behavior of people towards a specific item, merchandise, politics, sports, social media comments, review sites etc. Out of many issues in sentiment mining, analysis and classification, one major issue is that the reviews and comments can be in different languages like English, Arabic, Urdu etc. Handling each language according to its rules is a difficult task. A lot of research work has been done in English Language for sentiment analysis and classification but limited sentiment analysis work is being carried out on other regional languages like Arabic, Urdu and Hindi. In this paper, Waikato Environment for Knowledge Analysis (WEKA) is used as a platform to execute different classification models for text classification of Roman Urdu text. Reviews dataset has been scrapped from different automobiles sites. These extracted Roman Urdu reviews, containing 1000 positive and 1000 negative reviews, are then saved in WEKA attribute-relation file format (arff) as labeled examples. Training is done on 80% of this data and rest of it is used for testing purpose which is done using different models and results are analyzed in each case. The results show that Multinomial Naive Bayes outperformed Bagging, Deep Neural Network, Decision Tree, Random Forest, AdaBoost, k-NN and SVM Classifiers in terms of more accuracy, precision, recall and F-measure.Comment: This is a pre-print of a contribution published in Advances in Intelligent Systems and Computing (editors: Kohei Arai, Supriya Kapoor and Rahul Bhatia) published by Springer, Cham. The final authenticated version is available online at: https://doi.org/10.1007/978-3-030-03405-4_4

arXiv.org e-Print Archive

Crossref

Urdu Poetry Generated by Using Deep Learning Techniques

Author: Abbas Ali
Farooq Muhammad Shoaib
Publication venue
Publication date: 25/09/2023
Field of study

This study provides Urdu poetry generated using different deep-learning techniques and algorithms. The data was collected through the Rekhta website, containing 1341 text files with several couplets. The data on poetry was not from any specific genre or poet. Instead, it was a collection of mixed Urdu poems and Ghazals. Different deep learning techniques, such as the model applied Long Short-term Memory Networks (LSTM) and Gated Recurrent Unit (GRU), have been used. Natural Language Processing (NLP) may be used in machine learning to understand, analyze, and generate a language humans may use and understand. Much work has been done on generating poetry for different languages using different techniques. The collection and use of data were also different for different researchers. The primary purpose of this project is to provide a model that generates Urdu poems by using data completely, not by sampling data. Also, this may generate poems in pure Urdu, not Roman Urdu, as in the base paper. The results have shown good accuracy in the poems generated by the model.Comment: 11 pages, 2 figure

arXiv.org e-Print Archive

Adapting Deep Learning for Sentiment Classification of Code-Switched Informal Short Text

Author: Attia Mohammed
Devlin Jacob
Wang Xingyou
Wang Zhongqing
Publication venue: 'Association for Computing Machinery (ACM)'
Publication date: 04/01/2020
Field of study

Nowadays, an abundance of short text is being generated that uses nonstandard writing styles influenced by regional languages. Such informal and code-switched content are under-resourced in terms of labeled datasets and language models even for popular tasks like sentiment classification. In this work, we (1) present a labeled dataset called MultiSenti for sentiment classification of code-switched informal short text, (2) explore the feasibility of adapting resources from a resource-rich language for an informal one, and (3) propose a deep learning-based model for sentiment classification of code-switched informal short text. We aim to achieve this without any lexical normalization, language translation, or code-switching indication. The performance of the proposed models is compared with three existing multilingual sentiment classification models. The results show that the proposed model performs better in general and adapting character-based embeddings yield equivalent performance while being computationally more efficient than training word-based domain-specific embeddings

arXiv.org e-Print Archive

Crossref

Urdu News Content Classification Using Machine Learning Algorithms

Author: Khawar Iqbal Malik
Publication venue: Lahore Garrison University
Publication date: 30/03/2022
Field of study

As the world has become a global village, the flow of news in terms of volume and speed increases. It is necessary to engage computing machines for assisting people in dealing with this massive data. The availability of different types of news and such material on the Internet serves as a source of information for billions of users. Millions of people in our subcontinent speak and understand Urdu. There are several classification techniques that are available and are applied to classify English news like political, Education, Medical, etc. Plenty of research work has been done in multiple languages but Urdu is still to be worked on due to a lack of resources. This research evaluates the performance of twelve (12) different Machine learning classifiers for the Urdu News text Classification problem. The analysis was performed on a relatively big and recent collection of Urdu text that contains over 0.15 million (153,050) labeled instances of eight different classes. In addition, after applying pre-processing techniques, the TF-IDF weighting technique was adopted for feature selection and data extraction. After evaluating various machine learning methods, the SVM outperforms the other eleven algorithms with an accuracy of 91.37 %. We also compare its results with other classifiers like linear SVM, Logistic regression, SGD, Naïve bays, ridge regression, and a few others

Lahore Garrison University Research Journal of Computer Science and Information Technology