Search CORE

1 research outputs found

Feature extraction and selection for Arabic tweets authorship authentication

Author: A Abbasi
A Abbasi
A Pasha
Abdullateef Rabab’ah
AS Altheneyan
CC Aggarwal
E Stamatatos
E Stamatatos
F Mosteller
G Hirst
G Kanaan
GH Dunteman
H Sayoud
HC Chen
I Kononenko
JT Kent
M Hall
M Koppel
Mahmoud Al-Ayyoub
ML Brocardo
Monther Aldwairi
MS Khorsheed
N Cheng
O Vel De
P Juola
P Juola
P Kosmides
RS Baraka
T Helmy
W Deitrick
Yaser Jararweh
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/06/2017
Field of study

© 2017, Springer-Verlag Berlin Heidelberg. In tweet authentication, we are concerned with correctly attributing a tweet to its true author based on its textual content. The more general problem of authenticating long documents has been studied before and the most common approach relies on the intuitive idea that each author has a unique style that can be captured using stylometric features (SF). Inspired by the success of modern automatic document classification problem, some researchers followed the Bag-Of-Words (BOW) approach for authenticating long documents. In this work, we consider both approaches and their application on authenticating tweets, which represent additional challenges due to the limitation in their sizes. We focus on the Arabic language due to its importance and the scarcity of works related on it. We create different sets of features from both approaches and compare the performance of different classifiers using them. We experiment with various feature selection techniques in order to extract the most discriminating features. To the best of our knowledge, this is the first study of its kind to combine these different sets of features for authorship analysis of Arabic tweets. The results show that combining all the feature sets we compute yields the best results

ZU Scholars (Zayed University)

Crossref