Search CORE

1 research outputs found

Text classification using tree kernels and linguistic information

Author: Gonçalves Teresa
Quaresma Paulo
Publication venue: IEEE Computer Society
Publication date: 01/12/2008
Field of study

Standard Machine Learning approaches to text classification use the bag-of-words representation of documents to deceive the classification target function. Typical linguistic structures such as morphology, syntax and semantic are completely ignored in the learning process. This paper examines the role of these structures on the classifier construction applying the study to the Portuguese language. Classifiers are built using the SVM algorithm on a newspaper's articles dataset. The results show that syntactic structure is not useful for text classification (as initially expected), but a novel structured representation that uses document's semantic information has the same discriminative power over classes as the traditional bag-of-words one

Repositório Científico da Universidade de Évora