Search CORE

6 research outputs found

A New Feature Selection Method based on Intuitionistic Fuzzy Entropy to Categorize Text Documents

Author: Harish B S
Revanasiddappa M B
Publication venue: 'Universidad Internacional de La Rioja'
Publication date: 04/02/2022
Field of study

Selection of highly discriminative feature in text document plays a major challenging role in categorization. Feature selection is an important task that involves dimensionality reduction of feature matrix, which in turn enhances the performance of categorization. This article presents a new feature selection method based on Intuitionistic Fuzzy Entropy (IFE) for Text Categorization. Firstly, Intuitionistic Fuzzy C-Means (IFCM) clustering method is employed to compute the intuitionistic membership values. The computed intuitionistic membership values are used to estimate intuitionistic fuzzy entropy via Match degree. Further, features with lower entropy values are selected to categorize the text documents. To find the efficacy of the proposed method, experiments are conducted on three standard benchmark datasets using three classifiers. F-measure is used to assess the performance of the classifiers. The proposed method shows impressive results as compared to other well known feature selection methods. Moreover, Intuitionistic Fuzzy Set (IFS) property addresses the uncertainty limitations of traditional fuzzy set

Re-UNIR

Maximum Entropy Modeling with Feature Selection for Text Categorization

Author: Fei Song
Jihong Cai
Publication venue
Publication date: 01/01/2008
Field of study

Abstract. Maximum entropy provides a reasonable way of estimating probability distributions and has been widely used for a number of language processing tasks. In this paper, we explore the use of different feature selection methods for text categorization using maximum entropy modeling. We also propose a new feature selection method based on the difference between the relative document frequencies of a feature for both relevant and irrelevant classes. Our experiments on the Reuters RCV1 data set show that our own feature selection performs better than the other feature selection methods and maximum entropy modeling is a competitive method for text categorization

CiteSeerX

Maximum Entropy Modeling with Feature Selection for Text Categorization

Author: C. Manning
D..D. Lewis
G. Salton
T. Mitchell
Y. Yang
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/01/2008
Field of study

Crossref

Statystyczne metody klasyfikacji tekstów

Author: Idczak Adam
Korzeniewski Jerzy
Publication venue: 'Uniwersytet Lodzki (University of Lodz)'
Publication date: 01/01/2022
Field of study

W ostatnich latach, wraz z szybkim rozwojem technologii komputerowych i internetowych, coraz większego znaczenia nabierają komputerowe metody badania tekstu, w szczególności metody ustalania sentymentu czy też wydźwięku tekstu. Metody komputerowe mogą być później wykorzystywane w takich zagadnieniach, jak streszczanie tekstu, wyszukiwanie informacji z tekstu, sprawdzanie poprawności tekstu, maszynowe tłumaczenie tekstu i wielu innych. Niniejsza monografia zawiera przegląd metod analizy sentymentu dla dokumentów głównie anglojęzycznych, badanie efektywności wybranych metod analizy sentymentu w zastosowaniu do dokumentów polskojęzycznych, propozycje nowych metod, które mogą poprawić jakość klasyfikacji. W nowych propozycjach nacisk został położony na problemy klasyfikacji binarnej, niekorzystanie ze źródeł zewnętrznych, korzystanie w jak najmniejszym stopniu ze zbioru uczącego. Proponujemy przenieść ciężar klasyfikacji tekstów z obszernego zbioru uczącego na wyszukiwanie i analizowanie związków pomiędzy słowami tworzącymi dokument, a nawet grupami słów. Zaproponowana metoda ma prostą interpretację, może konkurować z metodami standardowymi oraz może być wykorzystana do innych problemów związanych z ustalaniem sentymentu tekstów

Repozytorium Uniwersytetu Łódzkiego (University of Lodz Repository)

Characterisation of business documents: an approach to the automation of quality assessment

Author: Thurlow Ian
Publication venue: UCL (University College London)
Publication date: 28/06/2018
Field of study

This thesis explores a new approach to automatic characterisation of business documents of different levels of document effectiveness. Supervised text categorisation techniques are used to derive text features that characterise a specific type of business document in accordance with pre-assigned levels of document utility. The documents in question are the executive summary sections of a representative sample of sales proposal documents. The executive summaries are first rated by domain experts against a quality framework comprising pre-selected dimensions of document quality. An automatic analysis of the texts shows that certain words, word sequences, and patterns of words have the capacity to discriminate between executive summaries of varying levels of document effectiveness. Function words, which are frequently ignored in many text classification tasks, are retained and are shown to provide an important element of the word patterns. Automatic text classifiers that utilise these features are shown to categorise previously unseen executive summaries at an acceptable level of classification performance. The outcomes of the research are applied to the development of a new computer application. The application identifies, in the text of a new executive summary, word patterns that discriminate between sets of summaries previously categorised into different levels of document utility. The action of highlighting the respective categories of discriminating word patterns directs authors to areas of text that may need further attention. A trial of a prototype of the application suggests that it provides an effective way to help sales professionals improve the content and quality of the text of this type of business document. Moreover, as the approach is suitably generic, it could be applied to different types of document in different domains

UCL Discovery