2 research outputs found

    Semi-Supervised Multiple Disambiguation

    Full text link
    Determining the true entity behind an ambiguousword is an NP-Hard problem known as Disambiguation. Previoussolutions often disambiguate a single ambiguous mention acrossmultiple documents. They assume each document contains onlya single ambiguous word and a rich set of unambiguous contextwords. However, nowadays we require fast disambiguation ofshort texts (like news feeds, reviews or Tweets) with few contextwords and multiple ambiguous words. In this research we focuson Multiple Disambiguation (MD) in contrast to Single Disambiguation(SD). Our solution is inspired by a recent algorithm developed for SD. The algorithm categorizes documents by first,transferring them into a graph and then, clustering the graphbased on its topological structure. We changed the graph-baseddocument-modeling of the algorithm, to account for MD. Also,we added a new parameter that controls the resolution of theclustering. Then, we used a supervised sampling approach formerging the clusters when appropriate. Our algorithm, comparedwith the original model, achieved 10% higher quality in termsof F1-Score using only 4% sampling from the dataset.QC 20160407</p