Search CORE

6 research outputs found

Evaluation of ILP-based approaches for partitioning into colorful components

Author: Bruckner S.
Hüffner F.
Komusiewicz Ch.
Niedermeier R.
Publication venue: 'Springer Fachmedien Wiesbaden GmbH'
Publication date: 01/01/2013
Field of study

The NP-hard Colorful Components problem is a graph partitioning problem on vertex-colored graphs. We identify a new application of Colorful Components in the correction of Wikipedia interlanguage links, and describe and compare three exact and two heuristic approaches. In particular, we devise two ILP formulations, one based on Hitting Set and one based on Clique Partition. Furthermore, we use the recently proposed implicit hitting set framework [Karp, JCSS 2011; Chandrasekaran et al., SODA 2011] to solve Colorful Components. Finally, we study a move-based and a merge-based heuristic for Colorful Components. We can optimally solve Colorful Components for Wikipedia link correction data; while the Clique Partition-based ILP outperforms the other two exact approaches, the implicit hitting set is a simple and competitive alternative. The merge-based heuristic is very accurate and outperforms the move-based one. The above results for Wikipedia data are confirmed by experiments with synthetic instances

Repository: Freie Universität Berlin (FU), Math Department (fu_mi_publications)

Automatic Taxonomy Construction from Keywords via Scalable Bayesian Rose Trees

Author: Member IEEE Haixun Wang
Member IEEE Yangqiu Song
Senior Member IEEE Shixia Liu
Xueqing Liu
Publication venue
Publication date: 31/03/2020
Field of study

Abstract-In this paper, we study a challenging problem of deriving a taxonomy from a set of keyword phrases. A solution can benefit many real-world applications because i) keywords give users the flexibility and ease to characterize a specific domain; and ii) in many applications, such as online advertisements, the domain of interest is already represented by a set of keywords. However, it is impossible to create a taxonomy out of a keyword set itself. We argue that additional knowledge and context are needed. To this end, we first use a general-purpose knowledgebase and keyword search to supply the required knowledge and context. Then we develop a Bayesian approach to build a hierarchical taxonomy for a given set of keywords. We reduce the complexity of previous hierarchical clustering approaches from O(n 2 log n) to O(n log n) using a nearest-neighbor-based approximation, so that we can derive a domain-specific taxonomy from one million keyword phrases in less than an hour. Finally, we conduct comprehensive large scale experiments to show the effectiveness and efficiency of our approach. A real life example of building an insurance-related Web search query taxonomy illustrates the usefulness of our approach for specific domains

CiteSeerX

Web Scale Taxonomy Cleansing

Author: 황승원
Publication venue: 'VLDB Endowment'
Publication date: 29/08/2011
Field of study

포항공과대학교

Web Scale Taxonomy Cleansing

Author: Haixun Wang
Seung-Won Hwang
Taesung Lee
Zhongyuan Wang
† Postech
Publication venue: 'VLDB Endowment'
Publication date: 01/01/2011
Field of study

11scopu

CiteSeerX

포항공과대학교

Web scale taxonomy cleansing

Author: Bunescu R.
Carlson A.
Chaudhuri S.
Cohen W. W.
Cucerzan S.
Monge A.
Song Y.
Publication venue: 'VLDB Endowment'
Publication date
Field of study

Crossref