7,712 research outputs found
No-But-Semantic-Match: Computing Semantically Matched XML Keyword Search Results
Users are rarely familiar with the content of a data source they are
querying, and therefore cannot avoid using keywords that do not exist in the
data source. Traditional systems may respond with an empty result, causing
dissatisfaction, while the data source in effect holds semantically related
content. In this paper we study this no-but-semantic-match problem on XML
keyword search and propose a solution which enables us to present the top-k
semantically related results to the user. Our solution involves two steps: (a)
extracting semantically related candidate queries from the original query and
(b) processing candidate queries and retrieving the top-k semantically related
results. Candidate queries are generated by replacement of non-mapped keywords
with candidate keywords obtained from an ontological knowledge base. Candidate
results are scored using their cohesiveness and their similarity to the
original query. Since the number of queries to process can be large, with each
result having to be analyzed, we propose pruning techniques to retrieve the
top- results efficiently. We develop two query processing algorithms based
on our pruning techniques. Further, we exploit a property of the candidate
queries to propose a technique for processing multiple queries in batch, which
improves the performance substantially. Extensive experiments on two real
datasets verify the effectiveness and efficiency of the proposed approaches.Comment: 24 pages, 21 figures, 6 tables, submitted to The VLDB Journal for
possible publicatio
Exploring Communities in Large Profiled Graphs
Given a graph and a vertex , the community search (CS) problem
aims to efficiently find a subgraph of whose vertices are closely related
to . Communities are prevalent in social and biological networks, and can be
used in product advertisement and social event recommendation. In this paper,
we study profiled community search (PCS), where CS is performed on a profiled
graph. This is a graph in which each vertex has labels arranged in a
hierarchical manner. Extensive experiments show that PCS can identify
communities with themes that are common to their vertices, and is more
effective than existing CS approaches. As a naive solution for PCS is highly
expensive, we have also developed a tree index, which facilitate efficient and
online solutions for PCS
Graph-based task libraries for robots: generalization and autocompletion
In this paper, we consider an autonomous robot that persists
over time performing tasks and the problem of providing one additional
task to the robot's task library. We present an approach to generalize
tasks, represented as parameterized graphs with sequences, conditionals,
and looping constructs of sensing and actuation primitives. Our approach
performs graph-structure task generalization, while maintaining task ex-
ecutability and parameter value distributions. We present an algorithm
that, given the initial steps of a new task, proposes an autocompletion
based on a recognized past similar task. Our generalization and auto-
completion contributions are eective on dierent real robots. We show
concrete examples of the robot primitives and task graphs, as well as
results, with Baxter. In experiments with multiple tasks, we show a sig-
nicant reduction in the number of new task steps to be provided
A Fast Quartet Tree Heuristic for Hierarchical Clustering
The Minimum Quartet Tree Cost problem is to construct an optimal weight tree
from the weighted quartet topologies on objects, where
optimality means that the summed weight of the embedded quartet topologies is
optimal (so it can be the case that the optimal tree embeds all quartets as
nonoptimal topologies). We present a Monte Carlo heuristic, based on randomized
hill climbing, for approximating the optimal weight tree, given the quartet
topology weights. The method repeatedly transforms a dendrogram, with all
objects involved as leaves, achieving a monotonic approximation to the exact
single globally optimal tree. The problem and the solution heuristic has been
extensively used for general hierarchical clustering of nontree-like
(non-phylogeny) data in various domains and across domains with heterogeneous
data. We also present a greatly improved heuristic, reducing the running time
by a factor of order a thousand to ten thousand. All this is implemented and
available, as part of the CompLearn package. We compare performance and running
time of the original and improved versions with those of UPGMA, BioNJ, and NJ,
as implemented in the SplitsTree package on genomic data for which the latter
are optimized.
Keywords: Data and knowledge visualization, Pattern
matching--Clustering--Algorithms/Similarity measures, Hierarchical clustering,
Global optimization, Quartet tree, Randomized hill-climbing,Comment: LaTeX, 40 pages, 11 figures; this paper has substantial overlap with
arXiv:cs/0606048 in cs.D
- …