Article

Probe, Cluster, and Discover: Focused Extraction of QA-Pagelets from the Deep Web

03/2004;
Source: CiteSeer

ABSTRACT In this paper, we introduce the concept of a QA-Pagelet to refer to the content region in a dynamic page that contains query matches. We present THOR, a scalable and efficient mining system for discovering and extracting QAPagelets from the Deep Web. A unique feature of THOR is its two-phase extraction framework. In the first phase, pages from a deep web site are grouped into distinct clusters of structurally-similar pages. In the second phase, pages from each page cluster are examined through a subtree filtering algorithm that exploits the structural and content similarity at subtree level to identify the QA-Pagelets.

0 0
 · 
0 Bookmarks
 · 
26 Views

Full-text (2 Sources)

View
1 Download
Available from
17 May 2013

Keywords

algorithm
 
contains query matches
 
discovering
 
distinct clusters
 
efficient mining system
 
extracting QAPagelets
 
first phase
 
QA-Pagelet
 
QA-Pagelets
 
scalable
 
two-phase extraction framework
 
unique feature
 
Web
 
web site