Sciweavers

Free Online Productivity Tools i2Speak i2Symbol i2OCR iTex2Img iWeb2Print iWeb2Shot i2Type iPdf2Split iPdf2Merge i2Bopomofo i2Arabic i2Style i2Image i2PDF iLatex2Rtf Sci2ools

199

PAKDD
2001
ACM

157views Data Mining» more PAKDD 2001»

Applying Pattern Mining to Web Information Extraction

15 years 11 months ago

Applying Pattern Mining to Web Information Extraction

Download winslab.cnu.ac.kr

Information extraction (IE) from semi-structured Web documents is a critical issue for information integration systems on the Internet. Previous work in wrapper induction aim to solve this problem by applying machine learning to automatically generate extractors. For example, WIEN, Stalker, Softmealy, etc. However, this approach still requires human intervention to provide training examples. In this paper, we propose a novel idea to IE, by repeated pattern mining and multiple pattern alignment. The discovery of repeated patterns are realized through a data structure call PAT tree. In addition, incomplete patterns are further revised by pattern alignment to comprehend all pattern instances. This new track to IE involves no human eﬀort and content-dependent heuristics. Experimental results show that the constructed extraction rules can achieves 97 percent extraction over fourteen popular search engines.

Chia-Hui Chang, Shao-Chen Lui, Yen-Chin Wu

Real-time Traffic

Data Mining | Multiple Pattern Alignment | PAKDD 2001 | Pattern Alignment | Wrapper Induction Aim |

claim paper

Related Content

» WebSets extracting sets of entities from the web using unsupervised information extraction

» Using the Web to Reduce Data Sparseness in PatternBased Information Extraction

» Web Usage Mining users navigational patterns extraction from web logs using antbased clust...

» Unsupervised Relation Extraction by Mining Wikipedia Texts Using Information from the Web

» LOGML Log Markup Language for Web Usage Mining

» Prioritization of DomainSpecific Web Information Extraction

» Extracting Author MetaData from Web Using Visual Features

» IEPAD information extraction based on pattern discovery

» Discovering User Access Pattern Based on Probabilistic Latent Factor Model

Post Info
More Details (n/a)

Added	30 Jul 2010
Updated	30 Jul 2010
Type	Conference
Year	2001
Where	PAKDD
Authors	Chia-Hui Chang, Shao-Chen Lui, Yen-Chin Wu

Comments (0)