Sciweavers

PODS
2002
ACM

Monadic Datalog and the Expressive Power of Languages for Web Information Extraction

14 years 11 months ago
Monadic Datalog and the Expressive Power of Languages for Web Information Extraction
Research on information extraction from Web pages (wrapping) has seen much activity in recent times (particularly systems implementations), but little work has been done on formally studying the expressiveness of the formalisms proposed or on the theoretical foundations of wrapping. In this paper, we first study monadic datalog as a wrapping language (over ranked or unranked tree structures). Using previous work by Neven and Schwentick, we show that this simple language is equivalent to full monadic second order logic (MSO) in its ability to specify wrappers. We believe that MSO has the right expressiveness required for Web information extraction and thus propose MSO as a yardstick for evaluating and comparing wrappers. Using the above result, we study the kernel fragment Elogof the Elog wrapping language used in the Lixto system (a visual wrapper generator). The striking fact here is that Elogexactly captures MSO, yet is easier to use. Indeed, programs in this language can be entirel...
Georg Gottlob, Christoph Koch
Added 08 Dec 2009
Updated 08 Dec 2009
Type Conference
Year 2002
Where PODS
Authors Georg Gottlob, Christoph Koch
Comments (0)