Information extraction challenges in managing unstructured data

16 years 7 months ago

Download pages.cs.wisc.edu

Over the past few years, we have been trying to build an end-to-end system at Wisconsin to manage unstructured data, using extraction, integration, and user interaction. This paper describes the key information extraction (IE) challenges that we have run into, and sketches our solutions. We discuss in particular developing a declarative IE language, optimizing for this language, generating IE provenance, incorporating user feedback into the IE process, developing a novel wikibased user interface for feedback, best-effort IE, pushing IE into RDBMSs, and more. Our work suggests that IE in managing unstructured data can open up many interesting research challenges, and that these challenges can greatly benefit from the wealth of work on managing structured data that has been carried out by the database community.

AnHai Doan, Jeffrey F. Naughton, Raghu Ramakrishna

Real-time Traffic

Database | Declarative Ie Language | IE Process | IE Provenance | SIGMOD 2008 |

claim paper

» Mapping enterprise entities to text segments

» Models and Indices for Integrating Unstructured Data with a Relational Database

» SQL Queries Over Unstructured Text Databases

» Predicting accuracy of extracting information from unstructured text collections

» Unsupervised information extraction from unstructured ungrammatical data sources on the Wo...

» A Mutually Beneficial Integration of Data Mining and Information Extraction

» Evaluating the Effectiveness of Information Extraction in RealWorld Storage Management

» Extracting unstructured data from template generated web documents

Post Info
More Details (n/a)

Added	08 Dec 2009
Updated	08 Dec 2009
Type	Conference
Year	2008
Where	SIGMOD
Authors	AnHai Doan, Jeffrey F. Naughton, Raghu Ramakrishnan, Akanksha Baid, Xiaoyong Chai, Fei Chen 0002, Ting Chen, Eric Chu, Pedro DeRose, Byron J. Gao, Chaitanya Gokhale, Jiansheng Huang, Warren Shen, Ba-Quy Vuong

Comments (0)

Sciweavers

Information extraction challenges in managing unstructured data

Database | Declarative Ie Language | IE Process | IE Provenance | SIGMOD 2008 |

Explore & Download

Productivity Tools

Sciweavers