Sciweavers

BIS
2010

Textractor: A Framework for Extracting Relevant Domain Concepts from Irregular Corporate Textual Datasets

14 years 18 days ago
Textractor: A Framework for Extracting Relevant Domain Concepts from Irregular Corporate Textual Datasets
Various information extraction (IE) systems for corporate usage exist. However, none of them target the product development and/or customer service domain, despite significant application potentials and benefits. This domain also poses new scientific challenges, such as the lack of external knowledge resources, and irregularities like ungrammatical constructs in textual data, which compromise successful information extraction. To address these issues, we describe the development of Textractor; an application for accurately extracting relevant concepts from irregular textual narratives in datasets of product development and/or customer service organizations. The extracted information can subsequently be fed to a host of business intelligence activities. We present novel algorithms, combining both statistical and linguistic approaches, for the accurate discovery of relevant domain concepts from highly irregular/ungrammatical texts. Evaluations on real-life corporate data revealed that Te...
Ashwin Ittoo, Laura Maruster, Hans Wortmann, Gosse
Added 29 Oct 2010
Updated 29 Oct 2010
Type Conference
Year 2010
Where BIS
Authors Ashwin Ittoo, Laura Maruster, Hans Wortmann, Gosse Bouma
Comments (0)