Sciweavers

ICDE
2005
IEEE

ProtChew: Automatic Extraction of Protein Names from Biomedical Literature

14 years 4 months ago
ProtChew: Automatic Extraction of Protein Names from Biomedical Literature
With the increasing amount of biomedical literature, there is a need for automatic extraction of information to support biomedical researchers. Due to incomplete biomedical information databases, the extraction is not straightforward using dictionaries, and several approaches using contextual rules and machine learning have previously been proposed. Our work is inspired by the previous approaches, but is novel in the sense that it is fully automatic and doesn’t rely on expert tagged corpora. The main ideas are 1) unigram tagging of corpora using known protein names for training examples for the protein name extraction classifier and 2) tight positive and negative examples by having protein-related words as negative examples and protein names/synonyms as positive examples. We present preliminary results on Medline abstracts about gastrin, further work will be on testing the approach on BioCreative benchmark data sets.
Amund Tveit, Rune Sætre, Astrid Lægrei
Added 24 Jun 2010
Updated 24 Jun 2010
Type Conference
Year 2005
Where ICDE
Authors Amund Tveit, Rune Sætre, Astrid Lægreid, Tonje Strommen Steigedal
Comments (0)