ProtChew: Automatic Extraction of Protein Names from Biomedical Literature

16 years 1 months ago

Download www-tsujii.is.s.u-tokyo.ac.jp

With the increasing amount of biomedical literature, there is a need for automatic extraction of information to support biomedical researchers. Due to incomplete biomedical information databases, the extraction is not straightforward using dictionaries, and several approaches using contextual rules and machine learning have previously been proposed. Our work is inspired by the previous approaches, but is novel in the sense that it is fully automatic and doesn’t rely on expert tagged corpora. The main ideas are 1) unigram tagging of corpora using known protein names for training examples for the protein name extraction classiﬁer and 2) tight positive and negative examples by having protein-related words as negative examples and protein names/synonyms as positive examples. We present preliminary results on Medline abstracts about gastrin, further work will be on testing the approach on BioCreative benchmark data sets.

Amund Tveit, Rune Sætre, Astrid Lægrei

Real-time Traffic

Biomedical | Database | ICDE 2005 | Incomplete Biomedical Information | Negative Examples |

claim paper

Post Info
More Details (n/a)

Added	24 Jun 2010
Updated	24 Jun 2010
Type	Conference
Year	2005
Where	ICDE
Authors	Amund Tveit, Rune Sætre, Astrid Lægreid, Tonje Strommen Steigedal

Comments (0)

Sciweavers

ProtChew: Automatic Extraction of Protein Names from Biomedical Literature

Biomedical | Database | ICDE 2005 | Incomplete Biomedical Information | Negative Examples |

Explore & Download

Productivity Tools

Sciweavers