Sciweavers

CLEF
2010
Springer

Creating a Persian-English Comparable Corpus

13 years 11 months ago
Creating a Persian-English Comparable Corpus
Multilingual corpora are valuable resources for cross-language information retrieval and are available in many language pairs. However the Persian language does not have rich multilingual resources due to some of its special features and difficulties in constructing the corpora. In this study, we build a Persian-English comparable corpus from two independent news collections: BBC News in English and Hamshahri news in Persian. We use the similarity of the document topics and their publication dates to align the documents in these sets. We tried several alternatives for constructing the comparable corpora and assessed the quality of the corpora using different criteria. Evaluation results show the high quality of the aligned documents and using the Persian-English comparable corpus for extracting translation knowledge seems promising.
Homa Baradaran Hashemi, Azadeh Shakery, Heshaam Fe
Added 06 Dec 2010
Updated 06 Dec 2010
Type Conference
Year 2010
Where CLEF
Authors Homa Baradaran Hashemi, Azadeh Shakery, Heshaam Feili
Comments (0)