Babylon Parallel Text Builder: Gathering Parallel Texts for Low-Density Languages

15 years 8 months ago

Download www.lrec-conf.org

This paper describes BABYLON, a system that attempts to overcome the shortage of parallel texts in low-density languages by supplementing existing parallel texts with texts gathered automatically from the Web. In addition to the identification of entire Web pages, we also propose a new feature specifically designed to find parallel text chunks within a single document. Experiments carried out on the Quechua-Spanish language pair show that the system is successful in automatically identifying a significant amount of parallel texts on the Web. Evaluations of a machine translation system trained on this corpus indicate that the Web-gathered parallel texts can supplement manually compiled parallel texts and perform significantly better than the manually compiled texts when tested on other Web-gathered data.

Michael Mohler, Rada Mihalcea

Real-time Traffic

Education | LREC 2008 | Parallel Text Chunks | Parallel Texts | Web-gathered Parallel Texts |

claim paper

Added	29 Oct 2010
Updated	29 Oct 2010
Type	Conference
Year	2008
Where	LREC
Authors	Michael Mohler, Rada Mihalcea

Sciweavers

Babylon Parallel Text Builder: Gathering Parallel Texts for Low-Density Languages

Education | LREC 2008 | Parallel Text Chunks | Parallel Texts | Web-gathered Parallel Texts |

Explore & Download

Productivity Tools

Sciweavers