The Czech Broadcast Conversation Corpus

16 years 1 months ago

Download www.mde.zcu.cz

Abstract. This paper presents the ﬁnal version of the Czech Broadcast Conversation Corpus released at the Linguistic Data Consortium (LDC). The corpus contains 72 recordings of a radio discussion program, which yield about 33 hours of transcribed conversational speech from 128 speakers. The release not only includes verbatim transcripts and speaker information, but also structural metadata (MDE) annotation that involves labeling of sentence-like unit boundaries, marking of non-content words like ﬁlled pauses and discourse markers, and annotation of speech disﬂuencies. The annotation is based on the LDC’s MDE annotation standard for English, with changes applied to accommodate phenomena that are speciﬁc for Czech. In addition to its importance to speech recognition, speaker diarization, and structural metadata extraction research, the corpus is also useful for linguistic analysis of conversational Czech.

Jáchym Kolár, Jan Svec

Real-time Traffic