Automatic capitalisation generation for speech input

14 years 2 months ago

Download mi.eng.cam.ac.uk

Two different systems are proposed for the task of capitalisation generation. The first system is a slightly modified speech recogniser. In this system, every word in the vocabulary is duplicated: once in a decapitalised form and again in capitalised forms. In addition, the language model is re-trained on mixed case texts. The other system is based on Named Entity (NE) recognition and punctuation generation, since most capitalised words are the first words in sentences or NE words. Both systems are compared when every procedure is fully automated. The system based on NE recognition and punctuation generation shows better results by word error rate, by F-measure and by slot error rate than the system modified from the speech recogniser. This is because the latter system has a distorted language model and a sparser language model. The detailed performance of the system based on NE recognition and punctuation generation is investigated by including one or more of the following: the refer...

Ji-Hwan Kim, Philip C. Woodland

Real-time Traffic

Automated Reasoning | Capitalisation Generation | CSL 2004 | NE Recognition | Punctuation Generation |

claim paper

Post Info
More Details (n/a)

Added	17 Dec 2010
Updated	17 Dec 2010
Type	Journal
Year	2004
Where	CSL
Authors	Ji-Hwan Kim, Philip C. Woodland

Comments (0)

Sciweavers

Automatic capitalisation generation for speech input

Automated Reasoning | Capitalisation Generation | CSL 2004 | NE Recognition | Punctuation Generation |

Explore & Download

Productivity Tools

Document Tools

Image Tools

Sciweavers