Sciweavers

Free Online Productivity Tools i2Speak i2Symbol i2OCR iTex2Img iWeb2Print iWeb2Shot i2Type iPdf2Split iPdf2Merge i2Bopomofo i2Arabic i2Style i2Image i2PDF iLatex2Rtf Sci2ools

116

ACL
2008

favoriteEmaildiscussreport

133views Computational Linguistics» more ACL 2008»

Assessing Dialog System User Simulation Evaluation Measures Using Human Judges

15 years 3 months ago

Assessing Dialog System User Simulation Evaluation Measures Using Human Judges

Download www.aclweb.org

Previous studies evaluate simulated dialog corpora using evaluation measures which can be automatically extracted from the dialog systems' logs. However, the validity of these automatic measures has not been fully proven. In this study, we first recruit human judges to assess the quality of three simulated dialog corpora and then use human judgments as the gold standard to validate the conclusions drawn from the automatic measures. We observe that it is hard for the human judges to reach good agreement when asked to rate the quality of the dialogs from given perspectives. However, the human ratings give consistent ranking of the quality of simulated corpora generated by different simulation models. When building prediction models of human judgments using previously proposed automatic measures, we find that we cannot reliably predict human ratings using a regression model, but we can predict human rankings by a ranking model.

Hua Ai, Diane J. Litman

Real-time Traffic

ACL 2008 | Automatic Measures | Computational Linguistics | Human | Simulated Dialog Corpora |

claim paper

Related Content

» Acquisition and Evaluation of a Dialog Corpus through WOz and Dialog Simulation Techniques

» Pilot Evaluation Study of a Virtual Paracentesis Simulator for Skill Training and Assessme...

» Applying Automated Metrics to Speech Translation Dialogs

» Making memories applying user input logs to interface design and evaluation

» Collaborative simulation interface for planning disaster measures

» Evaluating audio skimming and frame rate acceleration for summarizing BBC rushes

» A Search LogBased Approach to Evaluation

» Use study on a home video editing system

» Current Developments in Information Retrieval Evaluation

Post Info
More Details (n/a)

Added	29 Oct 2010
Updated	29 Oct 2010
Type	Conference
Year	2008
Where	ACL
Authors	Hua Ai, Diane J. Litman

Comments (0)