Obtaining DataUsing DataProviding DataCreating Data
About LDCMembersCatalogProjectsPapersLDC OnlineSearchContact UsUPennHome

LDC Catalog | By Type and Source | By Year | Top Ten | Projects | Catalog Search



1997 HUB5 English Evaluation

Item Name: 1997 HUB5 English Evaluation
Authors: NIST Multimodal Information Group
LDC Catalog No.: LDC2002S23
ISBN: 1-58563-233-3
Data Type: speech
Sample Rate: 8000 Hz
Sampling Format: 2-channel ulaw
Data Source(s): telephone conversations
Application(s): speech recognition
Language(s): English
Language ID(s): eng
Distribution: 1 CD
Member fee: $0 for 2002 members
Non-member Fee: N/A (Members Only)
Reduced-License Fee: N/A
Extra-Copy Fee: US $0.00
Member License: yes
Online documentation: yes
Licensing Instructions: Subscription Members, Standard Members, Non-Members
Citation: NIST Multimodal Information Group
2002
1997 HUB5 English Evaluation
Linguistic Data Consortium, Philadelphia

Introduction

This publication contains the speech data and transcripts used as the EvalSet data in NIST's 1997 Hub5 English Evaluation. The DevSet data for this evaluation consists of the EvalSet data from the CALLHOME American English Speech corpus.

The 1997 Hub5 English evaluation was part of an ongoing series of periodic evaluations conducted by NIST. These evaluations provide an important contribution to the direction of research efforts and the calibration of technical capabilities. They are intended to be of interest to all researchers working on the general problem of conversational speech recognition. To this end the evaluation was designed to be simple, to focus on core speech technology issues, to be fully supported, and to be accessible.

More information is available at NIST's web site for The 1997 Hub-5E Spring Evaluation.

Data

This publication contains 40 audio sphere files, each with a corresponding transcript. The sphere headers have been modified from the original evaluation data by the addition of sample checksums to the CALLHOME data files.

An included documentation table contains information on the speech segments to be processed as follows:

        
   
    ...
    

Updates

There are no updates at this time.

Content Copyright

Portions © 1996-2002 Trustees of the University of Pennsylvania.


About LDC | Members | Catalog | Projects | Papers | LDC Online | Search / Help | Contact Us | UPenn | Home | Obtaining Data | Creating Data | Using Data | Providing Data

Contact: ldc@ldc.upenn.edu

(c) 1992-2010 Linguistic Data Consortium, University of Pennsylvania. All Rights Reserved.