November 2017 Newsletter
New Publications
ASpIRE Development and Development Test Sets
IARPA Babel Kurmanji Kurdish Language Pack IARPA-babel205b-v1.0a
TAC KBP Chinese Cross-lingual Entity Linking - Comprehensive Training & Evaluation Data 2011-2014
Announcements
In addition to receiving new publications, current year LDC members also enjoy the benefit of licensing older data at reduced costs from our Catalog of over 700 holdings; current year for-profit members may use most data for commercial applications.
- Multilanguage conversational telephone speech: developed to support language identification research in related languages (Central Asian, Central European language groups)
- DIRHA (Distant-speech Interaction for Robust Home Applications): Wall Street Journal read speech with noise and reverberation, suitable for various multi-microphone signal processing and distant speech recognition tasks
- TRAD corpora: Chinese-French and Arabic-French parallel text (newswire, web data)
- IARPA Babel Language Packs (telephone speech and transcripts):languages include Cebuano, Guarani, Kazakh, Lithuanian, Telugu, Tok Pisin
- BOLT: discussion forum, SMS, word-aligned, and tagged data in all languages (Egyptian Arabic, English, Chinese)
- DEFT: Spanish Treebank (newswire, web data)
- RATS: Language Identification data set (Dari, Farsi, Levantine Arabic, Pashto, Urdu; degraded audio signals)
- TAC KBP: comprehensive English source and entity linked data (broadcast, telephone speech, newswire, web data)
- German children’s handwriting: longitudinal study of weekly writing in classroom setting with enhanced output for specific spelling patterns
And don’t forget, MY2017 and MY2016 are still open for joining. MY2016 can be joined through December 31, 2017 and includes data such as BOLT Chinese Discussion Forums, IARPA Babel Language Packs in multiple languages and Multi-Language Conversational Telephone Speech – Slavic Group. MY 2017 will remain open through December 31, 2018; among the year’s releases are 2010 NIST Speaker Recognition Evaluation Test Set, RATS Keyword Spotting, Noisy TIMIT Speech and BOLT Egyptian Arabic SMS/Chat and Transliteration. For full descriptions of these data sets, browse our Catalog.