Introduction
This file contains documentation on the Arabic Treebank: Part 4 v 1.0 (MPG Annotation), Linguistic
Data Consortium (LDC) catalog number LDC2005T30 and ISBN 1-58563-343-7.
The goal of the Arabic Treebank project is to support the development of data-driven approaches to natural language
processing (NLP), human language technologies, automatic content extraction
(topic extraction and/or grammar extraction), cross-lingual information
retrieval, information detection, and other forms of linguistic research on
Modern Standard Arabic in general, the LDC was sponsored to develop an
Arabic POS and Treebank of 1,000,000 words. This corpus is the fourth part of
that project. In this release, we provide annotation on part of speech
(POS), gloss, and word segmentation.
Samples
To view a example of this corpus, please review this sample POS file.
Content Copyright
Portions © 2004 Assabah Press Group, © 2005 Trustees of the University Pennsylvania. |