Linguistic resources for Arabic machine translation
In this chapter, we describe the linguistic resources particularly suitable for Arabic machine translation research and technology development that are published by the Linguistic Data Consortium (LDC). A significant number of the data sets developed by LDC are Arabic language resources, making LDC the leading source for such materials. LDC’s Arabic language resources represent all of the data types in LDC’s Catalog: speech, text, video and lexicons. Much of this data has already been used for machine translation research, as represented in this volume and throughout the field, and new datasets in the pipeline are expected to benefit machine translation work as well.