Speaker
Description
Historical treebanks — corpora annotated with part-of-speech and syntactic information — are essential resources for linguists studying language change, but they are labor-intensive to build. Existing historical treebanks are consequently small, typically around 1-2 million words, which limits the scope of research they can support. This talk gives an overview of work that explores whether that bottleneck can be broken by applying automatic syntactic annotation to large amounts of historical text containing OCR or OCR-like errors, dramatically increasing treebank size without manual annotation. The central question is whether automatic annotation on such imperfect historical text is accurate enough to serve as a reliable basis for modeling syntactic change. Current work focuses on the case of Yiddish texts from the 16th and 17th centuries.