Speaker
Description
The Penn Parsed Corpus of Historical Yiddish (PPCHY) is a treebank — text annotated with part-of-speech and syntactic information — developed for studying syntactic change in Yiddish. While PPCHY has been a valuable resource, at roughly 200K words it is quite small. Recent work has begun expanding the treebank, both with more modern (20th-century) text and with older material printed in Vaybertaytsh, a semi-cursive script used from the 16th to early 19th century. Substantial additional material survives in this script, and one goal of this work is to automatically annotate OCR'd texts, vastly increasing the treebank's size. This talk gives an overview of that work, including initial experiments applying OCR systems to Vaybertaytsh source texts in the treebank.