Using OCR to expand a treebank of historical Yiddish: A Progress Report

Sep 8, 2026, 10:00 AM
30m
Room 3

Room 3

Speaker

Seth Kulick (Linguistic Data Consortium, University of Pennsylvania)

Description

The Penn Parsed Corpus of Historical Yiddish (PPCHY) is a treebank — text annotated with part-of-speech and syntactic information — developed for studying syntactic change in Yiddish. While PPCHY has been a valuable resource, at roughly 200K words it is quite small. Recent work has begun expanding the treebank, both with more modern (20th-century) text and with older material printed in Vaybertaytsh, a semi-cursive script used from the 16th to early 19th century. Substantial additional material survives in this script, and one goal of this work is to automatically annotate OCR'd texts, vastly increasing the treebank's size. This talk gives an overview of that work, including initial experiments applying OCR systems to Vaybertaytsh source texts in the treebank.

Presentation materials

There are no materials yet.