Session

Working Group 6

Sep 8, 2026, 9:00 AM
University of Vienna

University of Vienna

Währinger Straße 29 1090 Vienna Austria

Presentation materials

There are no materials yet.

  1. Alexander O'Neill (Musashino University)
    9/8/26, 9:00 AM

    This presentation examines the applicability of ground truth for handwritten text recognition (HTR) on Nepalese manuscripts written in the Pracalit script (which is commonly used for Sanskrit and Newar manuscripts dating from approximately the 16th to 20th centuries) to use on an earlier Nepalese script, Bhujimol (which was prevalent from the 11th to 17th centuries). The presentation will...

    Go to contribution page
  2. Nikola Krisztian Czindrity (University of Vienna)
    9/8/26, 9:30 AM

    For an Old Icelandic manuscript-to-edition pipeline, we train a 10M-parameter PyTorch CharSeq2Seq Transformer to transform facsimile-like transcriptions into diplomatic transcriptions (facs2dipl task) and, in turn, into normalised ones (dipl2norm task; https://huggingface.co/NKCZ/old-icelandic-facs2dipl2norm).

    We train the model on around 30,000 line-level triples of facsimile-like,...

    Go to contribution page
  3. Seth Kulick (Linguistic Data Consortium, University of Pennsylvania)
    9/8/26, 10:00 AM

    The Penn Parsed Corpus of Historical Yiddish (PPCHY) is a treebank — text annotated with part-of-speech and syntactic information — developed for studying syntactic change in Yiddish. While PPCHY has been a valuable resource, at roughly 200K words it is quite small. Recent work has begun expanding the treebank, both with more modern (20th-century) text and with older material printed in...

    Go to contribution page
  4. Wayne de Fremery (Dominican University of California)
    9/8/26, 11:00 AM

    The bibliographer and literary critic Jerome McGann has written that “original documents are fictions we practice in order to manage their losses and our limits” (A New Republic of Letters, 5). This talk considers how a corpus of early twentieth-century Korean novels and periodicals is being practiced as a new kind of fiction within prototype software called Mo文oNExplorer. It describes how a...

    Go to contribution page
  5. Andrew Janco (Princeton University)
    9/8/26, 11:20 AM

    This demonstration builds on our experience with 19th-century archival documents from Chocó, Colombia. As researchers digitised these documents, there was a concurrent need for datafication to assess the collection's scope, content, and research potential. We developed a minimal tool to extract text using vision-language models (VLMs), identify and normalise entities, and create a search...

    Go to contribution page
  6. Daniel Tubb (Anthropology, University of New Brunswick, Canada)
    9/8/26, 11:40 AM

    In this talk, I demo Fichero. Fichero is a macOS app, a SwiftUI front end over a FastAPI Python engine, that puts machine-learning and AI workflows, using local and hosted models, in researchers' hands. Developed from a collaboration with Andy Janco (Princeton) and Ann Farnsworth-Alvear (UPenn), I demonstrate Fichero on the Circuit Court Archive of Istmina, Colombia. The archive is a British...

    Go to contribution page
  7. Martin Roček (Institute for Medieval Research, Austrian Academy of Sciences and Faculty of Arts, Charles University)
    9/8/26, 12:00 PM

    What if we reframe intertextuality detection as a retrieval task? In this talk, I will take quotations from an author and rank them against corpus of possible sources. This allows me to measure recall@10, precision@1, nDCG@10 and compare four different methods that are commonly used: BM25, dense embeddings (generic and Latin-specific), reciprocal-rank fusion, and reranking by cross-encoder or...

    Go to contribution page
Building timetable...