-
William Mattingly (Yale University)9/7/26, 11:00 AM
The rapid evolution of frontier Vision-Language Models (VLMs), such as the Gemini 3.5 family, has introduced unprecedented zero-shot capabilities in complex multimodal tasks, including high-fidelity document transcription and precise spatial reasoning via bounding box generation. As these massive, proprietary models increasingly solve general visual-linguistic tasks without task-specific...
Go to contribution page -
Colin Brisson (Ecolé pratique des hautes études)9/7/26, 11:20 AM
This presentation introduces AnandaSky, a compact vision-language model developed for line-level transcription of historical sinographic documents, and reports ongoing work toward a new page-level version designed for broader cross-domain use. AnandaSky combines a shallow, high-resolution visual encoder with a Qwen3-0.6B autoregressive decoder. Global visual attention, 10-pixel patches, an...
Go to contribution page -
Andy Stauder (READ COOP)9/7/26, 11:40 AM
It seems the age of thinking in tools is at an end. While Large Language Models offer rapid text processing, generic chat interfaces struggle with complex layout analysis, historical scripts, and verifiable ground truth, and especially control over workflows. As a toolbox, Transkribus follows a different path, aiming for a necessary alternative and an essential complement to LLM-driven...
Go to contribution page -
Doug Emery (University of Pennsylvania Libraries)9/7/26, 1:30 PM
Since 2023, members of Penn Libraries’ Cultural Heritage Computing, Schoenberg Institute for Manuscript Studies, and Research Data and Digital Scholarship groups have conducted workshops and pilot projects to explore workflows for handwritten text recognition (HTR) and its potential to support transcription of library collections for discovery, access, and computational use. This talk will...
Go to contribution page -
Wolfgang Göderle (University of Graz | University of Passau | MPI GEA), David Fleischhacker (University of Graz)9/7/26, 1:50 PM
The transcription of historical administrative sources presents a distinctive set of challenges for automated text recognition: dense, formulaic layouts, highly abbreviated Latin and German terminology, and the cumulative degradation typical of large serial document corpora. This paper presents a high-precision OCR pipeline developed within the Unlocking the Schematismus project, designed to...
Go to contribution page -
Alicia González Martínez (Hamburg University)9/7/26, 2:10 PM
This talk presents a scalable pipeline for converting digitised Arabic books from PDFs into structured mARkdown for integration into the OpenITI corpus. The workflow begins by preparing the PDFs and processing their opening pages to extract the bibliographic metadata required to identify works and editions. The documents are then transcribed in parallel by several complementary recognition...
Go to contribution page -
Elena Chepel (University of Vienna), Anton Repushko (Independent Researcher)9/7/26, 2:30 PM
In this paper, we present a new text recognition tool, Anagnostes, which
Go to contribution page
combines modern approaches in machine learning with optical character
recognition (OCR). It has been trained to read Greek literary and
documentary papyri from images, with a particular emphasis on
recognising not only neat and well-preserved fragments but also cursive
and damaged texts. Currently, the model achieves... -
Michal Racyn (Masaryk University)9/7/26, 3:30 PM
Arkindex (a document processing platform developed by TEKLIA) represents a promising open-source alternative to Transcribus and e-Scriptorium. Based on modular approach Arkindex implements machine learning algorithms that can be applied for layout analysis, transcription, structuration, named entity recognition, and information extraction of text documents as well as object segmentation,...
Go to contribution page -
Tristan Repolusk (Department of Digital Humanities, University of Graz)9/7/26, 3:50 PM
Medieval manuscripts, unlike modern printed documents, are typically characterized by non-uniform script and complex page layout, both with respect to larger regions (e.g., main text regions and marginal regions) and individual text lines. Especially manuscripts featuring extensive annotations such as glosses or reading signs tend to increase the page layout complexity significantly: this...
Go to contribution page -
Elise Wang (California State University, Fullerton)9/7/26, 4:10 PM
The English common law rests on a documentary record of staggering size and frustrating inaccessibility. At least eleven million pages of medieval legal manuscripts survive at the National Archives at Kew, the product of a documentary machine at Westminster and county courts that churned out records at a rate of tens of thousands per year. For comparison, roughly six to seven thousand literary...
Go to contribution page -
Tobias Hodel (University of Bern)9/7/26, 4:30 PM
As Automatic Text Recognition (ATR) pipelines rapidly evolve, the digital humanities are confronted with a highly fragmented landscape of transcription technologies. Researchers today must navigate between specialised, fine-tuned line-level engines, such as Kraken, TrOCR, and the emerging zero-shot capabilities of generalised Vision-Language Models (VLMs). All approaches have different...
Go to contribution page
Choose timezone
Your profile timezone: