Session

Working Group 3

WG3
Sep 7, 2026, 1:30 PM
Room 2

Room 2

Presentation materials

There are no materials yet.

  1. Benjamin Kiessling (ALMAnaCH, Inria Paris)
    9/7/26, 1:30 PM

    Computational methods have transformed many areas of historical research, yet
    paleography has seen comparatively few algorithms and tools that can operate at
    scale while remaining adaptable to scholarly practices across different
    scripts. This is not due to a lack of interest in digital methods. The growth
    of digital repositories, annotation platforms such as DigiPal, and...

    Go to contribution page
  2. Sajjad Nikfahm-Khubravan (Roshan Institute for Persian Studies, University of Maryland)
    9/7/26, 1:50 PM

    Arabic manuscript traditions present a constellation of challenges that continue to resist the assumptions underlying most contemporary Handwritten Text Recognition (HTR) systems. Unlike Latin scripts, Arabic is written right-to-left in a fully cursive hand in which most, but not all, letters are connected to their neighbors, and the graphical form of each letter changes depending on its...

    Go to contribution page
  3. Malamatenia Vlachou (IRHT/CNRS-ENPC)
    9/7/26, 2:10 PM

    The analysis of medieval handwriting lies at the heart of palaeographical research. While Automatic Text Recognition (ATR) systems are primarily developed to produce transcriptions, they also generate rich data that can be leveraged to study the scripts themselves. This presentation will explore how ATR data can be exploited to develop new computational tools for palaeographical analysis....

    Go to contribution page
  4. Dominique Stutzmann (IRHT-CNRS / HU Berlin)
    9/7/26, 2:30 PM

    Transcription levels, the identification of scribal hands, and workflow architectures are usually discussed in three separate conversations, one editorial, one palaeographical, one technical. I fold them into two decisions taken before any computational result exists: at what grain the written surface is read (allographetic, graphematic, normalised on the text side; glyph, patch, line, page on...

    Go to contribution page
  5. Jan Odstrčilík (Institute for Medieval Research, Austrian Academy of Sciences)
    9/7/26, 3:30 PM

    The rise of automated/handwritten text recognition has revived an old discussion that had seemed otherwise dormant: How should historical texts be transcribed? The main tension is between two approaches: the first calls for true-to-character transcriptions, while the second purposefully modifies the text, following long-established philological traditions. There are many confusing names used...

    Go to contribution page
  6. Anna Michalcová (Czech Language Institute, Czech Academy of Sciences; Institute for Medieval Research, Austrian Academy of Sciences)
    9/7/26, 3:50 PM

    Handwritten text recognition (HTR) is usually judged by how faithfully it reproduces what stands on the page. Czech editorial practice asks for something else. Old Czech orthography leaves consonants and vowel quantity underdetermined, so a single written form regularly supports several grammatically sound readings. Vowel quantity in Czech carries grammar, not just sound. An edition or a...

    Go to contribution page
  7. Ana Mihaljević (Institute for the Croatian language)
    9/7/26, 4:10 PM

    The digitization and scholarly editing of Glagolitic texts are entering a new phase in which traditional philological practices increasingly intersect with automated text recognition and artificial intelligence. This paper reflects on recent experiences with the transcription, transliteration, and digital editing of Croatian Glagolitic manuscripts, focusing on the opportunities and limitations...

    Go to contribution page
  8. Paweł Figurski (Polish Academy of Sciences)
    9/7/26, 4:30 PM

    Nearly eleven hundred sacramentaries copied up to c. 1100 — complete codices, fragments, and palimpsests — survive, dispersed across collections worldwide. The scale of this corpus has kept it outside granular analysis, and the tradition is still described through a handful of types established by earlier scholarship. RITUS+, a tool under development, addresses that scale by treating...

    Go to contribution page
  9. Michael Schonhardt (TU Darmstadt / Akademie der Wissenschaften und der Literatur | Mainz)
    9/8/26, 1:30 PM

    Automatic Text Recognition has become the most widely adopted application of machine learning in the humanities, yet projects often encounter challenges when integrating existing models into their workflow. This paper dicusses the source of these challenges as a mismatch between requirement profiles - a concept broader than transcription guidelines, constituted by three dimensions:...

    Go to contribution page
  10. Olaf Berg (Ruhr-Universität Bochum)
    9/8/26, 1:50 PM
    1

    My presentation addresses the technical challenges of applying automated text recognition (ATR) to demographic data preserved in late Ottoman census registers. It builds on the experience we gained in the LOOP research project—Late Ottoman Palestinians. Two issues are central. First, the registers are handwritten in Ottoman Turkish, an under-resourced historical language for which effective...

    Go to contribution page
  11. Osama Eshera (University of Maryland)
    9/8/26, 2:10 PM

    We study word-level text-to-image mapping in Arabic-script manuscripts: given a line image and its transcription, locate each word. At corpus scale, accurate word geometry opens new possibilities for digital paleographical analysis while improving ATR both upstream and downstream. This mapping is learned from 25,000 line crops drawn from 87 manuscripts, supervising word geometry only through...

    Go to contribution page
  12. Johannes Knüchel (Austrian National Library)
    9/8/26, 2:30 PM

    One of the Austrian National Library’s Strategic Goals for 2023–2027 is to create new ways for users to explore its collections. The improvement of ATR, especially OCR, falls within this context. During my presentation, I will talk about the currently ongoing OCR Project (2024–2027) and one of its main outcomes, an internal OCR Service. This Service contains a modular OCR pipeline using...

    Go to contribution page
Building timetable...