2ⁿᵈ Exchange Meeting of SCOOP: International Network for Automated Text Recognition of Historical Sources

Europe/Vienna
University of Vienna

University of Vienna

Währinger Straße 29 1090 Vienna Austria
Description

2nd Exchange Meeting, Vienna, September 7-9. 2026 

SCOOP (Source Codes of the Past) is an international network dedicated to the automatic transcription and analysis of handwritten historical sources. It brings together people from very different corners of the scholarly world: historians, philologists, and palaeographers, archivists and librarians; computer scientists and machine learning researchers; software engineers asf. What unites them is a shared challenge: how to use automatic/handwritten text recognition (ATR/HTR) to unlock the vast written heritage of the past, and how to do it well. (Read more about SCOOP.)

After a first smaller meeting at Princeton in June of 2025 (for more information see https://fsp-text-edition-blog.univie.ac.at/?p=128), the second SCOOP Exchange Meeting is taking place in Vienna on 7–9 September 2026 hosted by the Institute for Medieval Research at the Austrian Academy of Sciences and the Faculty for Cultural Historical Studies at the University of Vienna.

Over three days, it will bring together around a hundred members of the network for keynotes, parallel working group sessions, round tables, demonstrations, and open discussion formats, along with: as well as an internal working day devoted to the future of SCOOP:

September 7–8: Part of the conference open to registered external visitors

September 9: A dedicated internal working session for SCOOP members

September 10 (Associated Event): The OCR/HTR Workshop for Under-represented and Under-resourced Languages organized by Alíz Horváth - If you wish to join or present, please contact Alíz Horváth (HorvathA@ceu.edu).

Registration
Register to 2nd Exchange Meeting of SCOOP (7-8 September)
    • 9:00 AM 10:00 AM
      Registration 1h
    • 10:00 AM 10:15 AM
      Plenary: Welcome Room 1

      Room 1

    • 10:15 AM 11:30 AM
      Keynote Room 1

      Room 1

      • 10:15 AM
        The Transcription/Edition Distinction as a Foundation for Trustworthy Automatic Text Recognition 1h 15m
        Speaker: Thibault Clérice
    • 11:30 AM 12:00 PM
      Coffee break 30m
    • 12:00 PM 1:30 PM
      Working Group 1: WG1-1 - Technologies & Architectures Room 1

      Room 1

      • 12:00 PM
        Large Scale Approaches: Party 20m
        Speaker: Benjamin Kiessling (ALMAnaCH, Inria)
      • 12:20 PM
        VLM Input – Fine-tuning VLMS 20m
        Speaker: William Mattingly
      • 12:40 PM
        AnandaSky 20m
        Speaker: Colin Brisson (EPHE Paris)
      • 1:00 PM
        Metatools & Technological Agnosticism 20m
        Speaker: Andy Stauder (READ COOP)
    • 1:30 PM 2:30 PM
      Lunch 1h
    • 2:30 PM 4:00 PM
      Working Group 4: Language Challenges I Room 3

      Room 3

      • 2:30 PM
        Four Scripts, Five Languages, One Digitalization Challenge 20m
        Speaker: Ana Mihaljević (Institute for the Croatian language)
      • 2:50 PM
        Matchbox: creating a combined recognition model for early medieval Celtic languages and Latin 20m
        Speaker: Bernhard Bauer
      • 3:10 PM
        Promising results for under-resourced languages : use cases on Arabic, Greek and Chinese 20m
        Speaker: Baptiste Queuche
      • 3:30 PM
        Using ATR of Yiddish to study language change 20m
        Speaker: Seth Kulick
    • 2:30 PM 4:00 PM
      Working Group 1: Training as a continuous process Room 1

      Room 1

      • 2:30 PM
        Piloting HTR in the Library: Experiments and Groundwork 20m
        Speaker: Doug Emery (UPenn)
      • 2:50 PM
        Multi-Model OCR Arbitration: High-Precision OCR for 19th Century Administrative Mass Sources 20m
        Speaker: Wolfgang Goederle (Graz/Passau)
      • 3:10 PM
        From PDF to mARkdown: A Scalable Arabic OCR Pipeline for expanding the OpenITI Corpus 20m
        Speaker: Alicia González Martínez (Universität Hamburg)
      • 3:30 PM
        Anagnostes - Towards a Transformer-based OCR System for Ancient Greek Papyri 20m
        Speakers: Anton Repushko, Elena Chepel (University of Vienna)
    • 2:30 PM 4:00 PM
      Working Group 3: Transcription approaches I (Palaeography in focus) Room 2

      Room 2

      • 2:30 PM
        Computational Paleography through Automatic Text Recognition 30m
        Speaker: Benjamin Kiessling (ALMAnaCH, Inria)
      • 3:00 PM
        Towards Geometry-Based Scribal Hand Analysis: A Word-Level Graphometric Framework for Arabic Manuscripts 30m
        Speaker: Sajjad Nikfahm-Khubravan (Roshan Institute for Persian Studies)
      • 3:30 PM
        Leveraging HTR corpora for morphological and metrological paleographical analysis 30m
        Speaker: Malamatenia Vlachou (IRHT/LIGM IPParis)
    • 4:00 PM 4:30 PM
      Coffee break 30m
    • 4:30 PM 6:00 PM
      Working Group 4: Language Challenges II Room 3

      Room 3

      • 4:30 PM
        Dark Vellum, Dense Diacritics: Challenges in Training HTR Models for Medieval Icelandic 20m
        Speaker: Katrín Lísa van der Linde Mikaelsdóttir
      • 4:50 PM
        Impacts of Script Features on Text Recognition: Experiments with Greek and Arabic 20m
        Speaker: Christine Roughan
      • 5:10 PM
        Syriac Manuscripts That Can Talk: Giving Voice to Under-Resourced Scripts in HTR 30m
        Speaker: Ephrem Ishac
    • 4:30 PM 6:00 PM
      Working Group 1: Experimental approaches Room 1

      Room 1

      • 4:30 PM
        Utilization of the Open WebUI platform in Arkindex 20m
        Speaker: Michal Racyn (Masaryk University)
      • 4:50 PM
        Comparing HTR engines with Polyscriptor 20m
        Speaker: Achim Rabus
      • 5:10 PM
        Fichero and the Circuit Court Archive of Istmina: An AI Workflow App for Researchers 20m
        Speaker: Daniel Tubb
      • 5:30 PM
        Document Layout Analysis for Glossed Medieval Manuscripts: Comparing Kraken 6 and YOLO26 20m
        Speaker: Tristan Repolusk (University of Graz)
    • 4:30 PM 6:00 PM
      Working Group 3: Transcription approaches II (Transcription decisions, their effect, and related terminology) Room 2

      Room 2

      • 4:30 PM
        Otto Dei gram, rogat vram Clam. From facsimile and descriptive transcriptions to “quasi-diplomatic” and interpretative approaches in ATR/HTR 20m
        Speaker: Jan Odstrčilík (Institut für Mittelalterforschung, ÖAW)
      • 4:50 PM
        Lost in Transcription? Reconciling HTR Output with Old Czech Editorial Standards 20m
        Speaker: Anna Michalcová (Czech Academy of Sciences / IMAFO, ÖAW)
      • 5:10 PM
        HTR of Normalized Latin Texts: Insights from Liturgical Manuscripts 20m
        Speaker: Paweł Figurski (Polish Academy of Sciences)
      • 5:30 PM
        Reading Glagolitic in the Twenty-First Century: Between Philology, HTR, and AI 20m
        Speaker: Ana Mihaljević (Institute for the Croatian language)
    • 6:00 PM 6:15 PM
      Break 15m
    • 6:15 PM 7:30 PM
      Keynote Room 1

      Room 1

      • 6:15 PM
        New Epistemic Frontiers: LLMs for Linking Transcription, Editing, and Interpretation 1h 15m
        Speakers: Anna Dolganov, David Smith (Northwestern)
    • 7:30 PM 9:00 PM
      Reception 1h 30m
    • 10:00 AM 11:30 AM
      Working Group 5: Datasets and Institutions Room 2

      Room 2

      • 10:00 AM
        Pipelines and Workflows 20m
        Speaker: Jessie Dummer
      • 10:20 AM
        Manuscriptorium Full-Text Module – The newest component of the Manuscriptorium digital library 20m
        Speaker: Michael Lužný
      • 10:40 AM
        How do you revive a legacy dataset? 20m
        Speaker: Tim Geelhaar
      • 11:00 AM
        The current status of the field of HTR within the German manuscript centres 20m
        Speaker: Ursula Stampfer
    • 10:00 AM 11:30 AM
      Working Group 6: Leveraging Outputs: Text Reuse, NLP, and More (talks) Room 3

      Room 3

      • 10:00 AM
        Testing the Reuse Value of Pracalit Ground Truth for Bhujimol Manuscripts 30m
        Speaker: Alexander O'Neill
      • 10:30 AM
        From Old Icelandic HTR Outputs to Normalised Texts through Seq2Seq Transformers 30m
        Speaker: Nikola Krisztian Czindrity
      • 11:00 AM
        New Ways to Transform and Explore Older Korean Texts: From ATR to Reading Texts to Agentic Retrieval in Mo文oNExplorer 30m
        Speaker: Seth Kulick
    • 10:00 AM 11:30 AM
      Working Group 1: Roundtable Room 1

      Room 1

    • 11:30 AM 12:00 PM
      Coffee Break 30m
    • 12:00 PM 1:30 PM
      Working Group 4: Roundtable 🎤 - Language Challenges Room 2

      Room 2

    • 12:00 PM 1:30 PM
      Working Group 6: Leveraging Outputs: Text Reuse, NLP, and More (demos) Room 3

      Room 3

      • 12:00 PM
        New Ways to Transform and Explore Older Korean Texts: From ATR to Reading Texts to Agentic Retrieval in Mo文oNExplorer 30m
        Speaker: Wayne de Fremery
      • 12:30 PM
        From Document Images to Research Catalogue 30m
        Speaker: Andrew Janco
      • 1:00 PM
        TBD 30m
        Speaker: Martin Roček
    • 12:00 PM 1:30 PM
      Working Group 2: Handwriting Classification I Room 1

      Room 1

      • 12:00 PM
        Classifying Squeezes Again: Initial Results from an ICDAR Contest 30m
        Speaker: Aaron Hershkowitz
      • 12:30 PM
        Communities of Practice: Capturing the Aspect of Late Medieval Handwriting 30m
        Speakers: Sam Grieggs, Sebastian Sobecki
      • 1:00 PM
        Script classification and alphabet identification 30m
        Speaker: Giuseppe de Gregorio
    • 1:30 PM 2:30 PM
      Lunch 1h
    • 2:30 PM 4:00 PM
      Working Group 6: Unconference Room 3

      Room 3

    • 2:30 PM 4:00 PM
      Working Group 2: Handwriting Classification II Room 1

      Room 1

      • 2:30 PM
        [TBD - Handwriting classification] 30m
        Speakers: Asimina Paparrigopoulou, Paraskevi (Vivian) Platanou
      • 3:00 PM
        Classification of Armenian manuscripts without OCR/HTR preprocessing 30m
        Speaker: Tara Andrews
      • 3:30 PM
        [TBD - Explainability] 30m
        Speaker: Serena Ammirati
    • 2:30 PM 4:00 PM
      Working Group 3: Building ATR/HTR pipelines Room 2

      Room 2

      • 2:30 PM
        The Pragmatics of ATR Pipelines: Structural Dilemmas in Designing Workflows for Historical Corpora 20m
        Speaker: Michael Schonhardt (Univ. Freiburg)
      • 2:50 PM
        Building an ATR pipeline for tabular data and script translation 20m
        Speaker: Olaf Berg (Ruhr-Universität Bochum, AI:RUB, ERC-LOOP)
      • 3:10 PM
        Iterative HTR: A new pipeline for automatic text recognition of Arabic-script 20m
        Speaker: Osama Eshera (University of Maryland)
      • 3:30 PM
        Developing a pipeline for OCR at the Austrian National Library 20m
        Speaker: Johannes Knüchel (Austrian National Library)
    • 4:00 PM 4:30 PM
      Coffee Break 30m
    • 4:30 PM 6:00 PM
      Working Groups working on conclusions 1h 30m
    • 6:00 PM 6:15 PM
      Short Break 15m
    • 6:15 PM 7:45 PM
      Final Roundtable - Open data, open code, open minds in AI 1h 30m
    • 7:45 PM 8:00 PM
      Conclusion of the public part 15m
    • 10:00 AM 6:05 PM
      Internal meeting of SCOOP 8h 5m