2ⁿᵈ Exchange Meeting of SCOOP: International Network for Automated Text Recognition of Historical Sources
University of Vienna
2nd Exchange Meeting, Vienna, September 7-9. 2026
SCOOP (Source Codes of the Past) is an international network dedicated to the automatic transcription and analysis of handwritten historical sources. It brings together people from very different corners of the scholarly world: historians, philologists, and palaeographers, archivists and librarians; computer scientists and machine learning researchers; software engineers asf. What unites them is a shared challenge: how to use automatic/handwritten text recognition (ATR/HTR) to unlock the vast written heritage of the past, and how to do it well. (Read more about SCOOP.)
After a first smaller meeting at Princeton in June of 2025 (for more information see https://fsp-text-edition-blog.univie.ac.at/?p=128), the second SCOOP Exchange Meeting is taking place in Vienna on 7–9 September 2026 hosted by the Institute for Medieval Research at the Austrian Academy of Sciences and the Faculty for Cultural Historical Studies at the University of Vienna.
Over three days, it will bring together around a hundred members of the network for keynotes, parallel working group sessions, round tables, demonstrations, and open discussion formats, along with: as well as an internal working day devoted to the future of SCOOP:
September 7–8: Part of the conference open to registered external visitors
September 9: A dedicated internal working session for SCOOP members
September 10 (Associated Event): The OCR/HTR Workshop for Under-represented and Under-resourced Languages organized by Alíz Horváth - If you wish to join or present, please contact Alíz Horváth (HorvathA@ceu.edu).
-
-
9:00 AM
Registration
-
Plenary: Welcome Room 1
Room 1
-
Keynote Room 1
Room 1
-
1
The Transcription/Edition Distinction as a Foundation for Trustworthy Automatic Text RecognitionSpeaker: Thibault Clérice
-
1
-
11:30 AM
Coffee break
-
Working Group 1: WG1-1 - Technologies & Architectures Room 1
Room 1
-
2
Large Scale Approaches: PartySpeaker: Benjamin Kiessling (ALMAnaCH, Inria)
-
3
VLM Input – Fine-tuning VLMSSpeaker: William Mattingly
-
4
AnandaSkySpeaker: Colin Brisson (EPHE Paris)
-
5
Metatools & Technological AgnosticismSpeaker: Andy Stauder (READ COOP)
-
2
-
1:30 PM
Lunch
-
Working Group 4: Language Challenges I Room 3
Room 3
-
6
Four Scripts, Five Languages, One Digitalization ChallengeSpeaker: Ana Mihaljević (Institute for the Croatian language)
-
7
Matchbox: creating a combined recognition model for early medieval Celtic languages and LatinSpeaker: Bernhard Bauer
-
8
Promising results for under-resourced languages : use cases on Arabic, Greek and ChineseSpeaker: Baptiste Queuche
-
9
Using ATR of Yiddish to study language changeSpeaker: Seth Kulick
-
6
-
Working Group 1: Training as a continuous process Room 1
Room 1
-
10
Piloting HTR in the Library: Experiments and GroundworkSpeaker: Doug Emery (UPenn)
-
11
Multi-Model OCR Arbitration: High-Precision OCR for 19th Century Administrative Mass SourcesSpeaker: Wolfgang Goederle (Graz/Passau)
-
12
From PDF to mARkdown: A Scalable Arabic OCR Pipeline for expanding the OpenITI CorpusSpeaker: Alicia González Martínez (Universität Hamburg)
-
13
Anagnostes - Towards a Transformer-based OCR System for Ancient Greek PapyriSpeakers: Anton Repushko, Elena Chepel (University of Vienna)
-
10
-
Working Group 3: Transcription approaches I (Palaeography in focus) Room 2
Room 2
-
14
Computational Paleography through Automatic Text RecognitionSpeaker: Benjamin Kiessling (ALMAnaCH, Inria)
-
15
Towards Geometry-Based Scribal Hand Analysis: A Word-Level Graphometric Framework for Arabic ManuscriptsSpeaker: Sajjad Nikfahm-Khubravan (Roshan Institute for Persian Studies)
-
16
Leveraging HTR corpora for morphological and metrological paleographical analysisSpeaker: Malamatenia Vlachou (IRHT/LIGM IPParis)
-
14
-
4:00 PM
Coffee break
-
Working Group 4: Language Challenges II Room 3
Room 3
-
17
Dark Vellum, Dense Diacritics: Challenges in Training HTR Models for Medieval IcelandicSpeaker: Katrín Lísa van der Linde Mikaelsdóttir
-
18
Impacts of Script Features on Text Recognition: Experiments with Greek and ArabicSpeaker: Christine Roughan
-
19
Syriac Manuscripts That Can Talk: Giving Voice to Under-Resourced Scripts in HTRSpeaker: Ephrem Aboud Ishac
-
17
-
Working Group 1: Experimental approaches Room 1
Room 1
-
20
Utilization of the Open WebUI platform in ArkindexSpeaker: Michal Racyn (Masaryk University)
-
21
Comparing HTR engines with PolyscriptorSpeaker: Achim Rabus
-
22
Fichero and the Circuit Court Archive of Istmina: An AI Workflow App for ResearchersSpeaker: Daniel Tubb
-
23
Document Layout Analysis for Glossed Medieval Manuscripts: Comparing Kraken 6 and YOLO26Speaker: Tristan Repolusk (University of Graz)
-
20
-
Working Group 3: Transcription approaches II (Transcription decisions, their effect, and related terminology) Room 2
Room 2
-
24
Otto Dei gram, rogat vram Clam. From facsimile and descriptive transcriptions to “quasi-diplomatic” and interpretative approaches in ATR/HTRSpeaker: Jan Odstrčilík (Institut für Mittelalterforschung, ÖAW)
-
25
Lost in Transcription? Reconciling HTR Output with Old Czech Editorial StandardsSpeaker: Anna Michalcová (Czech Academy of Sciences / IMAFO, ÖAW)
-
26
HTR of Normalized Latin Texts: Insights from Liturgical ManuscriptsSpeaker: Paweł Figurski (Polish Academy of Sciences)
-
27
Reading Glagolitic in the Twenty-First Century: Between Philology, HTR, and AISpeaker: Ana Mihaljević (Institute for the Croatian language)
-
24
-
6:00 PM
Break
-
Keynote Room 1
Room 1
-
28
New Epistemic Frontiers: LLMs for Linking Transcription, Editing, and InterpretationSpeakers: Anna Dolganov, David Smith (Northwestern)
-
28
-
7:30 PM
Reception
-
9:00 AM
-
-
Working Group 5: Datasets and Institutions Room 2
Room 2
-
29
Pipelines and WorkflowsSpeaker: Jessie Dummer
-
30
Manuscriptorium Full-Text Module – The newest component of the Manuscriptorium digital librarySpeaker: Michael Lužný
-
31
How do you revive a legacy dataset?Speaker: Tim Geelhaar
-
32
The current status of the field of HTR within the German manuscript centresSpeaker: Ursula Stampfer
-
29
-
Working Group 6: Leveraging Outputs: Text Reuse, NLP, and More (talks) Room 3
Room 3
-
33
Testing the Reuse Value of Pracalit Ground Truth for Bhujimol ManuscriptsSpeaker: Alexander O'Neill
-
34
From Old Icelandic HTR Outputs to Normalised Texts through Seq2Seq TransformersSpeaker: Nikola Krisztian Czindrity
-
35
New Ways to Transform and Explore Older Korean Texts: From ATR to Reading Texts to Agentic Retrieval in Mo文oNExplorerSpeaker: Seth Kulick
-
33
-
Working Group 1: Roundtable Room 1
Room 1
-
11:30 AM
Coffee Break
-
Working Group 4: Roundtable 🎤 - Language Challenges Room 2
Room 2
-
Working Group 6: Leveraging Outputs: Text Reuse, NLP, and More (demos) Room 3
Room 3
-
36
New Ways to Transform and Explore Older Korean Texts: From ATR to Reading Texts to Agentic Retrieval in Mo文oNExplorerSpeaker: Wayne de Fremery
-
37
From Document Images to Research CatalogueSpeaker: Andrew Janco
-
38
TBDSpeaker: Martin Roček
-
36
-
Working Group 2: Handwriting Classification I Room 1
Room 1
-
39
Classifying Squeezes Again: Initial Results from an ICDAR ContestSpeaker: Aaron Hershkowitz
-
40
Communities of Practice: Capturing the Aspect of Late Medieval HandwritingSpeakers: Sam Grieggs, Sebastian Sobecki
-
41
Script classification and alphabet identificationSpeaker: Giuseppe de Gregorio
-
39
-
1:30 PM
Lunch
-
Working Group 6: Unconference Room 3
Room 3
-
Working Group 2: Handwriting Classification II Room 1
Room 1
-
42
[TBD - Handwriting classification]Speakers: Asimina Paparrigopoulou, Paraskevi (Vivian) Platanou
-
43
Classification of Armenian manuscripts without OCR/HTR preprocessingSpeaker: Tara Andrews
-
44
[TBD - Explainability]Speaker: Serena Ammirati
-
42
-
Working Group 3: Building ATR/HTR pipelines Room 2
Room 2
-
45
The Pragmatics of ATR Pipelines: Structural Dilemmas in Designing Workflows for Historical CorporaSpeaker: Michael Schonhardt (Univ. Freiburg)
-
46
Building an ATR pipeline for tabular data and script translationSpeaker: Olaf Berg (Ruhr-Universität Bochum, AI:RUB, ERC-LOOP)
-
47
Iterative HTR: A new pipeline for automatic text recognition of Arabic-scriptSpeaker: Osama Eshera (University of Maryland)
-
48
Developing a pipeline for OCR at the Austrian National LibrarySpeaker: Johannes Knüchel (Austrian National Library)
-
45
-
4:00 PM
Coffee Break
-
49
Working Groups working on conclusions
-
6:00 PM
Short Break
-
50
Final Roundtable - Open data, open code, open minds in AI
-
51
Conclusion of the public part
-
-
-
52
Internal meeting of SCOOP
-
52