Towards Geometry-Based Scribal Hand Analysis: A Word-Level Graphometric Framework for Arabic Manuscripts

Sep 7, 2026, 1:50 PM
20m
Room 2

Room 2

Speaker

Sajjad Nikfahm-Khubravan (Roshan Institute for Persian Studies, University of Maryland)

Description

Arabic manuscript traditions present a constellation of challenges that continue to resist the assumptions underlying most contemporary Handwritten Text Recognition (HTR) systems. Unlike Latin scripts, Arabic is written right-to-left in a fully cursive hand in which most, but not all, letters are connected to their neighbors, and the graphical form of each letter changes depending on its position within the word—initial, medial, final, or isolated. This positional morphology means that no letter has a single canonical shape; its geometry is always a function of its immediate lexical context. Compounding this, Arabic manuscript pages rarely exhibit a consistent baseline: words within a single line float at varying angles and heights, displaced by diacritics, sublinear descenders, and the calligrapher's deliberate compositional choices. At the level of scribal tradition, the challenge is further complicated—the same word written in Naskh, Thuluth, Maghrebi, or Ruq'a hands may share almost no visual geometry, meaning that any robust recognition system must account for inter-style variation before it can meaningfully generalize across corpora. These compounding factors make character-level and line-level recognition paradigms ill-suited as primary analytical units for Arabic manuscripts. We argue that the word is the natural unit of analysis, as it is the smallest graphically self-contained entity that encodes both lexical identity and scribal style simultaneously. In this paper, we present a word-level graphometric framework that addresses these challenges through geometric vectorization: each word image is binarized, skeletonized into a stroke graph, and characterized by a feature vector encoding structural complexity, stroke-width variation, path geometry, and ink density. We describe a quality-aware cluster profiling system that builds statistical fingerprints for each word type across manuscript instances, and a geometry-only classifier that assigns new, untranscribed word images to existing clusters without relying on prior lexical knowledge. We present the results of experiments conducted on a segmented Arabic manuscript corpus, demonstrating that computational paleography—the systematic, quantitative study of scribal hands—can expose the within-corpus variation that HTR models must learn to navigate, and that the graphometric profiles produced by our framework offer a principled basis for building training corpora that are representative of scribal diversity.

Presentation materials

There are no materials yet.