Speaker
Osama Eshera
(University of Maryland)
Description
We study word-level text-to-image mapping in Arabic-script manuscripts: given a line image and its transcription, locate each word. At corpus scale, accurate word geometry opens new possibilities for digital paleographical analysis while improving ATR both upstream and downstream. This mapping is learned from 25,000 line crops drawn from 87 manuscripts, supervising word geometry only through the transcription, and evaluated against a human-drawn gold-standard set of 1,276 words. We compare against a blind control that never sees the transcription.