Speaker
Description
At the first SCOOP meeting in June 2025 Aaron Hershkowitz and Nicholas Howe presented the results of their 2024-25 NEH-funded project to evaluate various approaches to automated text recognition. Two main text recognition approaches were explored: traditional sequential Optical Character Recognition (OCR), and Scene Text Detection. Out of the box OCR engines like Kraken were found not to handle squeeze images well, even after processing to convert them to “pseudo-ink”. Sequential OCR performed much better when retrained on squeeze images, but still lagged behind the performance of Scene Text Detection models like YOLO and DBNet. The most successful approach turned out to be using Scene Text Detectors to find characters, then assembling lines from those characters for sequential approaches like Convolutional Neural Networks. In 2025-26 Hershkowitz and Howe ran a contest through the International Conference in Document Analysis and Recognition (ICDAR) to attract outside researchers with fresh approaches to the problem. That contest concluded at the end of April, and initial analysis of the results has been conducted. As part of WG2, Aaron Hershkowitz will present the team’s conclusions from the ICDAR contest, and discuss potential next steps in terms of improving and utilizing squeeze ATR.