Speaker
Description
Calfa is developing general and custom AI models for researchers and cultural heritage professionals, dedicated to recognize the text and analyze documents in non-western languages, such as Arabic, Armenian, Georgian, Syriac, Chinese, Greek. Calfa also offers many open-source datasets for Arabic or Armenian languages.
We recently developed a promising hybrid approach, combining our specific OCR/HTR models with customized VLMs.
This approach allows for the processing of a very wide variety of languages and scripts with excellent accuracy rates on both handwritten documents and printed archives.
Researchers can now obtain a specialized AI solution tailored to their corpus, enabling not only OCR/HTR processing but also analysis and understanding of the documents layout according to their project specifications.
Several tools are offered free of charge online or for download by Calfa so that researchers can try these solutions on their corpus.