29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

Digitization of Adam Patačić's Manuscript Dictionary

1 Oct 2026, 10:20
20m
01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)

01 | Johannessaal

01 | Johannessaal, OeAW Main Seat, 1st floor

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speakers

Željko Jozić (Institute for the Croatian Language) Marijana Horvat (Institute for the Croatian Language) Martina Horvat (Institute for the Croatian Language)

Description

Adam Patačić is a Croatian lexicographer (1716–1784), Archbishop of Kalocsa and President of the Royal Council of the University of Buda and a promoter of music with his orchestra and choir. His extensive Dictionarium latino-illyricum et germanicum remained in manuscript, written after the example of Nomenclator omnium rerum by Hadrian Junius (Samardžija, 2019, p. 63).

The mentioned dictionary is a completed manuscript work written in ink, the pages are numbered, and on the inside title page the author is signed as D. ADAMUS L. B. PATACHICH De Zajezda. Although the manuscript does not contain information on the title page indicating the year the work dates from, the data relevant to the dating of the dictionary can be read from the preface, as well as from the title page. According to the data on the title page, the author's name is accompanied by a list of his titles and honors, some of which he achieved only towards the end of his life. In support of this dating, and according to the numbered third page of the dictionary, it can be concluded that the work was bound and completed in the late seventies or early eighties of the 18th century. According to the data from the preface, it can be stated with certainty that the dictionary was created in the period from 1772 to 1779 (Jonke, 1949, pp. 95–96). Patačić's dictionary was written in an era that, together with the previous century, is considered the golden age in the history of the Kajkavian Croatian literary language. It was at this time that major Kajkavian lexicographical works were created – Belostenec's, Jambrešić's and Patačić's dictionaries, as well as the first orthography and the first manuscript Kajkavian grammar (Štebih Golub, 2013, p. 258).

The dictionary is conceptually structured with a macrostructure of 13 thematic areas and several sub-areas, with encyclopedic interpretations in Latin. The headword side is Latin, and the explanations are given in Croatian, i.e. in the Kajkavian literary language, while the equivalent is also given in German. The dictionary contains a rich treasury of Kajkavian literary lexis (for example, kinship, zoological and botanical terminology), as well as the Kajkavian literary language of the 18th century.

The preface contains XIII pages, while the content consists of V pages. The dictionary material covers 1034 numbered pages. The dictionary is divided into four major parts. The first part entitled De Deo Spiritibus, Coelo, Elementi set Homine comprises 13 chapters on 220 pages and provides general information about the extralinguistic reality. The second part is entitled De iis, quae ad Hominis vitam tum recte agendam, tum tuendam pertinent: sive Sacrorum Praesides, Civilia et Militaria, and contains 15 chapters (pp. 221–415) and covers political and military duties, literature, music, and art. The third part is entitled Oeconomica and contains 14 chapters (pp. 416–734) and relates to the economy. The fourth part Technica (pp. 735–1034) consists of 16 chapters and relates to technical sciences.

Although Patačić's dictionary was not used as a corpus in the preparation of the Academy's historical multi-volume Dictionary of the Croatian or Serbian Language (1880–1976), it served as a corpus in the preparation of the Dictionary of the Croatian Kajkavian Literary Language (1984–present, kajkavski.hr) where, in structuring the dictionary entry, confirmations from well-known Kajkavian dictionaries of the 17th and 18th centuries are cited, along with Patačić's manuscript dictionary (Brlobaš and Horvat, 2019, p. 146).

The largest study of the dictionary is Ljudevit Jonke's Dikcionar Adama Patačića (1949) and the more recent one is Marijana Horvat's research on musical terminology in Croatian pre-revival dictionaries (Horvat, 2018, pp. 437–454).

Due to its exceptional importance for the Croatian language in general and the Kajkavian literary language, at the beginning of 2026, at the initiative of the director of the Institute for the Croatian Language, Željko Jozić, and with the help of collaborators on the Dictionary of the Croatian Kajkavian Literary Language project, and with other colleagues from the Department of the History of the Croatian Language, a project was launched to study Patačić's manuscript dictionary. The transliteration of the work has begun with the aim of further studying the extensive and the last historically extremely important Kajkavian dictionary. In the presentation, we will talk about the principles of digitization and the problems we encounter.

We are digitising Adam Patačić’s handwritten Dikcionar, a manuscript exceeding 1,000 pages, by combining philological transcription and handwritten text recognition (HTR) in Transkribus. As initial Ground Truth, we reused the manuscript text transcribed by Ljudevit Jonke in his study ‘Dikcionar’ Adama Patačića: studija iz hrvatske kajkavske leksikografije. This transcription corresponds to 25 manuscript pages of the Dikcionar. The text was imported into Transkribus and manually aligned with the corresponding manuscript images and baselines, yielding a first controlled Ground Truth set. Our very first training attempt—based on an incomplete subset of this Jonke-derived material and without an appropriate base model—produced poor recognition performance (Transkribus-reported CER 58.61%), showing that substantial script-specific adaptation was still needed. We therefore introduced transfer learning and retrained the model using the publicly available base model “Cosimo Bartoli’s Italian Humanistic and Cursive Scripts (1562–1572)” together with the Jonke-based Ground Truth. This markedly improved recognition, lowering the Transkribus CER to 11.52%. Since our research use case tolerates several predictable early-modern graphemic ambiguities, we additionally evaluated performance under a normalisation that ignores case and treats u/v, s/ſ, and æ/ae as equivalent. Under these conditions, an independent CER calculation yielded 6.01%, which we consider an excellent and practically usable starting point for large-scale digitisation. The next phase will proceed iteratively: a team of researchers will expand Ground Truth by correcting model output on newly processed pages, while the model will be periodically retrained to further improve robustness across the manuscript and to support subsequent linguistic and lexicographic analysis of the Dikcionar. This digitisation also marks an exceptional pilot step within the MONOGRAF project (Development and Implementation of a Model for Normalising the Orthography of Old Texts Printed in Latin Script), which otherwise focuses on printed texts but has, on a trial basis, extended its workflow to a handwritten manuscript in this case.

Presentation materials

There are no materials yet.