29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

From Field Recordings to an Online Dictionary: A Model for Processing Štokavian Dialect Material

30 Sept 2026, 14:00
1h 30m
Aula (Entrance Hall) & 1st floor (00/01 | Aula (Entrance Hall) & 1st floor, OeAW Main Seat)

Aula (Entrance Hall) & 1st floor

00/01 | Aula (Entrance Hall) & 1st floor, OeAW Main Seat

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speaker

Perina Vukša Nahod (Institute for the Croatian Language)

Description

This poster introduces a multi-modular model for developing an online Croatian dialect dictionary. Although Croatian dialect lexicography advanced significantly in the 20th century (Lisac 2006) and continues to develop in the 21st century, online Croatian dialect dictionaries remain uncommon. Most available resources are digital reproductions of print dictionaries or amateur compilations that lack adherence to established scientific principles of dialectological analysis.

The Croatian language exhibits a threefold dialectal division. However, no existing dictionary encompasses material from all three Croatian dialects, nor does any cover more than one dialect (Magaš 2026). Additionally, there is no Croatian dialect dictionary with a multi-modular structure similar to that developed for the Croatian standard language (Hudeček & Mihaljević 2024). Despite the predominance of Štokavian speakers in Croatia, dictionaries dedicated to individual Štokavian local varieties are limited.

To address this gap, the project focuses on Štokavian dialects and aims to develop a dialect dictionary that integrates multiple corpora, various levels of linguistic processing, and diverse target age groups. This dictionary is intended to serve as a model for processing the lexical material of other Croatian local varieties, groups of varieties, and dialects, especially those recognized as intangible cultural heritage.

Within the project Rijekom neretvanskih riječi – od kuće do škole (“Along the River of Neretva Words – From Home to School”), an online dialectological Dictionary of Neretva Dialects is under development. This dictionary includes varieties from two Štokavian dialects: Neo-Štokavian Ijekavian and Neo-Štokavian Ikavian. It represents the first attempt to present material from two Croatian dialects simultaneously and to implement transparent, functional lexicographic solutions. The dictionary features modules designed for pupils from the first grade of primary school through the fourth grade of secondary school. It is also the first Croatian online dialect dictionary to incorporate multiple fields, including headword, linguistic annotation (expandable), meaning, sentence example, idioms, audio recording, video recording, drawing, additional notes, and references.

Because dialect dictionaries frequently omit forms essential for precise phonological and morphological description (Kapović 2008), this project emphasizes the recording of grammatical forms relevant to determining the accentual paradigms of individual lexemes. Dictionary entries are compiled using TshwaneLex software, with three distinct modules developed:

  1. a module for grades 1–4 of primary school, which includes accented lemmas, meanings, drawings, additional notes, audio recordings, and video recordings;
  2. a module for grades 5–8 of primary school, which includes accented lemmas, grammatical forms, meanings, and audio recordings;
  3. a secondary-school module, which includes accented lemmas, grammatical forms, extended meanings, idioms, accented sentence examples, and audio recordings.

The core corpus of the dictionary comprises lexical items collected through a dialect survey questionnaire that covers a range of thematic domains, including family relations, fruits and vegetables, agricultural work, life by the river, children's games, recipes, and customs. To prevent restriction to predefined thematic fields, the corpus is systematically expanded with lexical items extracted from recordings of informants’ spontaneous speech. This thematic approach to data collection is particularly appropriate for primary school pupils in grades 1–4 and serves as a basis for further corpus expansion. Pupils in grades 5–8 actively participate in collecting and processing additional material, thereby contributing to both the educational and research aspects of the project.

Because the corpus includes material from two dialects, the dictionary records different phonological and morphological variants of lexical items. Each variant is accompanied by an abbreviation of the settlement where it was documented, enabling precise representation of local variation and systematic lexicographic treatment.

Secondary school students collect phraseological material using a targeted questionnaire developed from previous research on Neo-Štokavian dialects and through the investigation of selected concepts, such as strength, beauty, and stupidity. These concepts reveal conventional perceptions, stereotypes, and value systems characteristic of the speakers of a particular community. This approach enables systematic documentation of phraseological units and provides insights into the cultural and cognitive dimensions of dialectal language use.

By presenting these dictionary structures, the project offers insight into a contemporary and standardized approach to processing Croatian dialect material, ensuring both accessibility and searchability.

Presentation materials

There are no materials yet.