Speakers
Description
Objective
This study evaluates whether word-prevalence data can improve the selection and review of dictionary headword candidates. It focuses on Slovenian lemmata that combine higher corpus frequency with lower prevalence, because this mismatch may signal problems that frequency alone does not reveal.
Background and data
Word prevalence is the percentage of a population that knows a word (Keuleers et al., 2015). Large online lexical-decision studies have produced prevalence norms for Dutch, English, Spanish, Catalan, and Italian, based on tens of thousands of words and large participant samples (Keuleers et al., 2015; Brysbaert et al., 2019; Aguasvivas et al., 2018, 2020; Guasch et al., 2023; Amenta et al., 2025). These studies show that very frequent words are generally widely known, whereas words at lower frequencies vary considerably in prevalence. Lew & Wolfer (2024) represents a rare example of applying word-prevalence data to the field of lexicography.
The ongoing Slovenian megastudy has so far collected one hundred YES/NO responses for each of 35,000 words from 39,063 participants in 48,659 sessions; the complete list contains 79,413 words. In each session, participants classify 120 letter strings as Slovenian words or nonwords. Among the 120 letter strings given in each session, 84 were real Slovenian words and 36 were not, The list was compiled from existing dictionaries (Perdih et al., 2025), which provide verified words (as opposed to potentially unverified corpus lemmata) and to serve as a motivational element at the end of a session – lists of successfully and unsuccessfully recognized words linked to the Fran Slovenian dictionary portal (https://www.fran.si) are provided to participants at the end of each session. Consequently, frequent corpus lemmata absent from the dictionaries are outside the present analysis.
Selection and classification
We selected the 50 highest-frequency lemmata in the Gigafida reference corpus of Slovenian (Krek et al., 2020) whose prevalence values were below zero, meaning that fewer than half of the participants recognized them. The least frequent item, masen ‘pertaining to weight’ occurred 832 times (0.62 per million tokens). We then classified the likely sources of the frequency–prevalence mismatch.
Results
Lemmatization or tokenization errors account for 15 items (30%) and often inflate corpus frequencies. A further 17 items have dictionary forms that are uncommon in actual use: 13 adjectives, 3 verbs, and 1 pronoun. For example, the adjectival headword kaven ‘pertaining to coffee’ was recognized by 45.23% of respondents (prevalence value -0.118) despite a frequency of 4,094 (3.07 per million tokens); its more frequent definite form kavni would probably be easier to recognize. Similarly, the infinitives poiti ‘to run out’ and onemoči ‘to tire, to languish’ are much less common than their inflected forms.
Other mismatches reflect restricted use or ambiguous status. Eight items are terms, including obtežba ‘load, weight’; eight may be interpreted as foreign, although some are Slovenian loanwords, such as go ‘Japanese game’ and song ‘a song genre’, while others are now archaic or obsolete and only occur in the corpus as parts of foreign-language texts (il ‘clay’, but also Italian definite article). The two interjections hi and jah may be difficult to recognize as Slovenian words without context, and hi may also be associated with the English greeting. Two items are homographs of proper names. Some categories overlap: for example, Li may be lemmatized as li, while li may also result from incorrect tokenization of a masculine plural past-participle ending.
Implications for headword-list development
Word prevalence is useful as a diagnostic measure for high-frequency headword candidates. When combined with corpus frequency, it helps identify lemmatization and tokenization errors, specialized terms, foreign-language material, and items whose recognition depends strongly on context. These findings can support both the addition of new entries and the review or removal of existing entries: low prevalence may indicate corpus noise, opens a question of headword form user-friendliness, or suggests that a stylistic or terminological label may be required rather than simple exclusion of the headword.