Speakers
Description
This paper presents an experiment on generating definitions for Slovene headwords using four Large Language Models (LLMs): two commercial (Gemini and GPT) and two open-source (Gemma and the Slovenian model GaMS). Each model was provided with contextual data, including semantic indicators, collocations, and examples. We tested a zero-shot approach and two few-shot approaches using sample definitions from the Digital Dictionary Database of Slovene (DDDS) and the New Dictionary of Standard Slovene. The results show that while all models produce high-quality definitions, Gemini performs best across all word classes, particularly when utilizing DDDS sample definitions. The second part of the study evaluated LLMs as judges of definition quality, testing Gemini, GPT, and Claude-Sonnet. Gemini again achieved the highest agreement with human lexicographers. Notably, even when LLMs selected a different definition than the experts, their choices were usually viable. These results demonstrate the considerable potential of LLMs in lexicography, and we plan to deploy this pipeline to generate definitions for a much larger set of Slovenian headwords.