Speakers
Description
This paper investigates the potential of large language models (LLMs) as assistants with corpus analysis in lexicography, focusing on the assignment of register labels in the EKI Combined Dictionary (CombiDic). Inconsistencies in register labels become evident across synonym sets in CombiDic, which is the result of different lexicographers working on words at different times. We examine how LLMs from Anthropic, Google, and OpenAI handle the challenge of distinguishing between offensive and colloquial usage, which pose a challenge in lexicographic practice. Using a evaluation dataset of 297 words evaluated by five native Estonian-speaking annotators and three LLMs queried via API, we find that LLMs are quite stable tools for register categorisation and that corpus context significantly improves their performance. Depending on the LLM, 79.8โ88.6% of suggestions were deemed acceptable for lexicographic use, although since LLMs may be trained to detect offensive language, they tended to select the OFFENSIVE label more often than human annotators. Nevertheless, Gemini 3.1 Pro performed best overall. The experiment already enabled corrections to CombiDic, demonstrating practical value. While full systematisation has not yet been achieved โ the broader goal is labels grounded in usage data โ LLMs represent a promising and time-efficient aid in semi-automatic lexicographic workflows.