29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

Are Bears More Offensive than Dogs When Drunk? Using LLMs to Systematize the Labelling of Estonian Synonyms

3 Oct 2026, 09:00
30m
01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)

01 | Johannessaal

01 | Johannessaal, OeAW Main Seat, 1st floor

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speakers

Lydia Risberg (Estonian Language Institute, University of Tartu) Kristina Koppel (Estonian Language Institute) Margit Langemets (Estonian Language Institute) Hanna Maask (Estonian Language Institute) Esta Prangel (Estonian Language Institute) Maria Tuulik (Estonian Language Institute) Silver Vapper (Estonian Language Institute)

Description

This paper investigates the potential of large language models (LLMs) as assistants with corpus analysis in lexicography, focusing on the assignment of register labels in the EKI Combined Dictionary (CombiDic). Inconsistencies in register labels become evident across synonym sets in CombiDic, which is the result of different lexicographers working on words at different times. We examine how LLMs from Anthropic, Google, and OpenAI handle the challenge of distinguishing between offensive and colloquial usage, which pose a challenge in lexicographic practice. Using a evaluation dataset of 297 words evaluated by five native Estonian-speaking annotators and three LLMs queried via API, we find that LLMs are quite stable tools for register categorisation and that corpus context significantly improves their performance. Depending on the LLM, 79.8โ€“88.6% of suggestions were deemed acceptable for lexicographic use, although since LLMs may be trained to detect offensive language, they tended to select the OFFENSIVE label more often than human annotators. Nevertheless, Gemini 3.1 Pro performed best overall. The experiment already enabled corrections to CombiDic, demonstrating practical value. While full systematisation has not yet been achieved โ€“ the broader goal is labels grounded in usage data โ€“ LLMs represent a promising and time-efficient aid in semi-automatic lexicographic workflows.

Presentation materials

There are no materials yet.