29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

Agree to Disagree: Intra-Annotator Agreement in Semi-Automatic Vocabulary Building

2 Oct 2026, 16:30
30m
02 | Museumszimmer (02 | Museumszimmer, OeAW Main Seat, 2nd floor)

02 | Museumszimmer

02 | Museumszimmer, OeAW Main Seat, 2nd floor

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speakers

František Kovařík (Lexical Computing) Marek Blahuš (Lexical Computing) Miloš Jakubíček (Lexical Computing) Vojtěch Kovář (Lexical Computing)

Description

Intra-annotator agreement is the rate at which an annotator makes the same choices in the same situation. This paper examines the annotation process in a semi-automatic dictionary-making project, Czech Dictionary Express. In the vocabulary-building process, its annotators went through 100,000 headwords from a corpus frequency wordlist, marking each as incorrect or (partially) correct. The aim of the Dictionary Express projects is a high dictionary-making speed, so the annotators weren’t given the context of the headwords.

In this paper, we examine the shift in annotators’ opinions on headwords before and after context was provided. We argue against providing no context of the headwords, because, unlike the (partially) accepted words, which can still be cut from the vocabulary at later stages, the rejected words cannot be put back in, and this is the case of some correct words that have not been recognised by the annotators without context. We show that the context-based annotation process doesn’t take significantly longer. We also take a brief look at a study about using LLMs for the same annotation process.

Presentation materials

There are no materials yet.