Speakers
Description
Intra-annotator agreement is the rate at which an annotator makes the same choices in the same situation. This paper examines the annotation process in a semi-automatic dictionary-making project, Czech Dictionary Express. In the vocabulary-building process, its annotators went through 100,000 headwords from a corpus frequency wordlist, marking each as incorrect or (partially) correct. The aim of the Dictionary Express projects is a high dictionary-making speed, so the annotators weren’t given the context of the headwords.
In this paper, we examine the shift in annotators’ opinions on headwords before and after context was provided. We argue against providing no context of the headwords, because, unlike the (partially) accepted words, which can still be cut from the vocabulary at later stages, the rejected words cannot be put back in, and this is the case of some correct words that have not been recognised by the annotators without context. We show that the context-based annotation process doesn’t take significantly longer. We also take a brief look at a study about using LLMs for the same annotation process.