Speakers
Description
This paper investigates whether large language models can perform semantic clustering and categorisation of constructional collexemes to support the analysis of constructional meaning and the organisation of collexemes within constructicon entries. As a case study, we examine the collexemes of the Estonian Nominal Quantifier Construction, identified from lexicographic and corpus data. Using OpenAI’s ChatGPT-5.4, we conducted an informed free sorting task and a closed sorting task with five input types, consisting either in bare lemmas or lemmas accompanied by different types of context: corpus phrases, corpus sentences, dictionary definitions, and dictionary examples. Model outputs were evaluated against a human-created gold standard using Adjusted Mutual Information, Adjusted Rand Index, and a label quality rating scheme. A scaled closed sorting experiment was also conducted. The results of the free sorting tasks approached human agreement levels, with dictionary definitions yielding the most similar clustering and corpus sentence input the most acceptable labels. The results of the closed sorting task and the scaling experiment demonstrated that a bare list of lemmas provided sufficient input for assigning collexemes to predefined semantic categories.