Speaker
Description
The aim of this paper is to explore methods to utilize the linguistic knowledge accumulated in language models to evaluate corpus sentences as potential examples in dictionary entries. We examine the TPEX (Typical Example) measure in three variants in the task of assessing the typicality of a sentence based on how probable the appearance of the headword in the given context is, as per the selected language model. Then we present an extension, the UGTPEX (Usage-Guided Typical Example) measure, in six variants, where the contextual typicality is supplemented by data reflecting the similarity of the headword’s contextualized embedding to representations seen in known-good (dictionary-derived, sense-specific) examples. The combination of the two components, i.e. the contextual typicality and the embedding-based similarity, results in a score that is sensitive to both frequent usage patterns and the usage categories (sense categories) established by the lexicographer. We will examine the applicability of these measures in case studies involving five headwords using dictionary example sentences and corpus samples. The proposed measures are for ordering concordance lines, in combination with other ranking methods, as part of the lexicographic workflow.