Speaker
Description
In lexica of natural languages, polysemy is the rule rather than the exception. Understanding why it is so prevalent—given that the presence of multiple senses for a single form, from a purely semiotic point of view, constitutes ambiguity—and how it emerges is central to lexicology. Addressing this question, of course, requires proper operationalization of polysemy in the first place. A robust, straight forward, and commonly used way of measuring the polysemy of a word is simply counting the number of sense entries in a lexicographically curated resource. Such an approach, however, neglects the relative frequencies of a word’s senses as well as how similar senses are to each other. By integrating lexicographic and corpus data, I will discuss frequency and similarity aware measures of polysemy, originating from quantitative ecology (Leinster & Cobbold, 2012), and how they relate to ratings of subjectively perceived polysemy, gathered through crowdsourcing efforts. I will then discuss the roles that frequency and acquisition play in the emergence of polysemy (Baumann & Hartmann, 2026). Finally, I will zoom in on one specific dimension of meaning: lexical sentiment. I show that variation in human sentiment annotations correlates well with how emotional polysemy is subjectively perceived, but that purely NLP based approaches in fact struggle with appropriately capturing subjective emotional ambiguity. I take this to stress the relevance of ‘the human in the loop’ also in lexical and lexicographic research.