29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

Exploring Large Language Models in Word Sense Disambiguation: A Study of Corpus Examples for Perception Verbs in Learners' Dictionaries

2 Oct 2026, 17:00
30m
01 | Sitzungssaal (01 | Sitzungssaal, OeAW Main Seat, 1st floor)

01 | Sitzungssaal

01 | Sitzungssaal, OeAW Main Seat, 1st floor

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speaker

Sylwia Wojciechowska (Adam Mickiewicz University, Poznań)

Description

Recent developments in Large Language Models (LLMs) have opened new possibilities for supporting lexicographic work. Studies suggest that LLMs can assist in tasks such as generating definitions, drafting example sentences, and assigning examples to dictionary senses (Lew, 2023; Jakubíček and Rundell, 2023). At the same time, the status of word senses remains one of the most debated issues in lexicography, with scholars highlighting the fluidity and context-dependence of meaning. Traditional lexicographic practice involves abstracting discrete senses from potentially unlimited numbers of contextualised corpus attestations. This process is particularly challenging in the case of highly frequent and polysemous lexical items, where meanings often overlap and form complex semantic networks. Dictionaries may adopt different approaches to sense division, “lumping” or “splitting” senses, which leads to considerable variation in sense inventories (Bond et al., 2024).

Nowadays, apart from assigning corpus-based examples to particular senses, online learners’ dictionaries increasingly include large collections of automatically extracted corpus examples that are not explicitly linked to individual senses. While these collections provide valuable data, their pedagogical usefulness remains limited, as examples are not systematically organised. Recent advances in generative AI suggest that LLMs may help bridge this gap by analysing contextual meaning and aligning examples with sense definitions. Research indicates that such models are capable of handling metaphor, semantic relatedness, and graded meaning variation (Lin et al., 2024; Bond et al., 2024). Nevertheless, their performance in dealing with fine-grained sense distinctions and competing lexicographic models of word sense disambiguation remains underexplored.

The present study examines the potential of LLMs as tools for analysing sense granularity taking a cross-dictionary perspective. This paper investigates the ability of ChatGPT (GPT-5.2) to interpret corpus examples of English perception verbs and align them with the existing sense inventories. Both sense divisions and sections of unassigned corpus examples were drawn from two major online monolingual learners’ dictionaries: Cambridge Advanced Learner’s Dictionary and Longman Dictionary of Contemporary English. As the number of unassigned corpus examples for each examined entry differs between the dictionaries, the first twenty examples per verb were analysed. The dataset includes eleven high-frequency and highly polysemous perception verbs: see, hear, touch, taste, smell, look, watch, observe, notice, listen and feel. These verbs display systematic extensions from the physical domain into cognitive and evaluative domains, hence are suitable for examining sense granularity and semantic overlap. Given their systematic polysemy, the study examines how ChatGPT handles semantic shifts from physical to abstract and metaphorical meanings.

Presentation materials

There are no materials yet.