Speakers
Description
Recent empirical studies suggest that ChatGPT can assist users with a range of receptive and productive language tasks (Rees & Lew, 2023; Ptasznik, Wolfer & Lew, 2024). However, users consult dictionaries in many situations not covered in previous studies comparing ChatGPT with dictionaries, and ChatGPT’s usefulness is likely to vary depending on the task. This is plausible because ChatGPT does not appear to perform equally well when asked to produce different types of dictionary information: it has been found to produce definitions that are “practically indistinguishable” from those found in the COBUILD dictionary (Lew, 2023, p. 8), while ChatGPT-generated examples tend to be repetitive, unimaginative and unnatural (Jakubíček & Rundell, 2023; Lew, 2023). The present study focuses on a situation often faced by language learners but still underexplored in empirical comparisons of ChatGPT and dictionaries: resolving polysemy in L2 reception.
Polysemous L2 words pose a considerable challenge for language learners because they can be deceptively familiar: learners may assume they understand a word because they know one of its senses, although the context requires another. If they do consult a dictionary, they still have to determine which sense in the entry matches the context. As earlier studies have shown, many learners struggle to do so (Dziemianko, 2019; Kamiński, 2025).
Learners may now also use an LLM to help them understand which sense of a polysemous word is intended in context. LLMs may be useful here, but they may also present learners with false polysemy by proposing senses that are not clearly distinct from one another (Jakubíček & Rundell, 2023). This matters because false polysemy has been shown to make sense selection more time-consuming and to lower users’ confidence in the sense they selected (Michta & Frankenberg-Garcia, 2025).
In this study, we examine how successfully learners identify the contextually appropriate senses of polysemous L2 items. To this end, we compare the effectiveness of ChatGPT with that of a bilingual dictionary in an equivalent-selection task and with that of a monolingual learner’s dictionary in a definition-selection task. In addition to accuracy, we investigate learners’ confidence in their choices, time-on-task, and order effects.
The participants were 131 L1 Polish secondary-school students whose proficiency in English ranged from B1 to B2. Fifteen polysemous English nouns were selected as target items. Each item appeared in a single sentence taken from Sketch Engine for Language Learning, published books, or dictionaries other than those used by participants in the study. Some sentences were minimally edited to remove non-essential information. The task was administered via a worksheet in which sentences were presented in random order.
The study consisted of a pretest and a main test. In the pretest, all participants saw 15 sentences and attempted to translate or paraphrase each target item without tool support. In the main test, participants were assigned to one of four groups and used their phones to complete one of two tasks. In the equivalent-selection task, Group A used Diki.pl, an online English-Polish dictionary, while Group B used ChatGPT. Participants selected a Polish equivalent of the target word and rated their confidence in selecting an appropriate equivalent on a 1–5 Likert scale. In the definition-selection task, Group C used the online Oxford Advanced Learner’s Dictionary, while Group D used ChatGPT. Participants indicated their selected definition by copying three words from it and rated their confidence in their choice on a 1–5 Likert scale. They were then asked to provide a Polish equivalent of the English target word. Total time-on-task was recorded for the main test.
The data were analysed separately for each of the two tasks using logistic mixed-effects models for accuracy and ordinal mixed-effects models for confidence. Preliminary results suggest that, in the equivalent-selection task, Diki.pl users outperformed ChatGPT users and reported higher confidence. Order effects were also observed in this task. Full analyses will be presented at the conference.