Speakers
Description
A common challenge for language learners is finding appropriate L2 equivalents for L1 words. The difficulty is compounded when the L1 word is polysemous, because different senses typically call for different equivalents. When selecting an L2 equivalent for a polysemous L1 word, learners therefore face a twofold task: identifying the intended sense of the L1 word and choosing an L2 equivalent that fits the context.
Traditionally, learners have dealt with this problem by turning to bilingual dictionaries. High-quality bilingual dictionaries may provide glosses, labels, collocations, examples, and other information that helps users select the most appropriate equivalent from among those listed and then use it in context. Yet, as previous studies have shown, the mere availability of information does not guarantee that users will find it or use it effectively (Lew & Tokarek, 2010; Dziemianko, 2012).
With the arrival of ChatGPT and other large language models (LLMs), learners may now turn to AI tools instead of consulting dictionaries. A recent study by Ptasznik and Lew (2025) indicates that at least some learners already do so, although their preferences appear to be task-dependent. Learners’ willingness to use such tools is perhaps unsurprising, given evidence from user studies showing that LLMs can match and sometimes outperform dictionaries in receptive and productive tasks (Rees & Lew, 2023; Ptasznik, Wolfer & Lew, 2024). Yet expert evaluations of entries generated by LLMs have been mixed. An often-noted problem is false polysemy, which occurs when “the system enumerates multiple senses, with different definitions, in cases where there is really only one” (Jakubíček & Rundell, 2023, p. 525). A recent user study found that false polysemy can have a detrimental effect on dictionary users, increasing the time needed to select a definition and lowering users’ confidence in the definition chosen (Michta & Frankenberg-Garcia, 2025).
To the best of our knowledge, no study to date has specifically examined learners’ performance in selecting L2 equivalents for polysemous L1 words when consulting a bilingual dictionary versus an LLM, a task that is common yet under-researched. We investigate learners’ performance using the following measures: response accuracy, response time, and delayed recall of the target equivalents one week later. In addition, we examine learners’ confidence in the equivalents they selected and changes in their tool preferences after the task.
A total of 105 L1 Polish speakers, all enrolled in an English philology programme, took part in the study. Before completing the task, participants indicated which of the two tools, ChatGPT or the online Polish-English PONS dictionary, they would prefer to use for the task and rated how useful they expected each tool to be. They were then randomly assigned to one of the two conditions. Participants completed an online task involving 20 sentences presented in random order. Each item consisted of an English sentence with a gap and a Polish polysemous word in parentheses. Participants first indicated whether they were able to supply an English equivalent without consulting the assigned tool, and if so, provided it. Next, they consulted the assigned tool and filled in the gap using the information obtained during consultation; response time was recorded. After each response, they rated how confident they were that they had selected an appropriate equivalent. One week later, 70 participants completed a post-test in which the same Polish words and senses were presented to them in new contexts. For each item, they provided an English equivalent and rated their confidence. Participants also rated the overall helpfulness of the assigned tool after both the main task and the post-test.
The results suggest that the dictionary group provided more accurate responses, required less time to supply an equivalent, and reported higher confidence in their choices. After the task, participants more often selected the dictionary than ChatGPT as the tool they would choose for a similar task in the future. Full analyses will be presented at the conference.