Speakers
Description
A distinguishing feature of English dictionaries for learners is that most entries include example sentences or phrases that have been copied or adapted from corpora. Their function is to reinforce meaning by showing how words have been used in context, prioritising typical grammar patterns and lexical collocations (Fox 1987; Frankenberg-Garcia 2014; Frankenberg-Garcia, Rees and Lew 2021). While early corpus-based lexicographers had to scan concordances line by line to select good dictionary examples (Krishnamurthy 1987), nowadays tools like Word Sketches (Kilgarriff et al. 2014) and GDEX (Kilgarriff et al. 2008) shorten the time it takes for lexicographers to find suitable corpus examples. Despite these welcome developments, it is now arguably easier and faster to generate dictionary examples using LLMs. However, in early experiments evaluating AI-generated examples, experts found them to be redundant, unimaginative and inauthentic (Lew 2023, Jakubíček & Rundell 2023), although improvements could be seen after some prompt fine-tuning (Lew 2023). In this study, we wanted to find out how language learners rated AI-generated and corpus-based dictionary examples, to come to a better understanding of how they differ, and to consider the practical implications for lexicography.
Over 200 university students pursuing various degrees (language and non-language) participated in an online task involving 15 monosemous lexical items (from to different part-of-speech categories) that were likely to be unknown. For each lexical item, the participants were presented with the definitions from the Reverso English Dictionary (AI-generated) and the Oxford Advanced Learner’s Dictionary (corpus-based) followed by one example from each dictionary. As most entries contained more than one example, the ones presented to each student were randomly selected from the totality of examples available. The participants were asked to rate the 15 AI-generated and the 15 corpus-based examples they saw on a 1-5 scale (poor to excellent). The order of lexical items, definitions and examples shown to each student were all randomized. At the end of the task, the participants were asked to explain what made them give examples high or low ratings. They then responded to a few demographic questions and completed a standardized vocabulary test (LexTALE, Lemhöfer & Broersma 2012).
Prior to the data collection, the authors of this study conducted a data-driven appraisal of all 63 examples in the dataset. By comparing interpretive coding criteria and discussing divergences, we collaboratively negotiated a systematic analytical framework for describing dictionary examples. Using this framework, we noted a number of significant differences between the AI and the corpus-based examples (e.g., in the corpus-based examples, there was significantly more variability in the number of words and syntax used). However, not all differences were found to be significant (e.g., the presence of contextual cues about meaning).
With regard to the student ratings, overall preliminary findings indicate that there was no significant difference between examples in each resource. However, when only complete sentences were taken into consideration, the Oxford examples were rated significantly more positively.
The full results of the study, including triangulation of our analytical framework for describing dictionary examples along with user ratings and student demographics will be presented at the conference. We believe the study will help to shed new light on how corpus-based and AI-generated dictionary examples can be refined in the future.