Speakers
Description
As early as the eighteenth century, dictionary compilation has been perceived as a labor intensive and time-consuming endeavor, often characterized as meticulous and repetitive work (Lew, 2023). The creation of lexicographic articles, in particular, constitutes a complex process requiring the careful coordination of multiple parameters, including the target user group, users’ native language, the type and purpose of the dictionary, as well as the design of its macrostructure and the organization of its microstructure (Hartmann, 2001). In pedagogical lexicography, this complexity is further intensified, as definitions and supporting information must also account for learners’ age, language proficiency, learning objectives, and the conditions under which the dictionary is consulted, such as reception- or production-oriented use (Tarp, 2011). Against this background, contemporary lexicography has increasingly emphasized the need to reduce the time and cost of dictionary making, with automation widely regarded as a key factor in achieving this goal (Jakubíček & Rundell, 2023; Rundell, 2024). Recent advances in artificial intelligence, particularly over the past five years, have therefore prompted lexicographers to explore new solutions for both dictionary compilation and use. Notably, the emergence of Large Language Models (LLMs) has intensified interest in the potential role of chatbots in supporting—and possibly reshaping—various stages of the lexicographic process, including the generation and mediation of lexicographic articles (Phoodai & Rikk, 2023; Rees & Lew, 2023; Lew et al., 2024; Li & Tarp, 2024; Ptasznik et al., 2024; Ptasznik & Lew, 2025).
Against this backdrop, the present study investigates how three widely used LLMs generate dictionary entries and examines their potential contributions and limitations within pedagogical lexicography. Focusing on AI-generated dictionary entries produced by ChatGPT, Gemini, and DeepSeek, the study adopts a comparative, criteria-based evaluation framework grounded in learner lexicography, pedagogical principles, and AI-assisted language learning. Its primary objective is to assess the extent to which LLMs can simulate the writing practices and decision-making processes of human lexicographers. To this end, the models’ ability to generate definitions, examples, and idiomatic expressions is evaluated using twenty carefully selected entries from Helix, a bilingual illustrated lexicon designed for Greek heritage learners (Chadjipapa & Gavriilidou, 2024). All models were prompted under controlled conditions to produce dictionary entries including definitions, sense distinctions, usage notes, examples, collocations, register labels, and, where relevant, learner-oriented or cross-linguistic explanations.
The chatbot-generated lexicographic articles were evaluated by four expert lexicographers specializing in pedagogical lexicography. The evaluation combined qualitative and quantitative dimensions and was conducted under blind conditions, with assessors unaware of whether entries were produced by human lexicographers or AI systems. Entries were assessed along five main axes: (a) definitional accuracy and semantic precision, (b) internal consistency and sense discrimination, (c) pedagogical usefulness for learners, (d) transparency and explicability of lexical information, and (e) alignment with established learner lexicographic principles, including clarity, economy, controlled defining vocabulary, and relevance to communicative use. Particular attention was paid to whether the models reproduced established lexicographic conventions, extended them, or violated core norms. The results show that all three LLMs are capable of producing superficially well-formed dictionary entries that resemble the structure and tone of learner dictionaries. Nevertheless, systematic differences were observed. ChatGPT demonstrated stronger coherence in extended explanations and pedagogical scaffolding, particularly in usage notes and learner oriented paraphrasing. Gemini excelled in the range of examples and contextual variation, though sometimes at the expense of semantic precision, while DeepSeek offered comparatively limited learner-specific guidance. These findings align with patterns reported in previous research (Lew, 2024). Across models, recurrent weaknesses included inconsistent sense hierarchies, blurred distinctions between core meaning and pragmatic inference, occasional hallucinated usage restrictions, and limited sensitivity to frequency, register, and learner difficulty.
Despite these limitations, the findings point to a significant pedagogical opportunity: when guided by appropriate pedagogical prompts and lexicographic constraints, LLMs can meaningfully support aspects of dictionary compilation. The contribution of this paper is therefore threefold. First, it provides one of the first systematic and comparative evaluations of AI-generated lexicographic articles produced by multiple LLMs, employing a blind assessment procedure that enables a more nuanced evaluation of whether chatbot generated outputs meet established lexicographic standards beyond surface-level fluency. Second, it proposes an evaluative framework grounded in pedagogical lexicography, offering concrete criteria for assessing and improving chatbot-mediated dictionary entries, particularly with respect to learner-oriented definitions, sense structuring, and usage information. Third, the paper advances the methodological integration of chatbots into lexicographic workflows by examining prompt design as a key mediating factor in AI-assisted dictionary making. Focusing on few-shot prompting (Brown et al., 2020), it analyes how chatbots adapt their outputs in response to carefully curated lexicographic examples and identifies parameters necessary for the effective, reliable, and pedagogically sound use of prompts. Taken together, these contributions frame chatbots as potential partners in pedagogical lexicography, capable of extending traditional dictionary functions through interaction, personalization, and AI supported mediation when guided by lexicographic expertise.