Speakers
Description
The formulation of definitions is a core task of lexicography, traditionally carried out manually. Large language models now offer new opportunities to automate this workflow. While AI-assisted definition generation has been well studied for well-known vocabulary, this is not yet the case for lesser-known historical vocabulary, which is scarcely represented in training data. This paper therefore investigates how far targeted prompt engineering can improve the quality of AI-generated definitions for historical pandemic vocabulary, using data from the cholera corpus developed within the Pandemictionary project. Eight headwords of varying frequency and semantic complexity are analysed across three prompt configurations, focusing on the role of corpus evidence in generating lexicographical definitions. The results show that AI systems can generate formally correct, lexicographically structured definitions. Prompts without evidence draw on general model knowledge and often produce anachronistic or hallucinated content, whereas evidence-based prompts yield definitions that are more historically accurate and firmly grounded in the corpus. Corpus evidence proves a crucial quality factor, especially for rare or poorly documented lexemes, though challenges remain regarding hallucinations, polysemy, and reducing definitions to their essential elements. The study highlights the potential of prompt engineering for digital lexicography while underscoring the continuing central role of human expertise in quality assurance.