Speakers
Description
This paper investigates the potential of Large Language Models (LLMs) to support lexicographic work in low-resource contexts, focusing on Austrian Bavarian dialects documented in the Wörterbuch der bairischen Mundarten in Österreich (WBÖ). Using a structured prompt-engineering workflow and a three-stage data pipeline, 100 dictionary articles were generated with LLaMA 4 (Scout) and systematically compared to human-authored “gold” articles. Evaluation combined automated metrics, LLM-based judgement, and assessment by three human lexicographers. Results show that LLMs reliably reproduce formal lexicographic conventions, producing well-structured and coherent dictionary entries. Nevertheless, limitations emerge in semantic completeness and data fidelity: generated articles sometimes omit senses or misrepresent semantic structure and often inadequately integrate empirical evidence, especially regional information and attestations. While automated metrics indicate a relatively high similarity of original and generated senses based on their embeddings, both human evaluators and LLM-as-a-judge approaches reveal persistent weaknesses. However, the strong alignment between human judgements and the sense similarity metric suggests that embedding-based evaluation is a promising proxy for human assessment, whereas holistic LLM-as-a-judge scores appear comparatively strict and less transparent. The study argues for a hybrid evaluation framework and positions LLMs as assistive tools within human-in-the-loop lexicography.