Speakers
Description
The DiaLexPoL project is building a diachronic lexical database of Neo-Latin (ca. 1550–1800) from the semi-automatically compiled diachronic corpus of Polish Latin (DiaCorPoL). Its principal bottleneck is sense assignment: referring early-modern attestations to an inventory designed for the Dictionary of Medieval Latin in Poland (DMLP). We ask whether deployable LLMs – cheap commercial mid-tier and self-hostable open-weight models, rather than frontier systems a chronically underfunded project cannot reproduce – can draft this step in a human-in-the-loop workflow. Two experiments address it: reproducing DMLP's own sense assignments on its quotations, and assigning Neo-Latin concordance lines to the DMLP inventory against independent human gold-standard annotation. Self-hostable open-weight models give the strongest drafts, outperforming the single commercial mid-tier model at negligible cost, and the best corpus configuration matches human inter-annotator agreement overall, though not on every lemma. Gains over a most-frequent-sense baseline – always assigning a lemma's commonest sense – are modest and concentrate on genuinely polysemous lemmas, where editors most need help, while on dominant-sense lemmas the models over-split, marking where LLM drafting helps.