29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

DiaLexPoL: LLM-assisted Sense Assignment for a Diachronic Dictionary of Latin

2 Oct 2026, 16:30
30m
01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)

01 | Johannessaal

01 | Johannessaal, OeAW Main Seat, 1st floor

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speakers

Krzysztof Nowak (Institute of Polish Language PAS) Iwona Krawczyk (Institute of Polish Language PAS) Jagoda Marszałek (Institute of Polish Language PAS)

Description

The DiaLexPoL project is building a diachronic lexical database of Neo-Latin (ca. 1550–1800) from the semi-automatically compiled diachronic corpus of Polish Latin (DiaCorPoL). Its principal bottleneck is sense assignment: referring early-modern attestations to an inventory designed for the Dictionary of Medieval Latin in Poland (DMLP). We ask whether deployable LLMs – cheap commercial mid-tier and self-hostable open-weight models, rather than frontier systems a chronically underfunded project cannot reproduce – can draft this step in a human-in-the-loop workflow. Two experiments address it: reproducing DMLP's own sense assignments on its quotations, and assigning Neo-Latin concordance lines to the DMLP inventory against independent human gold-standard annotation. Self-hostable open-weight models give the strongest drafts, outperforming the single commercial mid-tier model at negligible cost, and the best corpus configuration matches human inter-annotator agreement overall, though not on every lemma. Gains over a most-frequent-sense baseline – always assigning a lemma's commonest sense – are modest and concentrate on genuinely polysemous lemmas, where editors most need help, while on dominant-sense lemmas the models over-split, marking where LLM drafting helps.

Presentation materials

There are no materials yet.