29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

The Role of LLMs in the Automatic Distribution of Meanings from Knapiusz's Polish–Latin–Greek Dictionary of 1643: A Pilot Study

30 Sept 2026, 14:00
1h 30m
Aula (Entrance Hall) & 1st floor (00/01 | Aula (Entrance Hall) & 1st floor, OeAW Main Seat)

Aula (Entrance Hall) & 1st floor

00/01 | Aula (Entrance Hall) & 1st floor, OeAW Main Seat

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speakers

Magdalena Majdak (Polish Academy of Sciences) Ewa Rodek (Polish Academy of Sciences) Katarzyna Kryńska (Polish Academy of Sciences) Jagoda Marszałek (Polish Academy of Sciences)

Description

Introduction
Grzegorz Knapiusz’s Thesaurus polonolatinograecus (second, expanded edition 1643; hereafter Kn) is the largest seventeenth-century dictionary that also records Polish vocabulary. It is one of the source materials for the Electronic Dictionary of Seventeenth- and Eighteenth-Century Polish (hereafter e-SXVII). One challenge for e-SXVII editors is assigning the content of Knapiusz’s entries—which contain, in addition to Polish, extensive Latin and Greek material provided as equivalents of Polish items—to the senses distinguished in e-SXVII entries. Although the Thesaurus was intended primarily as an extensive synonym dictionary supporting Latin learning, it also contains numerous observations on Polish vocabulary. It is regarded as the first work in Polish lexicography to address the semantics of Polish words. Knapiusz does not explicitly divide meanings within entries; instead, the Latin and Greek equivalents, or associated lexical items, of a Polish headword shift fluidly from one meaning to another, without clearly marked boundaries.

Aims
This study aimed to test whether LLMs can support the segregation and disambiguation of meanings in Kn entries, as well as their assignment to corresponding senses in e-SXVII. We designed and evaluated a procedure for the semi-automatic distribution of Kn material within e-SXVII entries using GPT-5.2 Thinking, Gemini 3 Pro Thinking, Perplexity Sonar Pro, and Claude Sonnet 4.5.

Method
We tested ten polysemous nominal entries, including mieczyk ‘short sword; gladiolus; yellow iris; bulbous iris; stinking iris; swordfish; Prussian coin’, dowód ‘proof; document; evidence; demonstration; reason’, laska ‘staff; mace; wand; rod; ridge; office; shepherd’s purse; unit of measure’, wyrok ‘oracle; decree; sentence; fate’, zapis ‘record; deed; bequest; legal act; treaty; statute; register; address; lodging assignment’, gwałt ‘assault; force; outrage; rape; corvée labour; violence; necessity; tumult’, wynalazek ‘idea; invention; discovery; decision; finding; result; decree; compromise’, bukowina ‘beech forest; woodland; beech trees; beech wood; beech grove; forest settlement’, motłoch ‘rabble; mob; worthless mass; commoners; riffraff’, and jabłko ‘apple; apple tree; apple-like fruit; pomegranate; truffle; pomander; heraldic charge’.

The trials differed in the amount and type of information provided to the models: (1) a basic instruction to translate, segregate and disambiguate Latin and Greek equivalents in Kn entries; (2) the same instruction with reference to the sense structure and definitions of model e-SXVII entries; (3) the instruction from trial 2 supplemented with quotations and established word combinations; (4) the instruction from trial 2 supplemented with the e-SXVII principles for defining senses and distinguishing polysemy from homonymy; and (5) the instruction from trial 4 supplemented with illustrative material from e-SXVII.

To reduce the influence of conversational context, each Kn entry was queried separately, in a new conversation. In evaluating the results, we took into account inter-entry variability resulting, among other factors, from different degrees of lexical polysemy and different entry structures. Outputs were compared with model e-SXVII entries and evaluated by lexicographers and classical philologists according to an agreed scale covering the correctness of sense division, the assignment of equivalents to senses, and agreement with e-SXVII definitions.

Results and conclusions
Perplexity Sonar Pro was excluded because its responses were not consistently generated by the same model. GPT-5.2 Thinking, Gemini 3 Pro Thinking, and Claude Sonnet 4.5 produced outputs of highly variable accuracy. Each model performed flawlessly on some entries but made significant errors in others. This variability appears to be related to differences in polysemy and entry structure. Increasing prompt specificity did not produce a straightforward improvement in accuracy, suggesting that output instability cannot be resolved merely by enlarging the sample without changing the validation procedures.

The trials also examined the minimum amount of information needed for an AI model to identify the semantic structure of a Kn entry correctly. For entries with highly typical mechanisms of polysemy, such as bukowina ‘a forest dominated by beech trees’ → ‘individual beech trees’ → ‘beech wood’, satisfactory results were obtained as early as the first trial. However, in comparable cases, such as jabłko ‘apple fruit’ → ‘apple tree’ or ‘round fruit’ → ‘round object’, the results were unsatisfactory both in the first trial and after additional context had been supplied. For some entries, subsequent trials even led to a deterioration in the disambiguation of Kn meanings, regardless of the AI model used.

The pilot study revealed clear limitations of the tested procedure. Verification of the method showed that any attempt to automatically divide meanings within loosely structured historical dictionary entries must take into account both their content and their intended function. In the case of Knapiusz’s Thesaurus, most illustrative material was taken from ancient authors and may therefore represent meanings not attested in Polish. Moreover, the study indicates that the current functioning of the selected models does not allow for standardised results or fully automatic support for lexicographic work. Nevertheless, under appropriate methodological assumptions, AI models with a “thinking” mode may offer support to lexicographers by generating hypotheses about senses, identifying problematic equivalents, and comparing Kn material with e-SXVII definitions

Presentation materials

There are no materials yet.