29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

PoliLex: Modelling Heterogeneous Lexical Resources: Lessons from the Dictionary of Polish Dialects

2 Oct 2026, 15:00
30m
01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)

01 | Johannessaal

01 | Johannessaal, OeAW Main Seat, 1st floor

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speakers

Krzysztof Nowak (Institute of Polish Language PAS) Dorota Mika (Institute of Polish Language PAS)

Description

The Institute of the Polish Language (Polish Academy of Sciences) maintains a body of dictionaries and card-file archives assembled over more than a century, whose differences in structure, editorial convention, and descriptive metalan-guage have long kept them difficult to search together. This paper reports on PoliLex, an effort to bring such material under a common, machine-queryable representation without erasing the editorial character of each source. We de-scribe an encoding and publication workflow that moves corrected dictionary text through TEI into an OntoLex-Lemon Linked Data representation, and a layered project ontology whose load-bearing decision is to keep a dictionary, as a published text, separate from the language system it records. We illustrate the approach with the Dictionary of Polish Dialects, the first resource processed through the full pipeline, and show how the same encoded source is served both as browsable full text and as a SPARQL-queryable graph. We then examine what such harmonisation is designed to achieve across resources, what it achieves today on a single resource, and where it stops — notably at cross-dictionary sense alignment.

Presentation materials

There are no materials yet.