29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

Large-Scale Semiautomatic Detection of New and Unregistered Words in Modern Ukrainian

2 Oct 2026, 10:20
20m
01 | Sitzungssaal (01 | Sitzungssaal, OeAW Main Seat, 1st floor)

01 | Sitzungssaal

01 | Sitzungssaal, OeAW Main Seat, 1st floor

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speakers

Vasyl Starko (Ukrainian Catholic University) Andriy Rysin (Independent Researcher)

Description

We have created and implemented a workflow for large-scale semiautomatic detection of new and previously unregistered words in Modern Ukrainian. Crucially, it involves a diverse corpus of Ukrainian (2 billion tokens), a large electronic morphological dictionary of Ukrainian, and a specially crafted NLP toolkit. The key methodological innovation lies in the dual use of the dictionary. On the one hand, it serves as a resource for lemmatizing and morphological tagging the corpus; on the other, it is cyclically enriched with new words detected in the corpus. The corpus is dynamic, as it is regularly expanded by adding a variety of texts from the 19th century to the present day. The dictionary, currently containing 444,000 lemmas, is dynamically reused with the help of a tailor-made NLP toolkit as an exclusion source (Janssen, 2009) against the growing corpus. The paper provides a detailed step-by-step description of the procedure. This semiautomatic pipeline has yielded thousands of new and unregistered Ukrainian words, laying the foundation for the systematic, large-scale exploration of the Modern Ukrainian lexicon.

Presentation materials

There are no materials yet.