Speakers
Description
We have created and implemented a workflow for large-scale semiautomatic detection of new and previously unregistered words in Modern Ukrainian. Crucially, it involves a diverse corpus of Ukrainian (2 billion tokens), a large electronic morphological dictionary of Ukrainian, and a specially crafted NLP toolkit. The key methodological innovation lies in the dual use of the dictionary. On the one hand, it serves as a resource for lemmatizing and morphological tagging the corpus; on the other, it is cyclically enriched with new words detected in the corpus. The corpus is dynamic, as it is regularly expanded by adding a variety of texts from the 19th century to the present day. The dictionary, currently containing 444,000 lemmas, is dynamically reused with the help of a tailor-made NLP toolkit as an exclusion source (Janssen, 2009) against the growing corpus. The paper provides a detailed step-by-step description of the procedure. This semiautomatic pipeline has yielded thousands of new and unregistered Ukrainian words, laying the foundation for the systematic, large-scale exploration of the Modern Ukrainian lexicon.