29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

Tracking Eight Centuries: The Czech Monitor Corpus as a Resource for Diachronic Lexical Research

2 Oct 2026, 17:00
30m
01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)

01 | Johannessaal

01 | Johannessaal, OeAW Main Seat, 1st floor

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speakers

Václav Cvrček (Charles University) Martin Stluka (Charles University) Klára Pivoňková (Charles University)

Description

This paper demonstrates the central role of richly annotated diachronic corpora in investigating historical lexical change, using the newly compiled Monitor Corpus of Czech as a case study. The corpus covers over eight centuries of the Czech written tradition and is designed as a general-purpose, consistently annotated resource for diachronic research. Unlike earlier diachronic corpora of Czech, the Monitor Corpus aims at systematic genre balance (where historical conditions permit) and at comparability of annotation across periods. To make the corpus maximally exploitable, it is accompanied by a dedicated application, Timeline, which supports queries and visualizations tailored to diachronic analysis.

The design of the Monitor Corpus seeks to represent, for each period, the broadest possible range of text types at the most general level: fiction, non-fiction, and journalism. Of course, this ideal cannot be met uniformly throughout the history of Czech. For some periods, no specialist texts have survived in Czech and journalism emerges only from the 18th century onwards. Nevertheless, in those historical stages where it is feasible, the corpus includes roughly comparable proportions of fiction, non-fiction, and journalistic writing, thereby enabling meaningful comparisons of lexical developments across time and genre.

For lexicographical and broader linguistic research, uniform, transparent and, above all, consistent annotation over time is essential. A key methodological challenge is how to handle both graphical and morphological variability so that it becomes possible to track lexical items across centuries. For example, the adjective estetický ‘aesthetic’ is attested in the 19th-century material in numerous spelling variants (e.g. aestetického, Esthetickému, aesthetickou). These forms must be recognized as instances of a single lemma in order to study its history in a principled way. To address this, we compiled three manually annotated etalon corpora representing different historical stages of Czech and used them as training data for a lemmatizer and POS/morphological tagger developed within the Universal Dependencies (UD) framework (de Marneffe et al., 2021; Zeman et al., 2021). The result is a diachronic corpus that is lemmatized and morphologically annotated according to a single, coherent scheme.

The benefits of this annotation become especially clear when investigating developmental tendencies that would be difficult or impossible to access using only raw text. By systematically linking historically and graphically divergent word forms under a single lemma and assigning each token a detailed morphological description, we can reliably estimate frequencies, identify collocational patterns, and detect changes in distribution across grammatical categories and constructions. The paper will describe how historical variants were handled in the annotation workflow, and how the resulting data can be exploited to study lexical change.

As concrete case studies illustrating how the Monitor Corpus and Timeline application can be used in practice, we focus on examples of nouns that have undergone semantic change or diachronic shifts in usage, such as nápad ‘attack’ / ‘idea’, puška ‘container’ / ‘rifle’, and národ ‘family’ / ‘nation’. For example, the word národ has undergone a complex semantic evolution since the earliest periods of Czech: attested meanings include ‘fruit’, ‘family/kin’, ‘pagans’, and progressively the now dominant sense ‘community of people sharing a territory or language’. Its sociopolitical salience in the 19th century is reflected in a marked increase in frequency in that period. The annotated corpus allows us to examine not only this overall frequency development, but also changes in the distribution of grammatical forms and constructions, and to relate them to semantic and discursive shifts.

In Old Czech, plural forms are strongly predominant, and one important meaning of the time—‘pagans’—is expressed exclusively in the plural (národové ‘nations’). From the early 19th century onward we observe a significant rise in genitive forms, and the lexical content of národ becomes increasingly abstract and context-dependent. The heads of noun phrases in which národ appears make an important contribution to its discursive interpretation. In the 19th century, typical genitive constructions include duch národa ‘the spirit of the nation’, and rozkvět národa ‘the flourishing of the nation’, reflecting nationalist and Romantic discourses. By contrast, in 21st-century texts, constructions such as společenství národů ‘community/commonwealth of nations’ and organizace národů ‘organization of nations’ become prominent, mirroring a shift towards international and institutional frames.

Using a small set of nouns as case studies, the paper will show how lemmatization enables the automatic extraction of statistically robust lemma-based collocations across periods and genres, and how morphological annotation supports the identification of typical semantic–syntactic patterns. Both lines of research use the Timeline application, which allows various frequency data and collocation profiles to be extracted and visualized on a timeline. The Monitor corpus and the Timeline application are currently in beta and are expected to be made publicly available by the end of 2026.

Presentation materials

There are no materials yet.