Speakers
Description
Like other lexicographic and language institutes in Europe, the Dutch Language Institute (Instituut voor de Nederlandse Taal, INT) is progressively integrating the compilation and management of its individual lexicographic products into a single shared lexicographic infrastructure with the aim of increasing internal efficiency and enable new forms of use, research and development (Depuydt, Tiberius & Heylen 2026). Comparable infrastructural trajectories can be observed for Estonian (Koppel et al. 2019), Danish (Pedersen et al. 2018), or Slovene (Gantar 2020). For the INT, this integration is particularly challenging because its historical and contemporary dictionaries were originally largely developed as stand-alone projects and differ substantially in their internal structure.
Over the past two decades, the INT has systematically invested in addressing this challenge by developing a shared infrastructural backbone for Dutch lexicography. Since 2007, this has resulted in the creation of GiGaNT (Groot Geïntegreerd Lexicon van de Nederlandse Taal), a centralized lexicon providing unified lemmatisation principles, part-of-speech tagging, and form-related information for both historical and contemporary Dutch dictionaries and lexical databases (Ruitenberg, de Does, Depuydt 2010). GiGaNT enables the alignment of lexicographic resources at the lemma level across time periods and projects, and already supports the reuse of formal and morphological information in multiple lexicographic end products.
The present paper focuses on the next infrastructural step: the development of a sense inventory that extends this integration from the lemma level to the sense level so as to share and combine information that is currently distributed across databases. However, these resources, including the historical Woordenboek der Nederlandsche Taal (WNT), the contemporary general Algemeen Nederlands Woordenboek (ANW), and the phraseological Woordcombinaties differ not only in sense granularity, but also in their organising principles and descriptive focus. Therefore, a sense inventory is introduced as an intermediate semantic layer that allows these resources to be linked through core senses, to which definitions, corpus attestations, collocational patterns, usage information, and sense relations can be attached.
The objectives of this infrastructural development are twofold. Internally, the sense inventory is intended to support more efficient lexicographic workflows, enable systematic consistency checking across resources, and facilitate the reuse of semantic information in multiple dictionary products and services. Externally, it forms part of a publicly accessible lexicographic infrastructure that can be used for research on the Dutch lexicon, for the development of language learning materials, and for language-technology applications. In addition, the Dutch sense inventory is explicitly designed to contribute to and represent Dutch within the European lexicographic infrastructure ELEXIS, which will be substantially upgraded in the recently started Horizon Europe infrastructure project ELEXAI (2026-2029). Within ELEXAI, a collaborative and dynamic sense inventory will play a central role in the further development of the Dictionary Matrix and the European lexicographic knowledge graph.
The design of the Dutch sense inventory builds on earlier work at the INT, most notably the DiaMaNT project (Depuydt & de Does 2018), which established a semantic lexicon for historical Dutch dictionaries by linking definitions across resources at a schematic level. This experience provides both methodological insights and reusable data for extending semantic integration to contemporary lexicography. At the same time, the current work takes into account international models and standards for lexicographic data modelling, including the DMLex data model (DMLex-1.0) and earlier proposals for unified sense-level organisation (e.g. Tavast et al. 2018).
The design of the Dutch sense inventory builds on earlier work at the INT, most notably the DiaMaNT project (Depuydt & de Does 2018), which established a semantic lexicon for historical Dutch dictionaries at a schematic level across resources by levering synonym definitions. This experience provides both methodological insights and reusable data for extending semantic integration to contemporary lexicography. At the same time, the current work takes into account international models and standards for lexicographic data modelling, including the DMLex data model (DMLex-1.0) and earlier proposals for unified sense-level organisation (e.g. Tavast et al. 2018).
This paper presents four interrelated components of the ongoing work. First, we introduce a preliminary data model for the Dutch sense inventory, which builds on the existing GiGaNT model while incorporating sense-level components already present in individual dictionaries. The model is designed to accommodate heterogeneous sense structures and to remain compatible with European interoperability efforts. Second, we discuss the principles for sense linking and sense lumping into core senses. These principles explicitly address large differences in sense granularity between dictionaries, the specific organising principles of Woordcombinaties, and the availability of explicit sense relations in the ANW, with attention to recurrent phenomena such as regular polysemy, for which earlier proposed (e.g. Pedersen et al. 2022) are evaluated and adapted.Third, we describe the compilation of a test dataset that serves both as a proof of concept and as an evaluation benchmark. The dataset includes lemmata with different parts of speech, varying degrees and types of polysemy, and contrasting treatments in the historical WNT, the contemporary ANW, and the pattern-based Woordcombinaties. It builds on previously curated material from DiaMaNT and on the Dutch component of the multilingual evaluation dataset for monolingual word sense alignment (Ahmadi et al. 2022). Finally, we discuss how this dataset will function as a baseline for future experiments in automatic sense linking, lumping and splitting using both open-weights models and commercial large language models.