Speakers
Description
At Euralex 2024, the initial design of a digital lexical infrastructure for Dutch Dialects was presented, developed in response to requests from dialect organizations in the Netherlands and Flanders to support the compilation and publication of dialect dictionaries. That contribution focused primarily on conceptual design choices and technical feasibility. Two years later, the project has moved from exploration to practice. The infrastructure has been implemented, tested, and refined through real-world use by two dialect communities. This paper reports on these developments, with particular attention to lexicographic equivalence, source integration, orthographic variation, and the role of non-professional lexicographers in collaborative dialect dictionary making.
Despite ongoing dialect loss, public engagement with regional varieties of Dutch remains strong. Dialect associations play a central role in documentation, education, and dissemination, often through dictionary projects or learning materials. This work is largely carried out by volunteers. While their commitment and local knowledge are indispensable, such initiatives typically operate with limited resources and without structural access to linguistic or lexicographic training. As a result, dialect dictionaries are often developed in isolation, follow divergent conventions, and are difficult to update, compare, or integrate. Lexical data are also at risk of being lost due to inadequate long-term preservation (Van Keymeulen et al. 2009). The project discussed here addresses these challenges by developing a shared digital infrastructure that supports sustainable dialect lexicography while accommodating differences in expertise, goals, working practices, and community-specific preferences.
Two pilot varieties were selected: Bildts, a Frisian-Dutch contact variety spoken in the province of Fryslân (Friesland), and West-Overijssels, a regional variety with a long lexicographic tradition. These cases were chosen not only for their linguistic characteristics, but also for the presence of active, volunteer-driven language communities. In both pilots, the first step consisted of establishing a lexical database as the backbone for future dictionary products, reflecting the lexicographic principle that a well-defined macrostructure is a prerequisite for consistent and reusable microstructural description.
The editorial workflow is centred on a lexical data editing environment (Lex’it) (fig. 1*) that supports structured data entry, linking, and revision. As an initial alignment device, a Dutch base word list derived from frequency-oriented learner resources was used. This list functions as a pragmatic onomasiological access structure, intended to facilitate dialect acquisition, translation activities, and comparison across dialects. Experience from the pilots confirmed, however, that frequency-based lists alone are insufficient for representing dialect lexicons. Many high-frequency standard-language words lack clear dialect equivalents, while culturally salient dialect words are often absent from such lists. Volunteers frequently prioritize local relevance and expressive richness over frequency.
In both pilots, existing dialect dictionaries (Buwalda 2013, Fien 2000, Kamman 1990, Kuijk & Van Baalen 2015) were integrated into the editing environment. For Bildts, one authoritative dictionary served as the primary source, supplemented with newly attested items stored in an extensible lexicon. For West-Overijssels, data from three independent dictionaries were combined, enabling the compilation of a compact, learner-oriented pocket dictionary suitable for a broader region. This process foregrounded classical lexicographic issues such as source criticism, sense delimitation, synonymy, and the treatment of variants.
Combining multiple sources substantially increased lexical coverage but also introduced ambiguity. Differences in spelling conventions, semantic scope, and levels of detail between dictionaries required explicit editorial decisions. Orthographic variation proved particularly prominent in volunteer-based projects, as contributors often follow different local, historical, or personal spelling systems. The infrastructure therefore supports the coexistence of multiple spelling variants without enforcing premature standardization, while still providing reliable search and access mechanisms (fig. 2*).
The pilots further demonstrated that lexical linking cannot be reduced to simple one-to-one equivalence. Partial equivalence, context-dependent meanings, and culture-specific concepts occur frequently in dialects. All editorial actions are logged within the system, enabling transparency, revision, and quality control (fig. 3*).
A key outcome after two years is that the same underlying system now supports two distinct public dictionary products (e.g. fig. 4 and 5*). The Bildts and West-Overijssels lexica are published on the respective websites of the collaborating organizations, each with its own interface and local identity, reflecting heterogeneous user groups ranging from experienced dialect speakers to beginners. At the infrastructural level, however, both products are generated from the same database and editorial workflow, illustrating a clear separation between lexicographic data, structure, and presentation.
The pilots show that creating digital dialect dictionaries involves not only linguistic data, but also volunteer practices, orthographic diversity, and diverse user needs. By embedding lexicographic theory into infrastructural design, the project offers a sustainable model that combines scholarly rigor with accessibility.
* see Book of Abstracts