Speakers
Description
This paper addresses the challenge of prioritizing dictionary revisions in corpus-based lexicography, specifically for the Digital Dictionary of German (DWDS), which relies heavily on legacy content. Given the massive data volume in modern corpora, manual selection of relevant words for updating is no longer feasible. We investigate two methods for identifying “trend words”—terms showing an overproportionally high present-day frequency relative to past intervals. The first method applies established linear regression to large monitor corpora, highlighting limitations due to reliance on the chosen time period. To mitigate this, the second method employs a dispersion-oriented keyness metric on a specialized “discourse corpus” of current online text. This metric scores lemmas based on their spread across unique sources, effectively prioritizing words relevant to contemporary discourse. Evaluation of the ranked candidate list confirms that the high-score group contains a significantly higher density of relevant trend words compared to lower-ranked groups. This system allows lexicographers to systematically prioritize legacy entries for revision, such as those that are outdated or missing a current sense, thereby optimizing the lexicographical workflow and ensuring the dictionary reflects current language use.