Speakers
Description
Based on OpenAI’s moderation model, we describe a method developed to automatically identify unwanted discriminatory content in the citations found in the printed edition of Den Danske Ordbog (DDO, 2003–2005). Today, the DDO is published online and is continuously revised and expanded with new headwords. We begin by testing whether the model can identify 234 citations that have been removed since 2015 because the editorial team deemed them discriminatory. The results are promising, so we apply the model to 71,000 citations from headwords that have not yet been revised. The model classifies 6.3% of these as harmful. In addition to identifying the expected types of citations, a large proportion are flagged because they involve violence or self-harm—a type of problematic content that had previously been overlooked. Of the 4,800 flagged instances, 37% have been manually validated. Regarding violent content, the precision is only 10%—which nevertheless yields a list of approximately 350 problematic cases. The editorial team intends to discuss this type of content and establish new guidelines. For discriminatory content, the precision is higher—33%—which likewise results in a total of approximately 350 citations that are candidates for removal from the DDO and replacement with others.