29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

Can AI-Assisted Normalisation Improve Access to Historical Lexicographic Data?

30 Sept 2026, 09:30
30m
01 | Sitzungssaal (OeAW Main Seat, 1st floor)

01 | Sitzungssaal

OeAW Main Seat, 1st floor

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speakers

Tarrin Wills (University of Copenhagen) Ellert Thor Johannsson (The Árni Magnússon Institute for Icelandic Studies)

Description

The topic of this article is AI-assisted normalisation and how it can be used in the context of a historical dictionary. We look at data from A Dictionary of Old Norse Prose, which is a lexicographic resource on the language of medieval Iceland and Norway. This dictionary has a long tradition of using non-normalised, diplomatic text editions as source for its lexicographic descriptions. As a result, the bulk of the citation material remains difficult to access for beginners and non-experts in the language. We tested several LLMs in order to assess whether it would be feasible to provide users with automatic normalisations. The article describes several phases in testing and comparing different LLMs and discussing the types of errors they made. The result show that from the initial experimental phase in 2025 the models have improved significantly and currently are able to provide fairly error-free normalisations. We also mention the option of using modern Icelandic as an intermediate stage in the normalisation process which would facilitate the use of NLP-tools. The conclusion is that LLM-assisted normalisation can support broader access to historical lexicographic data. Such tools combine automated processing with human oversight, allowing historical dictionaries to expand their usability while maintaining scholarly rigour.

Presentation materials

There are no materials yet.