29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

Session

Computational & AI-Based Lexicography

29 Sept 2026, 16:30
OeAW Main Seat

OeAW Main Seat

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Conveners

Computational & AI-Based Lexicography

  • Jelena Kallas (Institute of the Estonian Langauge)

Computational & AI-Based Lexicography

  • Iztok Kosem (University of Ljubljana & Joลพef Stefan Institute)

Computational & AI-Based Lexicography

  • Kristina ล . Despot (Institute for the Croatian Language)

Computational & AI-Based Lexicography

  • Ana Ostroลกki Aniฤ‡ (Institute for the Croatian Language)

Computational & AI-Based Lexicography

  • Gilles-Maurice de Schryver (Ghent University)

Computational & AI-Based Lexicography

  • Kris Heylen (Dutch Language Institute (INT))

Computational & AI-Based Lexicography

  • Miloลก Jakubรญฤek (Lexical Computing)

Presentation materials

There are no materials yet.

  1. Li Cui (Yonsei University), Seol Namkung (Yonsei University), Jun Lee (Yonsei University), Hae-Yun Jung (Kyungpook National University), Kilim Nam (Yonsei University)
    29/09/2026, 16:30

    The present study aims to empirically examine the possibilities and limitations of LLMs as a lexicographic tool for Korean, by automatically generating inflected forms of Korean verb and adjective headwords using Large Language Models and contrasting the results with corpus frequency data. Focusing on the description of inflected forms of Korean verbs and adjectives, the study analyses in...

    Go to contribution page
  2. Gilles-Maurice de Schryver (Ghent University)
    29/09/2026, 17:00

    At the time of the Euralex 2026 conference it will be nearly four years since the start of the GenAI revolution. The question that needs to be asked, and which will be answered, is thus: Is any of the past work in (academic) lexicography still of any relevance, or are LLM chatbots simply โ€˜dictionariesโ€™ in their own right and โ€˜lexicographersโ€™ plain and simple? The answer is a shocking one.

    Go to contribution page
  3. Igor Boguslavsky (IA.A.Kharkevich Institute for Information Transmission Problems), Vyacheslav Dikonov (A.A.Kharkevich Institute for Information Transmission Problems), Evgeniya Inshakova (A.A.Kharkevich Institute for Information Transmission Problems), Alexandre Lazursky (A.A.Kharkevich Institute for Information Transmission Problems), Svetlana Timoshenko (A. A. Kharkevich Institute for Information Transmission Problems), Tatiana Frolova (A.A.Kharkevich Institute for Information Transmission Problems)
    29/09/2026, 17:30

    This paper describes the lexicographic resources developed for the ETAP text analysis and generation system, with particular emphasis on its semantic module, SemETAP. In our approach, semantic analysis is viewed not merely as the construction of a semantic representation, but as the derivation of inferences licensed by linguistic and background knowledge. The semantic component operates in two...

    Go to contribution page
  4. Ivรกn Arias-Arias (Universidade de Santiago de Compostela), Marรญa Josรฉ Domรญnguez Vรกzquez (Universidade de Santiago de Compostela), Marรญa Teresa Sanmarco Bande (Universidade de Santiago de Compostela), Carlos Valcรกrcel Riveiro (Universidade de Vigo)
    01/10/2026, 14:00

    While recent scholarship highlights the generative capabilities of LLMs for drafting lexicographic entries, their โ€œblack boxโ€ nature poses challenges regarding training corpora, linguistic representativity, and semantic hallu-cination. We propose that this opacity is less problematic when LLMs are deployed not as linguistic experts, but as software engineers. By leveraging LLMs as agentic...

    Go to contribution page
  5. Hanna Fischer (Research Center Deutscher Sprachatlas), Alfred Lameli (Research Center Deutscher Sprachatlas), Nathalie Mederake (Research Center Deutscher Sprachatlas), Nico Urbach (Research Center Deutscher Sprachatlas)
    01/10/2026, 14:30

    Despite the growing availability of digital dictionaries, semantic interoperability often remains limited to headwords, metadata, or article structures. Dictionary senses remain difficult to compare because they are shaped by different editorial traditions, segmentation practices, and levels of semantic granularity. This paper asks whether taxonomy-guided LLM classification can generate...

    Go to contribution page
  6. Phoebe Nicholson (Oxford University Press), Will Rogers (Oxford University Press), Kate Wild (Oxford University Press)
    01/10/2026, 15:00

    The Oxford English Dictionary (OED) has long been recognized for its rigorous scholarship, yet it is equally notable for its readiness to adopt new technologies. As the volume and availability of linguistic evidence has grown in the digital era, so too have the challenges of efficiently revising a dictionary of this size and scope: in response, the OED has been undertaking a systematic...

    Go to contribution page
  7. Carlos Valcรกrcel Riveiro (Universidade de Vigo), Marรญa Teresa Sanmarco Bande (Universidade de Santiago de Compostela), Ivรกn Arias-Arias (Universidade de Santiago de Compostela), Marรญa Josรฉ Domรญnguez Vรกzquez (Universidade de Santiago de Compostela)
    01/10/2026, 16:00

    Digital dictionaries can become more than consultation tools when their internal data are sufficiently structured to be reused and transformed into embedded learning components. This paper evaluates whether large language models (LLMs) can convert PORTLEX entries, represented as reviewed JSON files, into interactive HTML didactic modules integrable into their corresponding dictionary entries....

    Go to contribution page
  8. Andrea Abel (Free University of Bolzano/Bozen & Eurac Research), Luca Ducceschi (Eurac Research), Federico Villa (Independent Researcher)
    01/10/2026, 16:30

    From a pluricentric perspective, German possesses several standard varieties, including the one used in South Tyrol (Italy). While automated lexicography has long focussed on differences of lexical forms across these varieties, meaning differences in formally identical words are more difficult to identify systematically. This study addresses such diatopic semantic variation by comparing usage...

    Go to contribution page
  9. Gilles-Maurice de Schryver (Ghent University)
    01/10/2026, 17:00

    This paper presents the background to a book on the visual history of academic lexicography, being a sequence of hundreds of one-page comic strips, created through conversations with LLM chatbots and GenAI image generators, each page representing a major publication in the field. Of course, I knew that this submission would be a gamble: Would a comic version of a conference paper be taken...

    Go to contribution page
  10. Elina Chadjipapa (Democritus University of Thrace), Zoe Gavriilidou (Democritus University of Thrace), Maria Mitsiaki (Democritus University of Thrace), Ifigeneia Dosi (Democritus University of Thrace)
    02/10/2026, 09:00

    As early as the eighteenth century, dictionary compilation has been perceived as a labor intensive and time-consuming endeavor, often characterized as meticulous and repetitive work (Lew, 2023). The creation of lexicographic articles, in particular, constitutes a complex process requiring the careful coordination of multiple parameters, including the target user group, usersโ€™ native language,...

    Go to contribution page
  11. Thordis Ulfarsdottir (The รrni Magnรบsson Institute for Icelandic Studies), Ellert Thor Johannsson (The รrni Magnรบsson Institute for Icelandic Studies)
    02/10/2026, 09:20

    This paper presents a pilot study exploring the use of artificial intelligence in reversing a bilingual dictionary. The Modern Icelandic-English Dictionary is a freely available online resource. The aim was to transform this Icelandic-English dictionary into a functional English-Icelandic dictionary while minimising manual labour. Using the DANTE English lexical database as a reference, the...

    Go to contribution page
  12. Tinatin Margalitadze (Ilia State University), Giorgi Meladze (Ilia State University), Ketevan Mchedlishvili (Ilia State University), Tamar Laluashvili (Ilia State University), Giorgi Okropiridze (Ilia State University)
    02/10/2026, 09:40

    The purpose of this paper is to present the results of a study investigating the effectiveness of AI in the compilation of the Dictionary of Georgian Neologisms, an ongoing project at the Centre for Lexicography and Language Technologies at Ilia State University. The data for the study were selected from a dataset of 1,700 neologisms identified in previous research that developed a...

    Go to contribution page
  13. Jesรบs Torres del Rey (University of Salamanca), M.ยช Teresa Fuentes Morรกn (University of Salamanca)
    02/10/2026, 14:00

    Electronic dictionaries pose particular accessibility challenges for blind screen reader users, yet the field lacks evaluation frameworks that account for the complexity of their information structures. This paper reports on a three-phase study comparing Automated Web Accessibility Evaluation Tools and Large Language Models in evaluating three online dictionary entries from different...

    Go to contribution page
  14. Iztok Kosem (University of Ljubljana & Joลพef Stefan Institute), Polona Gantar (University of Ljubljana, Faculty of Arts & Faculty of Computing and Information Science), Simon Krek (Joลพef Stefan Institute), Tjaลกa Arฤon (University of Ljubljana, Faculty of Computing and Information Science)
    02/10/2026, 14:30

    This paper presents an experiment on generating definitions for Slovene headwords using four Large Language Models (LLMs): two commercial (Gemini and GPT) and two open-source (Gemma and the Slovenian model GaMS). Each model was provided with contextual data, including semantic indicators, collocations, and examples. We tested a zero-shot approach and two few-shot approaches using sample...

    Go to contribution page
  15. Katharina Korecky-Krรถll (Austrian Academy of Sciences, ACDH), Philipp Stรถckle (Austrian Academy of Sciences, ACDH), Daniel Elsner (Austrian Academy of Sciences, ACDH), Wolfgang Koppensteiner (Austrian Academy of Sciences, ACDH), Elisabeth Eder (Austrian Academy of Sciences, ACDH)
    02/10/2026, 15:00

    This paper investigates the potential of Large Language Models (LLMs) to support lexicographic work in low-resource contexts, focusing on Austrian Bavarian dialects documented in the Wรถrterbuch der bairischen Mundarten in ร–sterreich (WBร–). Using a structured prompt-engineering workflow and a three-stage data pipeline, 100 dictionary articles were generated with LLaMA 4 (Scout) and...

    Go to contribution page
  16. Martina Vokรกฤovรก (Charles University), Anna Marklovรก (Charles University)
    02/10/2026, 16:00

    This paper investigates whether large language models can capture the embodied dimension of word meaning by comparing LLM-generated affordance norms with human-produced norms in Czech and English. Affordance norms are systematically collected data on the actions speakers associate with concrete objects. Human data were collected from 30 Czech native speakers in a free-production experiment...

    Go to contribution page
  17. Frantiลกek Kovaล™รญk (Lexical Computing), Marek Blahuลก (Lexical Computing), Miloลก Jakubรญฤek (Lexical Computing), Vojtฤ›ch Kovรกล™ (Lexical Computing)
    02/10/2026, 16:30

    Intra-annotator agreement is the rate at which an annotator makes the same choices in the same situation. This paper examines the annotation process in a semi-automatic dictionary-making project, Czech Dictionary Express. In the vocabulary-building process, its annotators went through 100,000 headwords from a corpus frequency wordlist, marking each as incorrect or (partially) correct. The aim...

    Go to contribution page
  18. Lydia Risberg (Estonian Language Institute, University of Tartu), Kristina Koppel (Estonian Language Institute), Margit Langemets (Estonian Language Institute), Hanna Maask (Estonian Language Institute), Esta Prangel (Estonian Language Institute), Maria Tuulik (Estonian Language Institute), Silver Vapper (Estonian Language Institute)
    03/10/2026, 09:00

    This paper investigates the potential of large language models (LLMs) as assistants with corpus analysis in lexicography, focusing on the assignment of register labels in the EKI Combined Dictionary (CombiDic). Inconsistencies in register labels become evident across synonym sets in CombiDic, which is the result of different lexicographers working on words at different times. We examine how...

    Go to contribution page
  19. Barbara Lewandowska-Tomaszczyk (University of Applied Sciences in Konin)
    03/10/2026, 09:30

    This paper investigates the accelerating influx of English slang into the Polish language, driven by digital globalization and social media usage. While historical Anglicisms often focused on technical or professional domains, modern borrowing increasingly penetrates the affective realmโ€”informal expressions of identity, emotion, and subcultural belonging. Central to this exploratory research...

    Go to contribution page
  20. Ines Rรถhrer (Bavarian Academy of Sciences and Humanities), Manuel Raaf (Bavarian Academy of Sciences and Humanities)
    03/10/2026, 10:00

    This article presents an application evaluation of the use of large language models (LLMs) for the semantic classification of lexicographical content in three German dialect dictionary projects. Building on earlier experiments with LLM-based semantic classification that showed hit rates of over 80%, the study investigates whether generative AI can support the assignment of meanings to...

    Go to contribution page
  21. Antonio San Martรญn (University of Quebec in Trois-Riviรจres), Catherine Trekker (University of Quebec in Trois-Riviรจres)
    03/10/2026, 10:30

    This paper proposes a human-centered artificial intelligence (HCAI) framework for AI-assisted lexicography. While generative AI offers significant opportunities to enhance lexicographic work, it also raises concerns regarding the future role of lexicographers and the preservation of linguistic and cultural diversity. Drawing on HCAI principles and previous applications in other language...

    Go to contribution page
Building timetable...