EURALEX 2026
OeAW Main Seat
XXII EURALEX INTERNATIONAL CONGRESS 2026: LEXICOGRAPHY IN THE AGE OF AI
The rapid development of artificial intelligence (AI) and natural language processing (NLP) has begun to transform the ways in which lexicographic data are compiled, analysed, and presented. Large language models, advanced corpus tools, and other machine learning methods offer unprecedented opportunities to accelerate tasks such as the extraction of lexical information, the induction of word senses, or the drafting of definitions. They also make possible new forms of user interaction, including conversational and personalized dictionaries that can adapt to different audiences or integrate multimodal content such as text, sound, and images.
At the same time, these technologies pose serious challenges that go beyond purely technical concerns. Automatically generated content raises questions of reliability, bias, and transparency, while the integration of opaque systems into lexicographic workflows risks undermining scholarly standards of quality and accountability. The ecological impact of large-scale AI models must also be taken into consideration, raising doubts about their long-term sustainability in resource-intensive projects such as lexicography. Finally, the introduction of AI into the field invites reflection on the role of the lexicographer: how can human expertise remain central when automation is increasingly pervasive, and what kinds of collaborations between humans and machines are both effective and responsible?
The conference Lexicography in the Age of AI will provide a forum for discussing these promises and risks. We will bring together professional lexicographers, publishers, researchers, software developers, and anyone interested in dictionaries to consider both the potential of AI for creating more dynamic, interactive dictionaries, and the ethical and ecological responsibilities that come with it.
-
-
Pre-Conference Workshop: Constructicography in Practice: Constructions, Challenges, and Opportunities 02 | Museumszimmer (02 | Museumszimmer, OeAW Main Seat, 2nd floor)
02 | Museumszimmer
02 | Museumszimmer, OeAW Main Seat, 2nd floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConveners: Voula Giouli (Aristotle University of Thessaloniki), Athina Sioupi (Aristotle University of Thessaloniki), Nina Böbel (Heinrich Heine University Düsseldorf), Alexander Ziem (Heinrich Heine University Düsseldorf)-
1
Constructicography in Practice: Constructions, Challenges, and Opportunities
Constructicography (Lyngfelt et al. 2008), the systematic description of constructions in Constructicons and construction-based lexicographic resources, has emerged as a key area at the intersection of Construction Grammar (i.e., Croft 2001), lexicography, and computational linguistics. While constructions have long been acknowledged as central units of linguistic knowledge, their consistent representation, annotation, and integration into lexicographic resources remain methodologically and technologically challenging.
The proposed workshop aims to bring together researchers working on theoretical, descriptive, and computational aspects of constructicography. It will focus on how constructions are identified, modeled, represented, and linked to lexical, semantic, and pragmatic information across languages and resource types. Particular attention will be paid to challenges such as construction granularity, variation, productivity, cross-linguistic comparability, corpus-driven discovery, and interoperability with existing lexical resources.
At the same time, the workshop highlights emerging opportunities offered by large corpora, annotation frameworks, and NLP techniques, including construction mining, semi-automatic Constructicon building, and the use of large language models in constructional analysis. Moreover, the workshop will explore applications of Constructicons, particularly in language teaching and learning.
By fostering dialogue between linguists, lexicographers, educational experts and computational researchers, the workshop seeks to advance constructicography as a mature and practically applicable field within lexicography.
-
1
-
Pre-Conference Workshop: Multi-Word Expressions and Phraseology: Corpus-Based and Computer-Processed 00 | Anton Zeilinger Salon (00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor)
00 | Anton Zeilinger Salon
00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConveners: Lian Chen (LLL, University of Orléans), Besim Kabashi (Eberhard Karls Universität Tübingen)-
2
Multi-Word Expressions and Phraseology: Corpus-Based and Computer-Processed
Phraseology (Cowie 1998; Granger & Meunier 2008; Mitkov 2017; Mel’čuk 2023; Polguère, 2002, 2014; Mejri 2018; Chen 2021), or multiword expressions (MWEs) (Savary 2008; Constant 2012) occupy a central position in linguistic description, language use, and language learning. Idioms, collocations, lexical bundles, and fixed or semi-fixed expressions constitute a significant part of natural language, yet they remain challenging to model, describe, and represent in lexicographic resources (Pecina, P. 2010; Mel’čuk 2011, Polguère 2014, Chen 2025). With the rapid development of new technologies—particularly corpus linguistics, Natural Language Processing (NLP), and Artificial Intelligence (AI)—phraseology and lexicography are currently undergoing a profound methodological and conceptual transformation.
From a modeling perspective, new technologies make it possible to move beyond traditional, intuition-based descriptions of phraseological units. Large-scale corpora (Mitkov 2017), both general and specialized, enable the systematic identification of MWEs through statistical, distributional, and syntactic approaches. Methods such as n-gram extraction, association measures (PMI, t-score), syntactic patterning, and embedding-based similarity allow researchers to capture degrees of fixedness, semantic compositionality, and contextual variability. These advances open new perspectives for representing phraseological knowledge in structured models, including ontologies, lexical networks, and standards such as OntoLex-Lemon (McCrae, Bosque-Gil et al. 2017 ; Bosque-Gil et al. 2019).
In terms of resource creation, technological tools have profoundly reshaped lexicographic practices (Atkins & Rundell 2008; Granger & Paquot 2012). Digital corpora, web-based data, and annotation platforms facilitate the semi-automatic extraction and validation of phraseological units (Mitkov 2017; Evert 2008; Gries 2008; Constant et al. 2017). New-generation lexical resources integrate rich metadata, usage examples, frequency information, and semantic relations, making them more dynamic and interoperable (Polguère 2014; Ci-miano et al. 2016; McCrae et al. 2017).. Moreover, computational approaches enable the development of multilingual and contrastive resources, which are essential for studying phraseological variation across languages and cultures (Paquot 2015; Mel’čuk 2011).
New technologies also play a crucial role in the design of pedagogical dictionaries and lear-ning-oriented resources (Bogaards & van der Kloot 2002; Lew 2012). Phraseology is often a major obstacle for language learners, as MWEs cannot always be interpreted compositionally (Wray 2002; Howarth 1998; Granger 1998). Corpus-based examples, learner-oriented definitions, and adaptive digital interfaces can significantly enhance the accessibility and usability of phraseological information (Sinclair 2004; Granger & Paquot 2015; Paquot 2015). AI-driven tools, such as intelligent tutoring systems or LLM-assisted lexicography, offer promising avenues for generating contextualized examples, usage notes, and difficulty pro-files tailored to learners’ needs (Heift & Schulze 2007; Godwin-Jones 2023; Bender & Koller 2020).
In translation lexicography, the impact of new technologies is equally significant (Hartmann 2007; Atkins & Rundell 2008; Tarp 2012). Phraseological units pose well-known challenges in translation due to their idiomaticity and phraseocultural specificity (Chen 2022a; 2022b). Parallel corpora, alignment tools, and machine translation systems provide valuable data for identifying translation equivalents, partial correspondences, and translation strategies (Chen et al. 2024). Digital translation dictionaries can now incorporate cross-linguistic mappings, semantic annotations, and real attested examples, bridging the gap between lexicographic description and actual translational practice.
Finally, the analysis of existing dictionaries through technological means (Béjoint 2010; Lew 2013) offers new insights into lexicographic traditions and practices. Computational methods allow for the large-scale comparison of dictionary entries, coverage, microstructure, and treatment of phraseological units. Such analyses contribute to a critical understanding of how dictionaries evolve in response to technological, linguistic, and societal changes (Hausmann 1989; Tarp 2012; Fuertes-Olivera & Bergenholtz 2011).
The intersection of phraseology, multi-word expressions, lexicography, and new technologies constitutes a fertile research domain, fostering innovative models, richer resources, and more effective tools for analysis, learning, and translation.
-
2
-
Pre-Conference Workshop: Postediting Lexicography with Sketch Engine and Lexonomy 01 | Sitzungssaal (01 | Sitzungssaal, OeAW Main Seat, 1st floor)
01 | Sitzungssaal
01 | Sitzungssaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Ondřej Matuška (Lexical Computing)-
3
Postediting Lexicography with Sketch Engine and Lexonomy
The tutorial aims to promote the understanding of state-of-the-art lexicographic practices based on the post-editing of corpus-generated and AI-generated content. It focuses on how contemporary lexicography can effectively combine automation with expert human judgement to achieve efficient, scalable, and cost-effective editorial workflows while ensuring that all dictionary data remain fully reviewed and validated by human editors.
The workshop provides an overview of current corpus and AI technologies relevant to dictionary making. Sketch Engine ( https://www.sketchengine.eu/ ) is exploited to support key lexicographic tasks such as sense analysis, collocation extraction, example selection, and pattern discovery. The session will also introduce Lexonomy ( https://guide.lexonomy.eu/ ) as a modern dictionary writing system designed for updating legacy dictionaries as well as creating new lexicographic works. Lexonomy demonstrates how corpus evidence and post-edited content can be efficiently structured, edited, and published within a collaborative editorial environment.
A central component of the tutorial is the presentation of the Dictionary Express method ( https://dictionary.express/ ), which combines automated content generation with systematic post-editing to accelerate dictionary production without compromising quality. Through practical workflows and concrete examples, the tutorial will show how corpus data, AI-assisted suggestions, and editorial guidelines can be integrated into a coherent process that supports consistency, transparency, and editorial control. The session is intended for lexicographers, editors, and researchers interested in applying corpus-based and AI-supported methods to contemporary dictionary projects.
-
3
-
Pre-Conference Workshop: The 8th Globalex Workshop on Lexicography and Neology (GWLN-8) 01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)
01 | Johannessaal
01 | Johannessaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConveners: Ilan Kernerman (Lexicala by K Dictionaries), Kris Heylen (Dutch Language Institute (INT))-
4
The 8th Globalex Workshop on Lexicography and Neology (GWLN-8)
In recent years, the Globalex Workshop on Lexicography and Neology series (GWLN) has evolved into an international forum for lexicographers, scholars, tool developers, and other practitioners interested in how new words and new senses emerge, are detected, are evaluated as candidates for inclusion, and described in lexicographic resources. Since its inauguration, the workshop has taken place in seven consecutive years in conjunction with annual conferences of the continental lexicography associations, including DSNA (2019), Euralex (2020 online, and 2022), Australex (2021), Asialex (2023), Afrilex (2024), as well as at eLex (2025). Each edition has brought together participants from diverse linguistic traditions – covering both high- and low-resource languages – to share their experience and exchange perspectives on the challenges, methodologies and processes for combining lexicography and neology. For the eighth iteration, GWLN is teaming up with the European Network on Lexical Innovation (ENEOLI COST Action) to further strengthen collaboration in Europe and beyond.
The main topic of GWLN-8 will be the shifting role of the lexicographer in the age of artificial intelligence (AI). While lexicography has always been an early adopter of technological advances in linguistics, corpus analysis, and dictionary production workflows, recent developments in large and custom language models and generative AI have fundamentally altered the ecology of lexical innovation, its description, and the possibilities for lexicographers. Neologisms now emerge and circulate not only through human linguistic communities, but also through AI-mediated content creation, algorithmic feedback loops, and hybrid forms of human–machine discourse. At the same time, AI-based technologies create new opportunities for lexicographers to automate detection pipelines, evaluate usage patterns, and compile entries grounded in different types of linguistic evidence and linked to highly structured lexical knowledge bases. Against this backdrop, the 2026 workshop will invite participants to reflect on lexicography’s mediating role between corpus-based empirical evidence, AI-generated content, and the need for reliable, curated and structured lexicographic data that can feed back into and support the emerging hybrid ecosystem of AI-enhanced human communication.
Besides publication in conference proceedings, GWLN papers have formed Special Issues of Dictionaries. Journal of the Dictionary Society of North America (2020), International Journal of Lexicography (2021), Lexicographica. Series Maior (2022), Lexicography. Journal of Asialex (2023), and a special section of Lexikos (2025). Selected GWLN-8 papers will be published in the form of a special issue of International Journal of Lexicography in 2027.
-
4
-
11:00
Coffee Break 00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna -
Pre-Conference Workshop: Constructicography in Practice: Constructions, Challenges, and Opportunities 02 | Museumszimmer (02 | Museumszimmer, OeAW Main Seat, 2nd floor)
02 | Museumszimmer
02 | Museumszimmer, OeAW Main Seat, 2nd floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConveners: Voula Giouli (Aristotle University of Thessaloniki), Athina Sioupi (Aristotle University of Thessaloniki), Nina Böbel (Heinrich Heine University Düsseldorf), Alexander Ziem (Heinrich Heine University Düsseldorf)-
5
Constructicography in Practice: Constructions, Challenges, and Opportunities
Constructicography (Lyngfelt et al. 2008), the systematic description of constructions in Constructicons and construction-based lexicographic resources, has emerged as a key area at the intersection of Construction Grammar (i.e., Croft 2001), lexicography, and computational linguistics. While constructions have long been acknowledged as central units of linguistic knowledge, their consistent representation, annotation, and integration into lexicographic resources remain methodologically and technologically challenging.
The proposed workshop aims to bring together researchers working on theoretical, descriptive, and computational aspects of constructicography. It will focus on how constructions are identified, modeled, represented, and linked to lexical, semantic, and pragmatic information across languages and resource types. Particular attention will be paid to challenges such as construction granularity, variation, productivity, cross-linguistic comparability, corpus-driven discovery, and interoperability with existing lexical resources.
At the same time, the workshop highlights emerging opportunities offered by large corpora, annotation frameworks, and NLP techniques, including construction mining, semi-automatic Constructicon building, and the use of large language models in constructional analysis. Moreover, the workshop will explore applications of Constructicons, particularly in language teaching and learning.
By fostering dialogue between linguists, lexicographers, educational experts and computational researchers, the workshop seeks to advance constructicography as a mature and practically applicable field within lexicography.
-
5
-
Pre-Conference Workshop: Multi-Word Expressions and Phraseology: Corpus-Based and Computer-Processed 00 | Anton Zeilinger Salon (00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor)
00 | Anton Zeilinger Salon
00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConveners: Lian Chen (LLL, University of Orléans), Besim Kabashi (Eberhard Karls Universität Tübingen)-
6
Multi-Word Expressions and Phraseology: Corpus-Based and Computer-Processed
Phraseology (Cowie 1998; Granger & Meunier 2008; Mitkov 2017; Mel’čuk 2023; Polguère, 2002, 2014; Mejri 2018; Chen 2021), or multiword expressions (MWEs) (Savary 2008; Constant 2012) occupy a central position in linguistic description, language use, and language learning. Idioms, collocations, lexical bundles, and fixed or semi-fixed expressions constitute a significant part of natural language, yet they remain challenging to model, describe, and represent in lexicographic resources (Pecina, P. 2010; Mel’čuk 2011, Polguère 2014, Chen 2025). With the rapid development of new technologies—particularly corpus linguistics, Natural Language Processing (NLP), and Artificial Intelligence (AI)—phraseology and lexicography are currently undergoing a profound methodological and conceptual transformation.
From a modeling perspective, new technologies make it possible to move beyond traditional, intuition-based descriptions of phraseological units. Large-scale corpora (Mitkov 2017), both general and specialized, enable the systematic identification of MWEs through statistical, distributional, and syntactic approaches. Methods such as n-gram extraction, association measures (PMI, t-score), syntactic patterning, and embedding-based similarity allow researchers to capture degrees of fixedness, semantic compositionality, and contextual variability. These advances open new perspectives for representing phraseological knowledge in structured models, including ontologies, lexical networks, and standards such as OntoLex-Lemon (McCrae, Bosque-Gil et al. 2017 ; Bosque-Gil et al. 2019).
In terms of resource creation, technological tools have profoundly reshaped lexicographic practices (Atkins & Rundell 2008; Granger & Paquot 2012). Digital corpora, web-based data, and annotation platforms facilitate the semi-automatic extraction and validation of phraseological units (Mitkov 2017; Evert 2008; Gries 2008; Constant et al. 2017). New-generation lexical resources integrate rich metadata, usage examples, frequency information, and semantic relations, making them more dynamic and interoperable (Polguère 2014; Ci-miano et al. 2016; McCrae et al. 2017).. Moreover, computational approaches enable the development of multilingual and contrastive resources, which are essential for studying phraseological variation across languages and cultures (Paquot 2015; Mel’čuk 2011).
New technologies also play a crucial role in the design of pedagogical dictionaries and lear-ning-oriented resources (Bogaards & van der Kloot 2002; Lew 2012). Phraseology is often a major obstacle for language learners, as MWEs cannot always be interpreted compositionally (Wray 2002; Howarth 1998; Granger 1998). Corpus-based examples, learner-oriented definitions, and adaptive digital interfaces can significantly enhance the accessibility and usability of phraseological information (Sinclair 2004; Granger & Paquot 2015; Paquot 2015). AI-driven tools, such as intelligent tutoring systems or LLM-assisted lexicography, offer promising avenues for generating contextualized examples, usage notes, and difficulty pro-files tailored to learners’ needs (Heift & Schulze 2007; Godwin-Jones 2023; Bender & Koller 2020).
In translation lexicography, the impact of new technologies is equally significant (Hartmann 2007; Atkins & Rundell 2008; Tarp 2012). Phraseological units pose well-known challenges in translation due to their idiomaticity and phraseocultural specificity (Chen 2022a; 2022b). Parallel corpora, alignment tools, and machine translation systems provide valuable data for identifying translation equivalents, partial correspondences, and translation strategies (Chen et al. 2024). Digital translation dictionaries can now incorporate cross-linguistic mappings, semantic annotations, and real attested examples, bridging the gap between lexicographic description and actual translational practice.
Finally, the analysis of existing dictionaries through technological means (Béjoint 2010; Lew 2013) offers new insights into lexicographic traditions and practices. Computational methods allow for the large-scale comparison of dictionary entries, coverage, microstructure, and treatment of phraseological units. Such analyses contribute to a critical understanding of how dictionaries evolve in response to technological, linguistic, and societal changes (Hausmann 1989; Tarp 2012; Fuertes-Olivera & Bergenholtz 2011).
The intersection of phraseology, multi-word expressions, lexicography, and new technologies constitutes a fertile research domain, fostering innovative models, richer resources, and more effective tools for analysis, learning, and translation.
-
6
-
Pre-Conference Workshop: Postediting Lexicography with Sketch Engine and Lexonomy 01 | Sitzungssaal (01 | Sitzungssaal, OeAW Main Seat, 1st floor)
01 | Sitzungssaal
01 | Sitzungssaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Ondřej Matuška (Lexical Computing)-
7
Postediting Lexicography with Sketch Engine and Lexonomy
The tutorial aims to promote the understanding of state-of-the-art lexicographic practices based on the post-editing of corpus-generated and AI-generated content. It focuses on how contemporary lexicography can effectively combine automation with expert human judgement to achieve efficient, scalable, and cost-effective editorial workflows while ensuring that all dictionary data remain fully reviewed and validated by human editors.
The workshop provides an overview of current corpus and AI technologies relevant to dictionary making. Sketch Engine ( https://www.sketchengine.eu/ ) is exploited to support key lexicographic tasks such as sense analysis, collocation extraction, example selection, and pattern discovery. The session will also introduce Lexonomy ( https://guide.lexonomy.eu/ ) as a modern dictionary writing system designed for updating legacy dictionaries as well as creating new lexicographic works. Lexonomy demonstrates how corpus evidence and post-edited content can be efficiently structured, edited, and published within a collaborative editorial environment.
A central component of the tutorial is the presentation of the Dictionary Express method ( https://dictionary.express/ ), which combines automated content generation with systematic post-editing to accelerate dictionary production without compromising quality. Through practical workflows and concrete examples, the tutorial will show how corpus data, AI-assisted suggestions, and editorial guidelines can be integrated into a coherent process that supports consistency, transparency, and editorial control. The session is intended for lexicographers, editors, and researchers interested in applying corpus-based and AI-supported methods to contemporary dictionary projects.
-
7
-
Pre-Conference Workshop: The 8th Globalex Workshop on Lexicography and Neology (GWLN-8) 01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)
01 | Johannessaal
01 | Johannessaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConveners: Ilan Kernerman (Lexicala by K Dictionaries), Kris Heylen (Dutch Language Institute (INT))-
8
The 8th Globalex Workshop on Lexicography and Neology (GWLN-8)
In recent years, the Globalex Workshop on Lexicography and Neology series (GWLN) has evolved into an international forum for lexicographers, scholars, tool developers, and other practitioners interested in how new words and new senses emerge, are detected, are evaluated as candidates for inclusion, and described in lexicographic resources. Since its inauguration, the workshop has taken place in seven consecutive years in conjunction with annual conferences of the continental lexicography associations, including DSNA (2019), Euralex (2020 online, and 2022), Australex (2021), Asialex (2023), Afrilex (2024), as well as at eLex (2025). Each edition has brought together participants from diverse linguistic traditions – covering both high- and low-resource languages – to share their experience and exchange perspectives on the challenges, methodologies and processes for combining lexicography and neology. For the eighth iteration, GWLN is teaming up with the European Network on Lexical Innovation (ENEOLI COST Action) to further strengthen collaboration in Europe and beyond.
The main topic of GWLN-8 will be the shifting role of the lexicographer in the age of artificial intelligence (AI). While lexicography has always been an early adopter of technological advances in linguistics, corpus analysis, and dictionary production workflows, recent developments in large and custom language models and generative AI have fundamentally altered the ecology of lexical innovation, its description, and the possibilities for lexicographers. Neologisms now emerge and circulate not only through human linguistic communities, but also through AI-mediated content creation, algorithmic feedback loops, and hybrid forms of human–machine discourse. At the same time, AI-based technologies create new opportunities for lexicographers to automate detection pipelines, evaluate usage patterns, and compile entries grounded in different types of linguistic evidence and linked to highly structured lexical knowledge bases. Against this backdrop, the 2026 workshop will invite participants to reflect on lexicography’s mediating role between corpus-based empirical evidence, AI-generated content, and the need for reliable, curated and structured lexicographic data that can feed back into and support the emerging hybrid ecosystem of AI-enhanced human communication.
Besides publication in conference proceedings, GWLN papers have formed Special Issues of Dictionaries. Journal of the Dictionary Society of North America (2020), International Journal of Lexicography (2021), Lexicographica. Series Maior (2022), Lexicography. Journal of Asialex (2023), and a special section of Lexikos (2025). Selected GWLN-8 papers will be published in the form of a special issue of International Journal of Lexicography in 2027.
-
8
-
9
Registration Clubraum (00 | Clubraum (Club Room), OeAW Main Seat, ground floor)
Clubraum
00 | Clubraum (Club Room), OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz-Seipel-Platz 2 1010 Vienna, Austria -
10
Opening & Welcome 01 | Festsaal (Festive Hall), OeAW Main Seat, 1st floor
01 | Festsaal (Festive Hall), OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna -
Keynote: With Bots Like These, Who Needs Dictionaries? 01 | Festsaal (Festive Hall), OeAW Main Seat, 1st floor
01 | Festsaal (Festive Hall), OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Robert Lew (Adam Mickiewicz University, Poznań)-
11
With Bots Like These, Who Needs Dictionaries?
Four years after the public launch of ChatGPT, Generative AI has entered many domains, both private and professional, and it is not going away. As anticipated in a prescient presentation at eLex 2019 (The Sintra variations), AI bots have begun to compete with dictionaries as reference tools for language-related tasks. Yet despite the pervasiveness of this emerging technology, dictionaries remain important consultation tools for certain types of language problems, at least for advanced learners, who appear able to choose critically between dictionaries and AI tools depending on the task at hand. At the same time, recent studies comparing the effectiveness of chatbots and dictionaries suggest that AI chatbots are not necessarily more effective than dictionaries—especially bilingual dictionaries—whether measured in terms of immediate success or longer-term learning. Against this empirical backdrop, the talk will sketch the current and near-future roles of dictionaries and AI tools.
Speaker: Robert Lew (Adam Mickiewicz University, Poznań)
-
11
-
16:00
Coffee Break 00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna -
Computational & AI-Based Lexicography 01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)
01 | Johannessaal
01 | Johannessaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Jelena Kallas (Institute of the Estonian Langauge)-
12
What LLMs Learn and What They Miss: Frequency Effects in the Automated Description of Korean Verbal Inflection
The present study aims to empirically examine the possibilities and limitations of LLMs as a lexicographic tool for Korean, by automatically generating inflected forms of Korean verb and adjective headwords using Large Language Models and contrasting the results with corpus frequency data. Focusing on the description of inflected forms of Korean verbs and adjectives, the study analyses in detail the extent to which LLMs reproduce corpus-based frequency rankings and the accuracy of inflected form generation, divided into the categories of regular, irregular, and defective predicates. The results show that LLMs consistently produce various errors in low-frequency inflected forms and reveal clear limitations in precisely reproducing frequency information. This empirical examination points to the possibilities and risks inherent in LLMs as lexicographic assistance tools from multiple angles. While LLMs hold meaningful value as an assistive means for drafting dictionary descriptions, lexicographic reliability can only be guaranteed when this is complemented by systematic corpus-based verification and the distinctive expertise of lexicographers. These findings reaffirm the fundamental role that lexicographers must play in the automation of dictionary compilation.
Speakers: Li Cui (Yonsei University), Seol Namkung (Yonsei University), Jun Lee (Yonsei University), Hae-Yun Jung (Kyungpook National University), Kilim Nam (Yonsei University) -
13
Chatbot Lexicography: Are GenAI LLMs Simply ‘Dictionaries’ In Their Own Right and ‘Lexicographers’ Plain and Simple?
At the time of the Euralex 2026 conference it will be nearly four years since the start of the GenAI revolution. The question that needs to be asked, and which will be answered, is thus: Is any of the past work in (academic) lexicography still of any relevance, or are LLM chatbots simply ‘dictionaries’ in their own right and ‘lexicographers’ plain and simple? The answer is a shocking one.
Speaker: Gilles-Maurice de Schryver (Ghent University) -
14
Lexicographic Resources for Inference-Based Semantic Text Processing System
This paper describes the lexicographic resources developed for the ETAP text analysis and generation system, with particular emphasis on its semantic module, SemETAP. In our approach, semantic analysis is viewed not merely as the construction of a semantic representation, but as the derivation of inferences licensed by linguistic and background knowledge. The semantic component operates in two stages. First, it constructs a Basic Semantic Structure (BSemS), which represents the meaning explicitly conveyed by the text. This representation is then enriched through ontological knowledge and inference rules, yielding an Extended Semantic Structure (EnSemS) that incorporates both explicit and implicit information. To support this process, ETAP relies on a set of large-scale, knowledge-rich lexicographic resources characterized by a high degree of reusability. Linguistic knowledge is distributed across a morphological dictionary, a combinatorial dictionary, and rule-based grammars, whereas background knowledge is represented in an ontology, including a repository of individuals, and a collection of inference rules. The paper describes the information encoded in these resources, shows formal languages used to represent it, and the interaction between linguistic and extralinguistic knowledge during semantic analysis.
Speakers: Igor Boguslavsky (IA.A.Kharkevich Institute for Information Transmission Problems), Vyacheslav Dikonov (A.A.Kharkevich Institute for Information Transmission Problems), Evgeniya Inshakova (A.A.Kharkevich Institute for Information Transmission Problems), Alexandre Lazursky (A.A.Kharkevich Institute for Information Transmission Problems), Svetlana Timoshenko (A. A. Kharkevich Institute for Information Transmission Problems), Tatiana Frolova (A.A.Kharkevich Institute for Information Transmission Problems)
-
12
-
Digital Lexicography: Design, Data Modelling & Theory 01 | Sitzungssaal (01 | Sitzungssaal, OeAW Main Seat, 1st floor)
01 | Sitzungssaal
01 | Sitzungssaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Ivana Filipović Petrović (Croatian Academy of Sciences and Arts)-
15
From Corpus to Dictionary: OntoLex-Lemon-based Modelling of Neology in Chinese–French and Vietnamese–French Computational Lexicography
This paper presents NeoLex, a multilingual electronic lexical resource dedicated to contemporary Chinese–French and Vietnamese–French neology. Based on diachronic media corpora (2015–2025) and pre-2015 reference lexicons, the project combines corpus linguistics, Natural Language Processing, and computational lexicography to detect emerging lexical units. Candidate neologisms were extracted using Out-of-Vocabulary detection, Vietnamese n-gram extraction, and a character-based method for Chinese, then manually validated according to lexicographic criteria. The validated entries were modelled using OntoLex-Lemon, SKOS, and vartrans to represent lexical entries, lexical senses, translation equivalents between Chinese–French and Vietnamese–French, semantic domains, and diachronic metadata. A Streamlit-based prototype was developed to explore and visualize the resulting resources. NeoLex demonstrates how corpus-based neology detection and ontology-driven lexicography can contribute to interoperable multilingual electronic dictionaries.
Speakers: Lian Chen (LLL, University of Orléans), Damien Nouvel (ERTILM), Huy-Linh Dao (CRLAO-CNRS), Alexander Delaporte (CNRS-CRLAO) -
16
Bridging Lexicography and NLP: The Architecture and Application of the Digital Dictionary Database of Slovene (DDDS)
In the era of data-driven linguistics, the transition from traditional, human-oriented lexicography to machine-readable and interoperable language resources is paramount. The Digital Dictionary Database of Slovene (DDDS), developed by the Centre for Language Resources and Technologies (CJVT) at the University of Ljubljana, represents a paradigm shift in how national lexicographical data is curated, stored, and disseminated. Rather than treating dictionaries as isolated digital products, the DDDS functions as a centralized, relational infrastructure designed to support a multi-layered ecosystem of language resources—ranging from the Thesaurus of Modern Slovene to the Collocations Dictionary and morphological database Sloleks. This paper will describe the structural design, the underlying data model, and the multifaceted API-based application of the DDDS, highlighting its role in the "responsive dictionary" framework.
At its core, the DDDS is built on a PostgreSQL database managed via the Django Object-Relational Mapping (ORM) framework. Unlike traditional lexicographical databases that often rely on XML hierarchies, the DDDS employs a highly normalized relational model that prioritizes the atomization of linguistic data. The model is organized into several clusters, with the primary entities being Lexical Units and Senses. A key innovation in the DDDS is its approach to "Lexemes," which correspond to tokens in a corpus but are treated as distinct structural components within the database. This allows for the precise representation of complex multi-word expressions, compounds, and syntactically varied forms. By utilizing a single primary key system and Django’s "content types" and "generic relations," the database maintains a flexible architecture where additional metadata—such as dictionary-specific labels or administrative information—can be associated with any table without necessitating disruptive schema changes. This modularity ensures that the database remains a "living" resource capable of accommodating various lexicographical workflows.
The DDDS serves as the "backbone" for a suite of modern Slovene resources. Most notably, it powers the Thesaurus of Modern Slovene, which introduced the concept of the responsive dictionary. In this model, the dictionary is not a static text but an evolving dataset where automatically generated candidates are continuously refined through expert validation and crowdsourcing. The database manages these dynamic relationships, tracking synonymy, sense disambiguation, and collocation frequency extracted from the Gigafida reference corpus.
A critical component of the DDDS project is its commitment to the FAIR (Findable, Accessible, Interoperable, and Reusable) principles. This is achieved through a comprehensive REST API that allows external developers, researchers, and NLP systems to interact with the database in real-time. The API design distinguishes between read-only public routes and restricted read-write routes.
A significant modern application of the DDDS is its role in the PoVeJMo project, which aims to optimize Large Language Models for the Slovene language. By providing high-quality, structured grammatical and lexical data, the DDDS enables models to learn the nuances of Slovene grammar and semantics more effectively than they would from unstructured web-crawl data alone. The extraction of sense-specific examples and sense-relation mappings from the database provides a "gold standard" dataset for training models in tasks such as Natural Language Inference (NLI) and Word Sense Disambiguation (WSD).
The development of the DDDS highlights a shift toward a "digital-first" lexicographic workflow. Traditional dictionary writing involved drafting entries that were later converted to digital formats. In the DDDS framework, the database is the dictionary. The interface used by lexicographers—a new dedicated online editor—is integrated with the database, ensuring that updates are immediately reflected in the underlying data structure. The Digital Dictionary Database of Slovene demonstrates that a central, well-structured database is essential for the survival and growth of a national language in the digital age. By decoupling the linguistic data from the final presentation layer (the web interface), CJVT aims to create a robust infrastructure that supports both human users and automated systems. Future developments involve the expansion of the data model to include more granular semantic relations and the continued integration of crowdsourced feedback, ensuring that the DDDS remains the authoritative, yet dynamic, reference for the modern Slovene language.
In the final paper we will describe the specialized data model with REST API, the online editor, and various online visualizations of DDDS data.
Speakers: Simon Krek (Jožef Stefan Institute), Iztok Kosem (University of Ljubljana & Jožef Stefan Institute), Polona Gantar (Faculty of Arts, University of Ljubljana) -
17
Revising the Usage-Labelling System of The Danish Dictionary
This article discusses the ongoing revision of the usage-labelling system in Den Danske Ordbog (DDO, ‘The Danish Dictionary’) in response to increasing public attention on discriminatory and otherwise socially marked language. In recent years, the dictionary has received growing criticism concerning the treatment of identity-related terms, revealing limitations in the existing inventory of stylistic and evaluative labels. The article examines the lexicographic challenges involved in balancing descriptive dictionary practice with users’ expectations for guidance and contextualisation. To address these issues, DDO is developing a revised system based on a closed inventory of evaluative labels, optional modifiers, and controlled combinations. The new model is supplemented by graphical warning symbols and expandable usage notes inspired by practices in dictionaries such as Merriam-Webster and Duden. The proposed revision seeks to establish a compromise between systematic description, user guidance, and the scientific legitimacy of descriptive lexicography.
Speakers: Lars Trap-Jensen (Society for Danish Language and Literature), Henrik Lorentzen (Society for Danish Language and Literature)
-
15
-
Learner's Lexicography & Dictionary Use 00 | Anton Zeilinger Salon (00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor)
00 | Anton Zeilinger Salon
00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Ivana Brač (Institute for the Croatian Language)-
18
Directions in Lexicography in the Global South
The AS Hornby Dictionary Research Awards (ASHDRA) have been running since 2019 and to date have received 160 applications from 47 countries. The awards fund dictionary research within English language teaching (ELT) and especially target projects in low-resource contexts. Thus, a large proportion of applications (although not all) are from researchers working in the Global South. This paper reviews three key trends that have emerged across those applications and suggests areas of particular focus in lexicography in these regions, with examples of ASHDRA-funded projects in each area.
Multilingual resources including local languages
The most significant trend to emerge among ASHDRA research proposals has been interest in bi- and multilingual reference resources for minority, local languages. Many countries of the Global South have a multi-layered linguistic patchwork in which people speak a local language at home as their L1, then learn a regional or national language as their L2 and lingua franca. English (effectively their L3) is commonly taught from primary school but typically via the L2. In one such project, Riadi (2021) created a multilingual dictionary for learners in schools in an underdeveloped region of West Borneo, Indonesia. The small-scale resource was based around photographs of the linguistic landscape (adverts, shop signs, etc.) near to the participating schools. The entry for each target word included equivalents in English, Bahasa Indonesia (the national language), Pontianak (a regional language), and Ketapang (a local language and most students' mother tongue). The resulting dictionary provided both a classroom resource enabling students to translate directly between English and their home language, something they had not previously done (or been encouraged to do) and also linked the English they were learning to their environment and lived experiences, something that classroom observation and feedback showed they found incredibly motivating.Other projects (proposed and ongoing) have focused on communities in rural India who speak a local language at home, then learn English at school through the medium of a regional language, such as Kannada, or indigenous students in Mexico who speak Zapotec, a largely oral local language, as their home language but are schooled in Spanish, then are expected to learn English. In some cases, schoolteachers do not come from the local communities and so do not even share their students' L1. This puts those learners, who may already come from disadvantaged communities, at a further disadvantage. With an increasing interest in 'translanguaging' and the benefits of using learners' L1 as a resource in language teaching, some national Ministries of Education are encouraging the use of L1 in the ELT classroom. Such an approach is hard to implement, however, without relevant resources for both learners and teachers. Around 25% of ASHDRA applications have been to develop multilingual resources for contexts where none yet exist. These projects often come up against the further barrier that local languages may be largely oral and even where written forms exist, there may be low levels of literacy within the community. This has led to creative solutions such as the use of images and audio (see more below).
English for (very) Specific Purposes
English for Specific Purposes (ESP) is a well-established area of lexicography, but many published resources, such as dictionaries of Medical English or English for Engineering, are quite broad in scope and often geared toward typical students and teaching contexts in countries in the Global North. Despite the massive growth in English as a medium of instruction (EMI) in higher education globally, the differing needs of learners in the Global South remain under-researched (Sahan et al., 2021) and existing ESP resources may not always be relevant or accessible to learners in low-resource, sometimes low-bandwidth, contexts with specific English needs. Another theme among ASHDRA proposals has been the development of very specific resources to target the needs of particular groups of English learners, such as a dictionary of semi-technical vocabulary aimed at Vietnamese medical students (Le, 2022), an app for Omani physics undergraduates (Mathew & McCallum, 2024), or a resource for Indonesian mining engineering students learning to work with the industry standard software in English (Nurfitriah, forthcoming). With limited resources available, these projects are not necessarily full ESP dictionaries but resources with a scope and format targeted at the specific and immediate needs of learners.Inclusive lexicographic resources
The final theme to emerge is the development of accessible resources for language learners who cannot easily access conventional dictionaries. This group includes learners with visual and hearing impairments. Two research projects due for completion in 2026 target, respectively, visually impaired students in Indonesian primary schools with a dictionary that includes both Braille and QR codes linked to audio (Karolina, forthcoming), and a project using sign language videos aimed at deaf learners in India (Zeshan, forthcoming). This category also includes learners with low levels of literacy. This overlaps with some of the projects under the first theme where local languages may be from a largely oral tradition, but also adult migrant and refugee learners who sometimes have low levels of literacy in their first language and who can more easily access English language resources with visual and/or audio components (such as O’Boyle, 2023, Oakey & Jones, forthcoming).Trends: focusing on the local and specific
What the most innovative projects in all three areas have in common is that they take as their starting point the needs of learners in a specific context and often work in collaboration with learners and educators to develop a lexicographic resource that directly addresses those needs, working with limited resources in creative ways. While it has to be acknowledged that ASHDRA have a specific focus and the proposals received reflect this, it seems that they may indicate a more general trend in lexicography and the projects they fund may provide models which could be replicated elsewhere.Speaker: Julie Moore (AS Hornby Trust) -
19
How German Language Textbooks Address the Digital Transformation of Lexicography
The digital transformation of lexicography has fundamentally changed lexicographic resources available for language learning and teaching. However, their integration into German language teaching – both as a first language (GL1) and as a foreign language (GFL) – remains limited. This paper investigates how the digital transformation of lexicography is reflected in contemporary German language textbooks through a qualitative and quantitative analysis (GFL: 2011–2024; GL1: 2003–2021). Drawing on a comparative textbook corpus, the study combines quantitative frequency analysis with qualitative content analysis of textbook tasks and their didactic functions. The findings show that lexicographic resources are underrepresented, predominantly conceptualised in traditional (print-based) ways, and rarely embedded in meaningful communicative or strategy-oriented tasks. Online lexicographic information systems, hybrid resources, and AI-supported language tools are almost entirely absent from the analysed textbooks. Furthermore, the development of lexicographic competence and media literacy is largely neglected. The paper argues that textbooks do not adequately reflect the digital transformation of lexicography and calls for a reconceptualisation of dictionary didactics that systematically integrates contemporary lexicographic information systems while also preparing learners to use AI-supported language tools critically and appropriately. In doing so, it contributes to current discussions on the digital transformation of lexicography and its implications for language education.
Speakers: Christine Ganslmayer (Friedrich-Alexander University of Erlangen–Nuremberg), Martina Nied Curcio (Università degli Studi Roma Tre) -
20
Types of Definitions in Croatian Primary School Science Textbooks
From the first grades of primary education, children encounter abstract and complex concepts, many of which cannot be adequately represented by traditional or conventionally formulated definitions. Consequently, younger learners may struggle with understanding abstract concepts like energy, community, or climate. Previous studies have shown that children’s ability to engage with scientific and abstract concepts is closely related to the form in which these concepts are presented, including sentence structure, vocabulary choice, and the type of context provided (Murray & Reuter, 2005; Gelman & Meyer, 2011). Research on concept acquisition further suggests that abstract concepts are acquired gradually and are initially grounded in emotionally relevant experiences and situations, particularly in early school years, age 6 to 9 (Caramelli, Setti & Maurizzi, 2004). Contemporary educational resources already reflect these findings, which is evident in different types of presentation of complex concepts in textbooks, particularly in the domain of science education, where the interdisciplinarity of concepts is more prominent than in other school subjects.
The aim of this study is to examine the types of definitions in Croatian primary school science textbooks, Grades 3 to 8, in order to analyze how key concepts are systematically presented to different age level learners. The data includes definitions and definitional explanations extracted from textbooks for the subjects Nature and Society (Grades 3 and 4), Nature (Grades 5 and 6), and Biology (Grades 7 and 8). The definitions analysed are part of a larger dataset of textbook definitions extracted for the purpose of creating an online, learner-oriented dictionary of scientific concepts for primary education. All textbooks have been officially approved and widely used in Croatian primary education. At the present stage, textbooks published by two major publishers, Školska knjiga and Alfa, are analysed. Textbooks published by a third approved publisher, Profil Klett, will be incorporated in the second stage of the study.
From each textbook, definitions or explanations of concepts highlighted in bold font were manually extracted. Currently, a total of 654 definitions and/or explanations have been independently annotated by two annotators trained in terminology management. The data has been annotated for a superordinate concept, if stated, salient concept characteristics, and conceptual relations. Inter-annotator agreement has not been measured yet, but a preliminary qualitative analysis based on 59 concepts, 45 definitions and 21 explanations (Ostroški Anić & Pavić, in press) showed an extremely high agreement of 97% in the annotation of superordinate concepts. A much lower agreement was found for identifying concept characteristics, 68%. The causes for lower agreement were identified to be avoided in further annotation process.
For this paper, definitions of the same concept were compared vertically across grade levels in order to identify differences in the type of definition (e.g., intensional, partitive, enumerative, descriptive, functional, etc.) and in concept characteristics (e.g., origin, composition, form, function, purpose, location, etc.). Qualitative analysis shows that concrete core science concepts like water and organism share rather a uniform approach to presenting the concepts in textbooks, resulting in widely accepted definitions, e.g. Water is a liquid without colour, taste, or smell – a variant of which was found in four out of six textbooks for Grades 3 and 4. All six textbooks refer to the states of water, either within the definition itself or in accompanying explanations. The process of enriching the conceptual knowledge is evident from a Grade 5 definition: Water is a good solvent for gases and mineral substances, whereas a Grade 7 explanation of water further expands the scope of concept characteristics by placing water within the context of human physiological processes.
Variation in definitional structures observed during the analysis shows clear vertical enrichment of concept-related information presented in definitions, which mostly follows the cognitive development primary school children. Discrepancies, however, do exist in some textbooks, where graphic information is the primary source of understanding, and it is not always matched with appropriate linguistic context. The results of this study will be used for developing a prototype of an online, learner-oriented science dictionary for primary school students, which will offer multiple definitions aligned to educational level and conceptual depth. A resource of this kind would support learners’ transition from experiential knowledge toward scientific conceptual network in line with similar approaches to knowledge description, e.g. the idea of flexible definitions (San Martín, 2022).
Speakers: Ana Ostroški Anić (Institute for the Croatian Language), Martina Pavić (Institute for the Croatian Language), Lidija Cvikić (University of Zagreb, Faculty of Teacher Education)
-
18
-
21
Reception 00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz-Seipel-Platz 2 1010 Vienna, Austria
-
-
-
Historical & Diachronic Lexicography 01 | Sitzungssaal (OeAW Main Seat, 1st floor)
01 | Sitzungssaal
OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Geraint Paul Rees (Pompeu Fabra University)-
22
Gender Representation and Editorial Responsibility in the OED: A Case Study of female, feminine, male and masculine
As a dictionary grounded in diachronic evidence, the Oxford English Dictionary (OED) occupies a central position in the tradition of historical lexicography, documenting not only the evolution of lexical meaning but also the cultural assumptions embedded in language over time. While the OED is explicitly descriptive and corpus-based, scholarship has shown that definitional framing and citation selection are shaped by editorial choices and socio-cultural ideologies (Baigent, Brewer & Larminie, 2005; Mugglestone, 2013; Norri, 2000; Vișan, 2021; Williams, 2023). Against this backdrop, this paper examines how gender is constructed and transmitted in the OED through a case study of the lemmas female, feminine, male, and masculine. The aim is to consider whether and to what extent gender representation has changed in the dictionary in light of the OED's continuous revision process. More specifically, it assesses whether sustained editorial updating has substantially altered patterns of gender representation previously identified in lexicographic research.
These four lemmas were selected because of their apparent semantic neutrality and their foundational role in structuring gender categories in English, in contrast to more overtly evaluative terms such as womanly, womanish, manly, or mannish. The study adopts a qualitative, interpretive methodology, using present-day definitions as a point of semantic orientation while placing particular emphasis on the diachronic evidence provided by illustrative citations. All data are drawn from the OED Online, Third Edition (June 2025 update). Senses referring to human gender were analysed in full, while technical or grammatical uses were excluded unless their figurative extensions revealed ideological relevance. Particular attention is paid to citations from the nineteenth century onwards, a period of dense documentary coverage in the OED and of major ideological consolidation in relation to gender.
The analysis reveals persistent asymmetries in the representation of femininity and masculinity in the OED. Entries for female and feminine are more extensive and semantically elaborated than those for male and masculine, reflecting the marked status of femininity within the dictionary’s treatment of gender. Across successive historical layers of citation, femininity is recurrently associated with emotional excess, passivity, superficiality, moral suspicion, and aesthetic or decorative qualities. Illustrative quotations frequently situate women in trivialised, domestic, or subordinate roles, or portray feminine agency as indirect or manipulative. Symbolic and metaphorical extensions—linking the feminine to flowers, planets, beauty products, or delicacy—further embed culturally inherited stereotypes within the historical record.
Masculinity, by contrast, is largely constructed as the unmarked and normative category. While male often functions as a default biological descriptor (as also in the case of female), masculine is consistently associated with strength, activity, authority, and intellectual or moral superiority. These qualities are rarely problematised in the historical citations and are frequently extended metaphorically to domains such as architecture, astrology, literature, and social organisation, reflecting long-standing models of hegemonic masculinity (Connell & Messerschmidt, 2005). When masculine traits are attributed to women, they tend to be evaluated positively, whereas feminine traits applied to men are typically framed as deviations—an asymmetry well documented also in feminist linguistic scholarship (Kramarae, Treichler & Russo, 1985).
Although more recent citations in the OED begin to reflect changing conceptions of gender, including gender fluidity and non-binary identities, these developments remain unevenly distributed across the entries. Many senses of male and masculine continue to rely on relatively dated historical material, allowing older ideological models to dominate the semantic profile of these lemmas. As a result, the historical layering of citations becomes a mechanism through which outdated representations are preserved and reactivated in contemporary dictionary use.
The paper argues that historical dictionaries do not merely record linguistic change but actively mediate it through editorial selection and presentation (Pullum, 2018). By foregrounding the persistence of gendered asymmetries within a continuously revised dictionary, and by showing how definitions and citation practices can perpetuate gender hierarchies even in ostensibly neutral entries, this study contributes to current debates in historical lexicography and metalexicography. It concludes by suggesting that, unless definitions and citations are more carefully contextualised and balanced with contemporary evidence, even continuous editorial revision may do little to alter entrenched patterns of gender representation. A more reflexive approach would allow authoritative dictionaries such as the OED to preserve their historical record while mitigating the continued circulation of inherited gender ideologies.
Speaker: Daniele Franceschi (Roma Tre University) -
23
Can AI-Assisted Normalisation Improve Access to Historical Lexicographic Data?
The topic of this article is AI-assisted normalisation and how it can be used in the context of a historical dictionary. We look at data from A Dictionary of Old Norse Prose, which is a lexicographic resource on the language of medieval Iceland and Norway. This dictionary has a long tradition of using non-normalised, diplomatic text editions as source for its lexicographic descriptions. As a result, the bulk of the citation material remains difficult to access for beginners and non-experts in the language. We tested several LLMs in order to assess whether it would be feasible to provide users with automatic normalisations. The article describes several phases in testing and comparing different LLMs and discussing the types of errors they made. The result show that from the initial experimental phase in 2025 the models have improved significantly and currently are able to provide fairly error-free normalisations. We also mention the option of using modern Icelandic as an intermediate stage in the normalisation process which would facilitate the use of NLP-tools. The conclusion is that LLM-assisted normalisation can support broader access to historical lexicographic data. Such tools combine automated processing with human oversight, allowing historical dictionaries to expand their usability while maintaining scholarly rigour.
Speakers: Tarrin Wills (University of Copenhagen), Ellert Thor Johannsson (The Árni Magnússon Institute for Icelandic Studies) -
24
Artificial Intelligence as Support for the Digitization and Compilation of the Dictionary of the Croatian Redaction of Church Slavonic
The Dictionary of the Croatian Redaction of Church Slavonic (RCJHR) is the first lexicographic description of the Croatian Church Slavonic language. Within the project Development of the Digital Infrastructure Model of the Old Church Slavonic Institute (DigiSTIN), the printed RCJHR is being digitized and adapted for online publication, while the RCJHR project focuses on compiling new entries and exploring a fully digital lexicographic workflow. This paper examines the potential of large language models (LLMs) to support the digitization and compilation of a dictionary documenting a low-resource historical language with an exceptionally complex microstructure involving five languages and four scripts. Six leading models (GPT-5.4 Thinking, DeepSeek-V3.2, Gemini 3 Pro, Claude Opus 4.6, Qwen3.6-Plus, and Copilot GPT-5) were evaluated across four tasks: OCR of previously published fascicles, HTR of handwritten index cards, defining entry structure, and compiling new dictionary entries. The results demonstrate that LLMs significantly accelerate digitization and assist in identifying foreign-language equivalents, suggesting sense divisions, and proposing cross-references. However, they remain unreliable in resolving subtle semantic distinctions and recognising the Glagolitic script, confirming that human philological expertise remains indispensable.
Speakers: Ana Mihaljević (Institute for the Croatian Language), Josip Mihaljević (Institute for the Croatian Language) -
25
Compiling a 16th-Century Croatian Dictionary in the Age of AI: Challenges and Perspectives
This paper presents the theoretical and methodological foundations for an online dictionary of sixteenth-century Croatian. It utilizes a corpus of ten Gospel translations in Glagolitic, Cyrillic, and Latin scripts, en-compassing missals, lectionaries, Bible translations, and postils. This work is a primary objective of the project CroBiLex: Language Variation in Croatian Books of the Early Modern Period (Based on the Gospels Corpus) (Croatian Science Foundation, IP-2025-02-2567; 2026–2029) at the University of Zagreb’s Faculty of Humanities and Social Sciences. The paper outlines a conceptual lexicographic architecture for processing this linguistically heterogeneous, multi-script corpus. It emphasizes a Human-in-the-Loop approach where deterministic processing is complemented by machine-assisted transcription using handwritten text recognition and multimodal models. It also explores the experimental use of large language models (LLMs) for candidate lemma grouping, provisional semantic clustering, and preliminary entry-template preparation. Be-cause highly variable historical texts pose severe risks of historical inaccuracy and hallucination, LLM out-puts are treated strictly as non-authoritative proposals requiring expert verification. As CroBiLex is newly launched, this framework is conceptual rather than empirically evaluated. It seeks to balance computational assistance with scholarly accountability, providing a foundation for a sustainable digital model in historical Croatian lexicography.
Speakers: Petra Bago (University of Zagreb), Ivana Eterović (University of Zagreb)
-
22
-
Lexical Semantics, Neology & Phraseology 01 | Johannessaal (OeAW Main Seat, 1st floor)
01 | Johannessaal
OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Hindrik Sijens (Fryske Akademy)-
26
Phraseological Theoretical Gaps in Digital Lexicography: A Contrastive Study of English–German Phraseologisms with Gender Representations
Dictionaries are widely regarded as authoritative linguistic resources that play a central role in shaping and legitimising language use in society (Haß-Zumkehr 2001). Beyond their descriptive function, they act as social agents influencing public perceptions and norms (Gouws 2022) and must therefore be understood as socially embedded artefacts reflecting linguistic, cultural, and societal values (Müller-Spitzer 2023). This perspective is particularly relevant for gender representation, as lexicographic descriptions of personal referents such as woman and man have repeatedly been shown to reproduce asymmetrical and stereotypical patterns across languages (Hidalgo 2000; Kotthoff & Nübling 2018; Flood 2019; Diewald & Nübling 2022). Corpus-based studies, for instance, reveal systematic gendered asymmetries in adjectival co-occurrences, with men frequently characterised as rich or strong and women as beautiful or pregnant (Müller-Spitzer 2024). Public debates surrounding dictionary revisions, notably those concerning the Oxford Dictionaries (Flood 2019), further underscore the social impact of lexicographic decisions.
While previous research has primarily focused on individual lemmas and their typical co-occurrence patterns, the phraseological level remains comparatively underexplored in studies on gender representation. In particular, the categorisation and lexicographic treatment of idioms, semi-idioms, and collocations have received limited systematic attention (Burger et al. 2007; Burger 2015; Schafroth 2019). Existing analyses largely rely on condensed printed dictionaries, where phraseologisms may be typographically highlighted within single-word entries or occur unmarked in example sentences, complicating their identification and classification (Moon 2007; Steyer 2008; Hollós 2017). Even in comparable dictionaries, the selection, classification, and presentation of phraseological units frequently diverge (Moon 2007).
These challenges are further intensified in digital lexicography, where spatial constraints no longer apply and lexicographic practices vary considerably across platforms and languages (Gouws 2014; 2018; 2022; Klosa-Kückelhaus & Gouws 2015; Klosa-Kückelhaus 2024). Based on a manually compiled corpus drawn from selected German and English online dictionaries, comparable phraseologisms are frequently labelled, defined, positioned, or exemplified differently, or remain entirely unlabeled, directly affecting their identification in contrastive research. A case in point is the German semi-idiom gefallenes Mädchen, which appears as a separate hyperlinked entry under Mehrwortausdrücke in the Digitales Wörterbuch der deutschen Sprache (DWDS), but occurs without any phraseological label in the example section of the Duden Online-Wörterbuch entry Mädchen. Its English equivalent fallen woman is treated equally inconsistently, appearing as a separate entry under Other results in Oxford Learner’s Dictionaries and under More results in Longman English Dictionary, again without explicit phraseological labelling.
Against this background, the present research addresses a theoretical and methodological gap in digital lexicography by proposing a contrastive framework for the identification, classification, and analysis of idioms, semi-idioms, and collocations with gender representations in German and English online dictionaries. The study is based on a manually curated contrastive lexicographic corpus compiled from selected online dictionaries and structured around a predefined set of key lemmas referring to socially defined personal roles, including general designations, family relations, romantic relationships, and social roles. The corpus-building process combines lemma-based extraction with manual phraseological annotation and cross-linguistic alignment. Phraseologisms were identified according to semantic opacity, degree of lexical fixation, and recurrent co-occurrence patterns, while the comparison across dictionaries focused on four analytical dimensions: phraseological labelling, positioning within entries, definitional treatment, and exemplification practices.
By foregrounding these methodological steps, the contribution outlines the compilation of a contrastive lexicographic corpus and demonstrates how inconsistencies in online dictionary structures require interpretative decisions during corpus annotation. It then examines how phraseological units are treated in major German platforms (e.g., DWDS, Duden) and English learner’s dictionaries (e.g., Oxford Learner’s Dictionaries, Longman), highlighting systematic differences in lexicographic practice.
The analysis reveals recurring cross-linguistic and platform-specific divergences in the treatment of phraseological units and gender representation. To address these inconsistencies, the paper proposes a flexible comparative annotation model that distinguishes between the lexicographic treatment of a unit (e.g., labelled or unlabelled, embedded or listed separately) and its linguistic properties as a phraseologism. This approach allows formally comparable units to be analysed despite divergent dictionary structures. The findings demonstrate that collocations pose particular methodological difficulties because their status as phraseologisms often remains implicit and platform-dependent. More broadly, the study argues that contrastive phraseological research in digital lexicography requires transparent annotation criteria and platform-sensitive analytical categories in order to account for the variability of online dictionary practices.
Speaker: Akerke Yessenali (ATILF (CNRS)) -
27
AI-Supported Phraseography for Young Learners: Designing Definitions and Illustrations through AI – Opportunities and Challenges
This paper presents a case study exploring the potential of Large Language Models for learner-oriented phraseography through the generation of COBUILD-style Full-Sentence Definitions and AI-generated dual illustrations representing both the literal and phraseological meanings of idioms. The study is situated within Phrasolino, an electronic Italian-German phraseological learner’s dictionary for German-speaking children aged 7–11 learning Italian as an L2. A corpus-based dataset of 100 Italian idioms containing the body-part noun testa (‘head’) was compiled from the Italian Web Corpus itTenTen20 and complementary lexicographic resources. ChatGPT-5.5 was prompted to generate child-oriented Full-Sentence Definitions and corresponding dual illustrations. The outputs were evaluated qualitatively with regard to semantic accuracy, phraseological adequacy and pedagogical suitability. The findings show that ChatGPT successfully reproduces the formal conventions of COBUILD-style definitions, although recurrent difficulties emerge in the semantic simplification and age-appropriate adaptation of phraseological meaning. The image-generation analysis further reveals both the potential and the limitations of AI-assisted multimodal phraseography, particularly regarding semantic disambiguation, literal interpretation and representational bias. Overall, the study highlights the opportunities of generative AI for learner-oriented lexicography while emphasizing the continuing necessity of lexicographic supervision and post-editing.
Speaker: Linda Prossliner (University of Innsbruck) -
28
A Critical Lexicographical Discourse Analysis of New Adjectival Entries in the Standard Dictionaries for Norwegian
The study presented in the article examines the lemma selection of newly added adjectives in the Norwegian standard dictionaries Bokmålsordboka (BOB) and Nynorskordboka (NOB). Anchored in Critical Lexicographical Discourse Studies (CLDS), the study investigates whether the new adjectives and their sentence examples in the dictionary reflect potential gender-related asymmetries. Methodologically, the new adjectives are classified and analyzed using a socio-semantic category system, combined with concordance analysis of sentence examples to examine patterns of gender representation. The findings show largely comparable patterns across the two dictionaries. There is a clear tendency for new adjectives to be associated with masculine-labelled subcategories versus feminine ones. Furthermore, the sentence examples in the dictionary entries for the new adjectives are predominantly gender-neutral, though a weak tendency towards male overrepresentation is observed, particularly in NOB. Interpreted through a CLDS perspective, the study demonstrates how lexicographic practices related to lemma inclusion may subtly reproduce or cautiously negotiate gendered representations, offering insights relevant to ongoing revision work and lexicographic methodology.
Speaker: Emma Josefin Ölander Aadland (University of Bergen) -
29
Ranking Collocations in Woordcombinaties
This paper describes efforts to extract frequency information from corpora and to rank the collocates in Woordcombinaties, a Dutch phraseological lexicographic resource, in a uniform and systematic way, taking various factors such as the size and composition of the corpus, linguistic processing and the mathematical properties of various association measures into account. A small sample of 2032 combinations, divided over 19 verbs and 19 nouns, was used, and their frequencies were extracted from the Corpus Contemporary Dutch, using surface co-occurrence, textual co-occurrence, and syntactic co-occurrence. Results were verified manually. It was found that: 1) corpus size increases coverage; 2) syntactic co-occurrence leads to loss of coverage; 3) combination of a narrow window span and within-sentence context seems to have the highest coverage; 4) separate treatment is needed to deal with multi-word collocates. However, given that surface/textual co-occurrence still produces a substantial amount of false positives, we intend to use syntactic co-occurrence to retrieve the frequency information for the combinations in Woordcombinaties, because it has a higher precision and offers the greatest scope for improvement.
Speakers: Carole Tiberius (Instituut voor de Nederlandse Taal/Leiden University Centre for Linguistics), Martin Kroon (Instituut voor de Nederlandse Taal), Lut Colman (Instituut voor de Nederlandse Taal), Dalí Chirino (Dutch Language Institute)
-
26
-
Multilingualism, Language Varieties & Contact Lexicography 00 | Anton Zeilinger Salon (OeAW Main Seat, ground floor)
00 | Anton Zeilinger Salon
OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Kristina Koppel (Estonian Language Institute)-
30
A Dictionary for All Canadians: Establishing an Inclusive Standard for the Canadian English Dictionary
To create ‘a dictionary for all Canadians’ is one of the stated goals of the new Canadian English Dictionary (CED). But who are ‘all’ Canadians, and how can a single resource meet their varied needs? To help answer these questions and develop lexicographical principles and protocols, the team behind the CED is exploring the concept of an ‘inclusive standard’. In this talk, we discuss our ongoing exploration of inclusivity and consider specific issues that arise in the drafting of entries.
Canada has been without an updated English dictionary for over twenty years (since Canadian Oxford Dictionary, 2nd ed., Barber, 2004). To address this gap, in 2022 a team of academics and editors launched the Society for Canadian English, a non-profit working group committed to creating a new resource. One goal is to produce an updated dictionary with changes to the language over the past twenty years; another important goal is to interrogate the ‘standard’ on which dictionaries are typically based and create a resource for ‘all’ that reflects the linguistic and cultural diversity of the country.
In Canada, the ‘standard’ variety of English typically characterizes the speech of urban, formally-educated people in the Central-West region from families with multiple generations in the country (Chambers, 1998). This conception of the standard comprises only 36% of the population (Dollinger, 2010) and excludes communities in the East and North, rural areas, those with less formal education, speakers of Indigenous Englishes, Anglophones in Québec, people in heritage language communities, and recent immigrants. A dictionary for ‘all’ must do better.
Our inclusivity goals are also necessarily decolonial. Dictionaries have historically been profoundly colonial in their construction and implementation, and projects like the CED must begin the work of practicing decolonial lexicographical methodologies. As well, an inclusive dictionary must consider sensitives around lexical items relating to specific social groups, including people with disabilities, racialized individuals and the LGBTQ+ community.
Representing the full range of Canadian English varieties and accents across different linguistic and social groups might achieve maximum inclusivity, but it could come at the cost of producing a resource with clear guidelines on how English is used in Canada, for those who seek such clarity. Dictionaries already struggle with a tension between descriptive and prescriptive aims and are used differently by editors, linguists, students and learners. An inclusivity goal makes balancing the needs and goals of diverse users even more challenging.
Given these issues, how do we define an ‘inclusive standard’, and how do we put our ideas into action in the dictionary-making process? After presenting a brief overview of the CED’s history, methodological procedures and corpus research, we focus our discussion on inclusivity, providing examples in the realms of defining, pronunciations and etymology.
In the area of defining, one issue of consideration is our practice for the inclusion and presentation-order of variant spellings. In the case of kayak, for example, do we include the spelling qajaq, which reflects the word’s Inuktitut roots, and if so, should it appear as the headword or variant? Another issue is the treatment of ‘offensive’ terms. This includes our use of labels to clarify offensive historical terms (e.g., half blood, paraplegic) as well as our implementation of usage notes for terms that are historically (and in some contexts still) offensive, but that have been reclaimed as neutral or positive (e.g., queer, Indian).
Regarding pronunciations, the diversity of accents in Canada means that inclusivity considerations affect all aspects of the transcriptions. One issue is the choice of transcription system. We explain our decision to use the IPA — a divergence from North American tradition — to increase accessibility, especially for learners. Another issue is on whose speech we base the main pronunciation and, relatedly, how much variation we include. We discuss the challenges of choosing a ‘model speaker’ who is broadly representative of Canadian English as well as how to accommodate alternative pronunciations without embracing a level of detail better left to a dialect dictionary.
With etymologies, we have the opportunity to ground our inclusive standard in the decolonial practice of acknowledgment, similar to acknowledgments of unceded Indigenous land that settler-colonial peoples occupy. Many Canadian English words have an origin in an Indigenous language that may not be readily understood to the average speaker, for example, kayak/qajaq, woodchuck, and squash (vegetable). This etymological work highlights the debt of Canadian English to Indigenous languages of North America. It also acknowledges the view amongst Indigenous scholars such as Emma LaRocque (Cree and Métis) that “English is the new Native language” and that “it is English that is now serving to de-colonize and to unite” Indigenous Peoples (1990, 57).
In an inclusive dictionary, every Canadian should be able to open the resource and see their variety of English represented and treated with respect. We hope that our development of an inclusive standard will benefit not only the CED but serve as a model for future lexicographical projects.
Speakers: Emma Ferrett (Queen's University), Amir Ghorvei (Queen’s University), Anastasia Riehl (Queen's University) -
31
Reversing the New English–Irish Dictionary Database to Create Entry Frameworks for Irish–English and Monolingual Irish Dictionaries
The first major monolingual Irish dictionary, An Foclóir Nua Gaeilge [The New Irish Dictionary] was launched in December 2025, when an initial tranche of 40,000 senses was published on focloir.ie. The dictionary will be added to incrementally in 2026 and 2027 to bring it to its content target of 80,000 senses. Along with the monolingual dictionary, a new bilingual Irish-English dictionary is also being produced and will be fully published in 2027. The entry frameworks for the new project were created by reversing the New English-Irish Dictionary (NEID) (2013-2017) database. This paper explores the decision to reverse the NEID database, the challenges involved in specifying the task, the resulting Irish-language entry frameworks and their use, the unintended consequences and related challenges, and the solutions found to those challenges.
Speakers: Cormac Breathnach (Foras na Gaeilge), Pádraig Ó Mianáin (Foras na Gaeilge) -
32
Spanish Loanwords in Modern Greek: Linguistic Analysis and Lexicographic Treatment in the Dictionary of Standard Modern Greek
This study aims to provide a systematic linguistic and lexicographic analysis of Spanish-derived loanwords in Modern Greek, addressing the limited attention given to Spanish as a donor language.
Methodologically, it draws on a dataset of 130 items, combining 105 entries from the Dictionary of Standard Modern Greek (DSMG) with 25 additional items collected from contemporary digital discourse. The analysis adopts a typological framework distinguishing loanwords, loanblends, and loanshifts, alongside grammatical, morphophonological, and semantic-domain analyses.
The results indicate a clear predominance of loanwords, with loan translations and semantic extensions occurring less frequently. Nouns constitute the dominant grammatical category, while morphophonological integration is largely systematic, aligning with Greek inflectional patterns. Semantically, borrowings cluster in culturally salient domains such as food, music, and material culture, whereas core vocabulary remains largely unaffected. The findings also highlight the predominantly indirect transmission of Spanish borrowings through intermediary languages and the accelerating role of global media in recent diffusion.
The study contributes by documenting Spanish as a secondary but meaningful donor language in Modern Greek and by identifying gaps in the DSMG, advocating for more systematic etymological attribution and a usage-based, dynamic approach to lexicographic practice.
Speaker: Zoe Gavriilidou (Democritus University of Thrace) -
33
From Pilot to Practice: Two Years of Progress in Developing a Shared Infrastructure for Dutch Dialect Dictionaries
At Euralex 2024, the initial design of a digital lexical infrastructure for Dutch Dialects was presented, developed in response to requests from dialect organizations in the Netherlands and Flanders to support the compilation and publication of dialect dictionaries. That contribution focused primarily on conceptual design choices and technical feasibility. Two years later, the project has moved from exploration to practice. The infrastructure has been implemented, tested, and refined through real-world use by two dialect communities. This paper reports on these developments, with particular attention to lexicographic equivalence, source integration, orthographic variation, and the role of non-professional lexicographers in collaborative dialect dictionary making.
Despite ongoing dialect loss, public engagement with regional varieties of Dutch remains strong. Dialect associations play a central role in documentation, education, and dissemination, often through dictionary projects or learning materials. This work is largely carried out by volunteers. While their commitment and local knowledge are indispensable, such initiatives typically operate with limited resources and without structural access to linguistic or lexicographic training. As a result, dialect dictionaries are often developed in isolation, follow divergent conventions, and are difficult to update, compare, or integrate. Lexical data are also at risk of being lost due to inadequate long-term preservation (Van Keymeulen et al. 2009). The project discussed here addresses these challenges by developing a shared digital infrastructure that supports sustainable dialect lexicography while accommodating differences in expertise, goals, working practices, and community-specific preferences.
Two pilot varieties were selected: Bildts, a Frisian-Dutch contact variety spoken in the province of Fryslân (Friesland), and West-Overijssels, a regional variety with a long lexicographic tradition. These cases were chosen not only for their linguistic characteristics, but also for the presence of active, volunteer-driven language communities. In both pilots, the first step consisted of establishing a lexical database as the backbone for future dictionary products, reflecting the lexicographic principle that a well-defined macrostructure is a prerequisite for consistent and reusable microstructural description.
The editorial workflow is centred on a lexical data editing environment (Lex’it) (fig. 1*) that supports structured data entry, linking, and revision. As an initial alignment device, a Dutch base word list derived from frequency-oriented learner resources was used. This list functions as a pragmatic onomasiological access structure, intended to facilitate dialect acquisition, translation activities, and comparison across dialects. Experience from the pilots confirmed, however, that frequency-based lists alone are insufficient for representing dialect lexicons. Many high-frequency standard-language words lack clear dialect equivalents, while culturally salient dialect words are often absent from such lists. Volunteers frequently prioritize local relevance and expressive richness over frequency.
In both pilots, existing dialect dictionaries (Buwalda 2013, Fien 2000, Kamman 1990, Kuijk & Van Baalen 2015) were integrated into the editing environment. For Bildts, one authoritative dictionary served as the primary source, supplemented with newly attested items stored in an extensible lexicon. For West-Overijssels, data from three independent dictionaries were combined, enabling the compilation of a compact, learner-oriented pocket dictionary suitable for a broader region. This process foregrounded classical lexicographic issues such as source criticism, sense delimitation, synonymy, and the treatment of variants.
Combining multiple sources substantially increased lexical coverage but also introduced ambiguity. Differences in spelling conventions, semantic scope, and levels of detail between dictionaries required explicit editorial decisions. Orthographic variation proved particularly prominent in volunteer-based projects, as contributors often follow different local, historical, or personal spelling systems. The infrastructure therefore supports the coexistence of multiple spelling variants without enforcing premature standardization, while still providing reliable search and access mechanisms (fig. 2*).
The pilots further demonstrated that lexical linking cannot be reduced to simple one-to-one equivalence. Partial equivalence, context-dependent meanings, and culture-specific concepts occur frequently in dialects. All editorial actions are logged within the system, enabling transparency, revision, and quality control (fig. 3*).
A key outcome after two years is that the same underlying system now supports two distinct public dictionary products (e.g. fig. 4 and 5*). The Bildts and West-Overijssels lexica are published on the respective websites of the collaborating organizations, each with its own interface and local identity, reflecting heterogeneous user groups ranging from experienced dialect speakers to beginners. At the infrastructural level, however, both products are generated from the same database and editorial workflow, illustrating a clear separation between lexicographic data, structure, and presentation.
The pilots show that creating digital dialect dictionaries involves not only linguistic data, but also volunteer practices, orthographic diversity, and diverse user needs. By embedding lexicographic theory into infrastructural design, the project offers a sustainable model that combines scholarly rigor with accessibility.
* see Book of Abstracts
Speakers: Veronique De Tier (Dutch Language Institute), Katrien Depuydt (Dutch Language Institute), Koen Mertens (Dutch Language Institute), Mathieu Fannee (Dutch Language Institute)
-
30
-
11:00
Coffee Break 00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna -
Keynote: Beyond the Entry: The Evolution of Historical Lexicography at the Oxford English Dictionary 01 | Festsaal (Festive Hall), OeAW Main Seat, 1st floor
01 | Festsaal (Festive Hall), OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Kate Wild (OUP)-
34
Beyond the Entry: The Evolution of Historical Lexicography at the Oxford English Dictionary
As the Oxford English Dictionary approaches the centenary of the completion of its first edition, this talk reflects on how the practice and possibilities of historical lexicography have changed over the past century. The first edition of the OED was produced letter by letter and entry by entry, culminating in the publication of its final volume in 1928. Alphabetical order is no longer a constraint, and for the past two decades entries have been selected for revision according to criteria such as cultural significance, semantic complexity, and user feedback, enabling more dynamic and targeted updates. In recent years there has been a further development in that the individual entry is no longer the only, or even the primary, unit of revision. Alongside traditional entry-by-entry revision, the OED is now undertaking systematic improvements to components across the text, sometimes supported by computational methods and AI. Key goals for the centenary include the addition of recent quotation evidence to all senses in current use, the modernization of outdated defining language across the dictionary, and the inclusion of etymologies and variant forms lists in all entries.
Looking beyond the text itself, the OED is also engaged in initiatives that position it increasingly as a digital dataset as well as a dictionary. These include the development of a historical corpus both as a tool to support editorial work and as a resource to complement OED data; and (in collaboration with the Historical Thesaurus team at the University of Glasgow) the further development and improvement of the Historical Thesaurus taxonomy, both as a structured dataset linked to the OED and as a tool which can support historical word sense disambiguation and other types of data analysis.
The argument advanced in this talk is not that historical lexicography should abandon its entry-by-entry craft, but that its future depends on combining that craft with methods designed for the dynamic, linked, data-rich digital environments in which dictionaries now live.
Speaker: Kate Wild (OUP)
-
34
-
12:30
Lunch Break 00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna -
Poster Session Aula (Entrance Hall) & 1st floor (00/01 | Aula (Entrance Hall) & 1st floor, OeAW Main Seat)
Aula (Entrance Hall) & 1st floor
00/01 | Aula (Entrance Hall) & 1st floor, OeAW Main Seat
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna-
35
A Definition Walks into a Dictionary: The Lexicography of Humor
Different disciplines approach the definition of humor from different angles: philosophy through theoretical frameworks that explain why people find things funny, psychology through behavioral and dispositional dimensions that describe how humor is experienced and expressed. This paper examines the treatment of humor in terms of its lexicographic definition, drawing on entries from six Greek dictionaries currently in circulation. Through systematic microstructural analysis, definitional components are identified and categorized, revealing significant variation in how dictionaries conceptualize humor. Definitions are broken down into components that reveal lexicographers’ attitudes towards humor’s nature (what it is), manifestation (how it appears), subject (what its target is), features (what characterizes it), and function (what its goal is). Results demonstrate that while lexicographers employ diverse definitional strategies, ranging from analytical to functional approaches, they align on certain core elements, particularly humor’s cognitive dimension and social function. This research proposes a systematic methodology for analyzing how dictionaries handle abstract, culturally embedded concepts, providing a replicable framework applicable to other affective, cognitive, or socially constructed concepts such as irony, empathy, and creativity.
Speaker: Anna Vacalopoulou (Athena RC / ILSP) -
36
A Lexicographical Analysis of Neologism Dictionaries Based on the Multilingual Dictionary of New Words (2010–2024)
This paper examines the mega-, macro-, and microstructure of all eight editions (2010–2024) of the Multilingual Dictionary of New Words (MDNW), published by the European Parliament’s Directorate-General for Translation, focusing on the treatment of Latvian and Lithuanian neologisms. Combining qualitative source analysis with quantitative and qualitative analysis of the extracted material, the study compiled a dataset of 184 neologism occurrences, corresponding to 87 unique lexical units (69 Latvian, 18 Lithuanian). The analysis shows that the dictionary’s macrostructure evolved from an early, hierarchically nested entry format toward a flatter, integrated one, and that its editorial model shifted from a cumulative index that retained every Latvian headword extracted from the 4th through the 7th editions (2013–2019) to a non-cumulative index in the 8th edition (2024) that included only the most recently attested words. These findings indicate that the MDNW applies no single, stable criterion of lexical “novelty” throughout its publication history, raising the broader question of how lexicographers should define the temporal scope of a “new word” and communicate it to dictionary users.
Speaker: Silga Sviķe (Ventspils University of Applied Sciences) -
37
Accessible Digital Language Resources (ADLaR)
A good deal of lexicographic research is concerned with evaluating or improving the accessibility of lexicographic data by making it easier for users to find and interpret. However, there is another sense of accessibility that has received less attention. This sense of accessibility concerns the design of resources so that they can be used by any person irrespective of personal features such as disability or impairment. Recent EURALEX congresses suggest a growing recognition of the importance of making lexicographic resources accessible in this sense. For instance, the XIX congress (in Alexandroupolis, Greece, 2021) was themed Lexicography for Inclusion while the XXI congress (in Cavtat, Croatia, 2024) featured a pre-conference workshop on Lexicography and Accessibility (Rees et al., 2024). Similarly, a handful of recent studies have demonstrated that existing dictionaries display a range of accessibility problems (Arias-Badia & Torner, 2023; Rees, 2023; Torres del Rey & García Garmendia, 2025).
However, in general, most research on the latter sense of accessibility in lexicography has suffered from several limitations. Firstly, the majority of studies have focused on a restricted set of users, mostly centring on persons with disabilities—especially visual impairments (Rees, 2025)—while the needs of those with cognitive or intellectual disabilities, as well as other disabilities, and more generally other vulnerable groups, have received little attention. Secondly, these studies have revealed little about how, or indeed if, users actually employ the lexical resources under evaluation. Thirdly, they have relied on compliance with guidelines such as the Web Content Accessibility Guidelines (WCAG) (World Wide Web Consortium (W3C), 2023)—a set of technical standards—as a proxy for accessibility evaluation rather than carrying out evaluation directly with users. Finally, existing research has dealt with monolingual resources for widely used languages, predominantly English and Spanish, while the accessibility needs of users of lower-resourced languages have not been directly addressed.
This poster outlines plans for the ADLaR (Accessible Digital Language Resources) project, which aims to address these gaps in previous research on accessibility in lexicography. Widening the range of participants considered, the project focuses on prototypical users who are at risk of digital exclusion, including persons with cognitive disabilities, many older adults—who remain among those “least likely to tap the potential of innovative technologies” worldwide (UN, n.d., para. 6)—and various sub-groups of migrants and refugees. These groups are among the target users of Easy-to-understand (E2U) texts, that is, texts produced using specific techniques intended to facilitate their comprehension by these populations (Jiménez-Andrés & Orero, 2022).
In the first, exploratory, phase of the project, a series of focus groups with users representing these cohorts is planned in order to establish whether their members make use of digital language reference resources (DLRRs), which resources they use, and how they use them. It is envisaged that digital dictionaries, online language forums, machine translation, and AI chatbots may be among the resources mentioned. Following this, semi-automatic web accessibility evaluation (c.f. Rees, 2023) of the resources identified in the focus groups will also be undertaken in order to give an initial indication of their accessibility.
In the second, experimental, phase of the project, participants will be directly involved in evaluating the resources identified in the exploratory phase. In this second phase, a series of screen-recorded look-up and text production experiments is planned with the aim of assessing the extent to which DLRRs help users resolve language doubts. These will be followed by standardised usability testing conducted with the System Usability Scale (SUS) questionnaire (Brooke, 1996) to assess the usability (and by extension accessibility) of the DLRRs identified. Finally, follow-up interviews with a sample of participants will allow further insight into user experiences of each DLRR.
The results of these procedures will inform guidelines for the creation of accessible DLRRs. Most of the work will take place in Catalonia, allowing the investigation of DLRRs not only for English, a global lingua franca, and Spanish, a widely spoken global language, but also for Catalan, a language that has received scarce attention in research on accessibility in lexicographic resources.
Given the barriers to education, employment, and more generally to participation in society that exclusion from DLRRs could produce, it is important that E2U language users have access to suitable resources. It is hoped that the ADLaR project will go some way to fomenting such access. By presenting the project at EURALEX, the poster aims to invite feedback, foster collaboration, and engage the EURALEX community in the project.
Speakers: Geraint Paul Rees (Pompeu Fabra University), Blanca Arias-Badia (Universitat Pompeu Fabra), Elisenda Bernal (Universitat Pompeu Fabra), Marta Brescia-Zapata (Universitat Pompeu Fabra), Clara Anush Del Rey Adanalian (Universitat Pompeu Fabra), Ana Frankenberg-Garcia (University of Surrey), Marta Garcia-Casado (Pompeu Fabra University), Paula Igareda González (Universitat Pompeu Fabra), Daniel Jiménez Casas (Universitat Pompeu Fabra), Robert Lew (Adam Mickiewicz University, Poznań), Alba Milà Garcia (Universitat Pompeu Fabra), Guiomar Salvat (Universitat Pompeu Fabra), Aurora Troncoso Ruiz (Radboud University Nijmegen), Sergi Torner (Universitat Pompeu Fabra), Patrick Zabalbeascoa (Universitat Pompeu Fabra) -
38
AI Used to Detect Problematic Content in Corpus Examples
Based on OpenAI’s moderation model, we describe a method developed to automatically identify unwanted discriminatory content in the citations found in the printed edition of Den Danske Ordbog (DDO, 2003–2005). Today, the DDO is published online and is continuously revised and expanded with new headwords. We begin by testing whether the model can identify 234 citations that have been removed since 2015 because the editorial team deemed them discriminatory. The results are promising, so we apply the model to 71,000 citations from headwords that have not yet been revised. The model classifies 6.3% of these as harmful. In addition to identifying the expected types of citations, a large proportion are flagged because they involve violence or self-harm—a type of problematic content that had previously been overlooked. Of the 4,800 flagged instances, 37% have been manually validated. Regarding violent content, the precision is only 10%—which nevertheless yields a list of approximately 350 problematic cases. The editorial team intends to discuss this type of content and establish new guidelines. For discriminatory content, the precision is higher—33%—which likewise results in a total of approximately 350 citations that are candidates for removal from the DDO and replacement with others.
Speakers: Sanni Nimb (Society of Danish Language and Literature), Ida Flörke (Society of Danish Language and Literature), Agnes Aggergaard Mikkelsen (Society of Danish Language and Literature), Nathalie Hau Norman (University of Copenhagen), Nicolai Hartvig Sørensen (Society of Danish Language and Literature) -
39
AI-Assisted Information Extraction from Larramendi’s Trilingual Dictionary (1745)
This poster presents the research on the application of different approaches to the information extraction process and encoding of the 18th century Larramendi’s Trilingual Dictionary. The digitisation and computational analysis of historical dictionaries present significant challenges due to archaic typefaces, complex layouts, and inconsistent lexicographic structures. Manuel de Larramendi's Diccionario trilingüe castellano, bascuence y latin (1745, hereafter LAR), a foundational text for modern Basque lexicography, remained largely inaccessible in a structured digital format, available only as scanned facsimile images.
Initial digitisation efforts employed a semi-automatic workflow combining machine learning (ML) for Optical Character Recognition (OCR) using Kraken and rule-based information extraction. The OCR for LAR was completed, and the resulting transcription was uploaded to Wikisource, a crowdsourcing platform that provides an interface for collaborative facsimile transcriptions. That process is ongoing and will very soon lead to an error-free digital version of the original text. Wikisource transcription quality is ensured by its rigorous, community-driven validation process, as completed works usually go through a two-step human review. The rule-based information extraction process facilitated the identification of Spanish headword candidates using positional data (line indentation) and orthographic features (capitalisation corresponding to alphabet sections). The extracted data was further enriched by linking Spanish lemmata to a contemporary reference, the Diccionario de Autoridades (1726-1739), and mapping historical Basque forms to modern standard lexicons. This work demonstrated the feasibility of converting a historical printed dictionary into structured datasets and linked data, though it highlighted limitations in fully capturing the complex microstructure of entries through rule-based methods alone (Lindemann & Alonso, 2021).
The first approach to information extraction from LAR involved a hybrid methodology. For a more nuanced structural analysis, it was then employed the Elexifier toolchain, a ML-based application. A sample of dictionary columns was manually annotated with tags from the TEI-Lex0 schema and used for training a model to predict and segment lexicographic components—such as headwords, Basque and Latin equivalents—across the entire dictionary text. This ML approach proved complementary to the rule-based method, showing high precision for Spanish headwords and Basque translations but lower accuracy for Latin equivalents and cross-references, revealing the complexity of the dictionary’s microstructure.
Advancing beyond specialised ML tools like Elexifier, contemporary research explores the application of general-purpose Large Language Models (LLMs) with vision capabilities for similar tasks. Inspired by methodologies applied to Estonian-German dictionaries (Jürviste & Jakobson, 2025) or to Swedish historical patent cards using the GPT-4o model (Xie et al., 2025), a third approach involves using DeepSeek, an advanced LLM, for the information extraction and encoding of LAR. The extraction of key microstructural elements like headwords (Spanish), translation equivalents (Basque, Latin), part-of-speech indicators, and example phrases and sentences is demanded from standard files, such as XML or JSON, or from the wikitext format text from Wikisource, with the few-shot prompting strategy, which basically consists of showing the model a few examples of inputs and desired outputs before asking it to solve the task. A fourth approach involves employing Microsoft Copilot, another multimodal LLM, for the same extraction task. The process consists of feeding Copilot with the same input and few-shot prompting, requesting structured TEI-Lex0 output that represents LAR’s tripartite entries.
Following an experimental framework used for historical dictionaries (Jürviste & Jakobson, 2025), this allows for an evaluation of different model performances on identical material. Key evaluation metrics include the structural correctness of the generated markup (adherence to TEI-Lex0) and the model’s ability to correctly classify and relate the Spanish, Basque, and Latin lexical units within each entry. Preliminary analyses of the results suggest that DeepSeek encodes TEI-Lex0 schema with almost no syntax errors and can correct the dictionary’s microstructural inconsistencies in the encoded output file, whereas Microsoft Copilot tends to fail to follow consistently the encoding schema showed in the given examples.
This comparative analysis extends the results of the mentioned studies about the strengths and weaknesses of different LLMs in handling the specific challenges posed by early modern bilingual/multilingual dictionary digitisation, particularly for languages with less digital representation. Nevertheless, it would be interesting for further digitisation processes to conduct experiments providing the tools with scanned images, because unlike the earlier workflow requiring separate OCR and segmentation stages, a vision-enabled LLM could potentially perform text recognition and semantic structuring in a single, integrated step. This approach promises a more direct path from facsimile to structured data, potentially reducing pipeline complexity and adapting more flexibly to layout irregularities.
In conclusion, the initial work on LAR established a viable pipeline and valuable datasets from rule-based and specialised ML tools, although the evolution to general-purpose (vision-enabled) LLMs represents a significant shift in historical dictionary digitisation methodology. Approaches leveraging LLMs aim to integrate OCR, segmentation, and semantic encoding into a more efficient workflow. This research not only advances the digitisation of a key Basque lexicographic resource but also can contribute to computational historical lexicography by evaluating and refining modern AI-driven extraction techniques.
Speaker: Mikel Alonso (University of the Basque Country UPV/EHU) -
40
Analysis of Synthetic Clinical Neologisms: A Corpus-Based Analysis of Large Language Model-Generated Content in the Context of Artificial Intelligence-Assisted Clinical Documentation
With generative artificial intelligence (AI) tools becoming more widely available and integrated into clinical environments, physicians are more frequently exploring ways to integrate AI into their clinical documentation and decision-making processes (Brender et al., 2025). Research already shows that large language models (LLMs) have the capability for summarization and analysis of clinical documents but carry also risks towards introduced hallucinated or biased content that can potentially skew further diagnostic steps and treatment options (Busch et al., 2025). Previous research suggested that the addition of neologisms to texts ended up decreasing machine translation quality by an average of 43% (Zheng et al., 2024).
This current investigation examined synthetically LLM-generated German clinical corpora and the use of neologisms and lexical patterns that differ from established controlled vocabularies, such as the Unified Medical Language System (UMLS) and SNOMED CT. Four synthetic German clinical training datasets were generated using GPT-4o and GPT-5 models. Each dataset comprising 900 snippets was constructed to compare zero-shot and few-shot prompting strategies, where both prompts instructed the generation of short, approximately 100-character length clinical-style snippets. LLM outputs were designed to reflect realistic clinical documentation features, such as abbreviated syntax, incomplete sentences, and laboratory-style reporting. Words or multiword expressions absent from reference biomedical corpora were flagged as candidate neologisms. These were then analyzed based on their morphological and structural properties.
To assess semantic relatedness between candidate neologisms and established biomedical terminology, embedding-based similarity analysis was performed using a pretrained multilingual transformer model (sentence-transformers/paraphrase-multilingual-mpnet-base-v2). Each candidate term and each controlled vocabulary entry was converted into a vector representation. The nearest-neighbor approach via cosine similarity allowed each candidate form to be linked to its most semantically similar standardized form. The method compared candidate terms that closely match existing medical concepts linked in terminologies, and those with low similarity scores, which were more likely to represent hallucinated expressions. The similarity scores were furthermore combined with frequency information to highlight commonly occurring but semantically distant forms for further analysis. This embedding-based comparison supported also frequency-based filtering of candidate terms for manual validation and review.
Preliminary analysis of synthetically generated German clinical narrative snippets revealed that the majority of candidate neologisms were not entirely novel lexical terms with reference to SNOMED CT and UMLS, but rather orthographic variants, abbreviation modifications, and fragmented forms derived from existing biomedical terminology. Co-occurrence analysis demonstrated clustering around conventional documentation templates when few-shot prompting was implemented in the synthetic generation pipeline. This suggested that neologisms primarily emerged through recombination and modification of established documentation patterns. Embedding-based similarity analysis further indicated that most candidate forms remained semantically close to controlled-vocabulary concepts, while only a small subset showed characteristics consistent with hallucinations. The comparison between prompting strategies revealed differences in lexical variability. Few-shot prompting led towards repetitive documentation templates, e.g., repeating the same structure and information found within the example prompt though varied slightly by changing abbreviations or contents. This meant that the few-shot method introduced rigidity with prompt examples and reduced lexical diversity, whereas zero-shot prompting produced a higher degree of structural variability and a greater proportion of candidate non-standard forms.
From a lexicographic perspective, this approach offers a systematic way to identify emerging terminology and non-standard usages in specialized corpora. By mapping novel forms to established vocabularies, lexicographers can better document language innovation, improve term standardization, and support the creation of more comprehensive, up-to-date lexical resources for clinical, technical and scientific domains.
Speaker: Amila Kugic (Medical University of Graz) -
41
Beyond Word Lists – But Not Beyond Dictionaries
The rapid rise of generative AI has prompted claims that traditional dictionaries, and perhaps even dictionary use itself, are becoming obsolete. We acknowledge that generative language models are transforming many aspects of information access and language-related tasks. However, we argue that they do not remove the need for high‑quality, human‑curated lexical resources. Generative systems, which rely on stochastic sampling and lack explicit provenance and systematic source control, cannot guarantee information that is fully accurate, transparent, and verifiable, and this seems unlikely to change. In this landscape, carefully edited lexical data remain not only relevant but essential. The challenge for lexicography is therefore not survival, but adaptation: competing for users’ attention and integrating lexicographic strengths into an AI‑saturated environment.
In this poster, we present the redesign of a large national dictionary portal that now brings together twelve dictionaries, all meaning‑oriented rather than purely spelling‑oriented. Our central claim is that online dictionaries have relied too long on opaque lists of word forms, especially in search results and autocompletion interfaces. Conventional interfaces typically display bare lemma forms, at most accompanied by a part‑of‑speech label. Such lists presuppose that users already know the words, or else force them to click through to each article before they can judge relevance. This is not only inefficient; it underuses one of lexicography’s core assets: sense information.
We propose moving “beyond word lists” by integrating compact meaning information directly into all key navigational entry points. In the redesigned interface of our portal site, autocompletion results and other word lists are now always enriched with short snippets derived from dictionary definitions. Users thus see, from the outset, not only which headwords are attested in the dictionaries but also a concise indication of what each candidate means.
For instance, in our main contemporary dictionary, typing malle now automatically autocompletes with (cf. figure 1):
• malle1 sb. “1. benfisk der ofte har skægtråde og en lille fedtfinne …”
• malle2 sb. “lille bøjle eller ring af metal der sammen med en tilsvarende hægte bruges til at lukke fx en vest e. …”
• mallemuk sb. “ca. 45 cm lang, mågelignende stormfugl med grå ryg og tykhals”This approach allows us to dispense with separate, formal disambiguation pages for homographs: instead of presenting a list of identical word forms that must be selected blindly, we display each homograph with an immediate sense‑based cue while typing, ideally shortening the path from the query to the relevant entry.
A further design challenge concerns inflected forms. Many existing systems exclude inflected forms from autocompletion, often out of concern that the same lemma will appear repeatedly. In our case, the autocompletion list is intended to assume much of the work of a separate search results page, so excluding inflected forms would significantly weaken its usefulness. We address this by allowing exact matches on inflected forms while still providing the same definition snippets, enabling users to navigate efficiently from the forms they actually type to the relevant lexical entries, cf. figure 2.
In cases where the user types a string that does not correspond to a lemma in the dictionary, the autocompletion list starts showing words that similar in spelling to the string, i.e. a Did you mean? result.
Implementing this snippet strategy across a heterogeneous set of twelve dictionaries, including several retro‑digitized works with only typographical markup and no explicit semantic structure, required developing methods for automatic snippet generation. In practice, we simply chose to present the first sense, or, in case of retro-digitized dictionaries, the beginning of the body of the entry, letting the sense number and the common abbreviation for ellipses, “...”, suggest that the entry include more sense information than presented in the autocompletion. We argue that even these automatically derived, simplistic sense cues represent a substantial usability gain. Even these support more informed navigation, reduce fruitless clicks, and make the semantic richness of the dictionaries visible earlier in the user journey.
We also utilize these snippets when supplying links to matches of the current query in the 11 other dictionaries, cf. figure 3.
We contend that the future of dictionaries lies not in mimicking generative AI, but in foregrounding what lexicography uniquely provides: curated, structured, and interpretable information that guides users through the lexicon more intelligently than word lists alone ever can. In the future, when the new platform has been live for a sufficient time, we hope to have data to support this.
In this poster, we present the motivation for the new design, and in-depth examples of the queries, before and after the redesign, and illustrative search scenarios highlighting how snippets support cross-dictionary navigation.
Speakers: Nicolai Hartvig Sørensen (Society of Danish Language and Literature), Andrea Stengaard (The Danish Dictionary) -
42
Bilingual Dictionaries with Italian in the 19th and 20th Centuries: The ALON Portal
The aim of this poster is to present ALON – Archivio della Lessicografia dell’Otto-Novecento (Archive of the Lexicography of the Nineteenth-Twentieth Century), an online portal dedicated to Italian lexicography. More specifically, this poster focuses on the section of the portal dedicated to bilingual dictionaries, particularly Italian-German ones.
During the 19th and 20th centuries, Italian lexicography underwent a process of renewal and reached its peak in terms of quantity, typology, and quality. This development also concerned bilingual dictionaries with Italian. The ALON project, funded by the European Union’s Next Generation EU programme as an Italian PRIN project from 2023 to 2026, aimed at studying and documenting the dictionaries of this period. The online portal described above constitutes the core of the project and contains presentations and detailed entries on the single dictionaries. In addition to a section dedicated to one of the Grandi Dizionari of the time, the Tommaseo-Bellini, and another devoted to monolingual dictionaries, the portal also includes a section on bilingual dictionaries from the same period.
The state of research on Italian bilingual lexicography varies greatly depending on the different language pairs. For Italian-French and Italian-Spanish dictionaries, extensive repertoires already exist in Lillo (2020) and in the Contrastiva portal (directed by San Vicente). For English-Italian dictionaries, an important reference is O’Connor (1990) with a detailed bibliography of dictionaries until 1990. In the case of Italian-German lexicography, the most comprehensive work remains an unpublished thesis (Bruna 1983), alongside the overview by Bruna, Bray & Hausmann (1991) and single studies (e.g. Schweickard, 2000; Giacoma, 2012; A scuola di Tedesco, 2016; Gärtig, 2016). ALON initially focused on this language pair to fill a gap in research and to develop a model for the database structure that could later be extended to Italian bilingual dictionaries with other languages.
At the present time, the ALON portal includes descriptions of 20 Italian-German dictionaries in their first edition (48 when considering further editions), ranging from Domenico Antonio Filippi’s Dizionario italiano-tedesco e tedesco-italiano. Italienisch-deutsches und deutsch-italienisches Wörterbuch (1817) to Martina Lucia Curcio’s Kontrastives Valenzwörterbuch der gesprochenen Sprache Italienisch-Deutsch (1999). Among the dictionaries currently described on the platform, there are some works that represent milestones in Italian-German lexicography. One is Francesco Valentini’s Vollständiges italienisch-deutsches und deutsch-italienisches grammatisch-praktisches Wörterbuch. Gran Dizionario grammatico-pratico italiano-tedesco, tedesco-italiano (1831–1836), one of the most important Italian-German bilingual dictionaries of the 19th century because of its many innovations, later adopted by other lexicographers. Another key work is Giuseppe Rigutini and Oskar Bulle’s Neues italienisch-deutsches und deutsch-italienisches Wörterbuch. Nuovo dizionario italiano-tedesco e tedesco-italiano (1896-1900), which marked the first collaboration between an Italian and a German lexicographer. Giuseppe Rigutini was also one of the most important Italian lexicographers of the century and the author of several dictionaries. His collaboration with Oskar Bulle on an Italian-German dictionary therefore highlights the important role played by the two languages in relations between the two nations. The platform also includes major dictionaries such as Henriette Michaelis’ Vollständiges Wörterbuch der italienischen und deutschen Sprache. Dizionario completo italiano-tedesco e tedesco-italiano (1879–1881), as well as so-called “pocket editions” such as Friedrich Ernst Feller’s Nuovo dizionario portatile italiano tedesco, tedesco Italiano. Neuestes Taschen-Wörterbuch der italienischen und deutschen Sprache (1851).
The database has a modular structure and includes both extensive texts and synthetic information on the dictionaries described. The level of detail varies according to the type of work considered. In addition to identification data for the dictionaries and an overview of their most important sources and research works on them, longer texts may describe their history (conception, publication history, reception, and circulation), structure, and content (paratexts, macro- and microstructure). Under the section “Materials and Resources,” links to digitized full texts are provided where available, together with documents such as images or archive material. This is followed by the bibliography. For bilingual dictionaries, the portal also provides additional summary information, including a transcription of the title page, a reconstructed table of contents, information on entry structure (besides the above-mentioned works cfr. Hausmann, 1989; Marello, 1989; Marello & Rovere, 1999), sources, distribution in Italy, and details of the consulted library. The portal takes into account not only the first edition of a dictionary, but also later editions, which may differ significantly.
The names of the dictionary authors are linked to a separate section describing their lives and works. This allows to illustrate the connections between lexicographical projects: for example, a person may appear as a reviser in one dictionary project and as an author in another.
All dictionary descriptions in the portal are assigned a DOI (Digital Object Identifier). The database is conceived as an expandable resource and offers external authors the opportunity to publish, after peer review, their research on Italian dictionaries from the 19th and 20th centuries in the form described above.
Speakers: Anne-Kathrin Gärtig-Bressan (Università degli Studi di Trieste), Pia Carmela Lombardi (Università degli Studi di Trieste) -
43
Choosing Optimal AI Prompt Instructions for Neologism Detection
The past few years have seen numerous attempts to make use of generative artificial intelligence (AI) in the field of lexicography. One such field is neology – the detection of new words and the preparation of their lexicographic entries for further study (cf. Zheng et al., 2024). When identifying neologisms in large collections of texts with the help of AI, existing neologism data – usage examples already collected by lexicographers – can be used for two purposes: (1) as positive examples included in AI prompts and (2) as material for evaluating model performance (cf. Tang et al., 2025). Since the number of neologisms in a dataset of usage examples is known, it is possible to calculate not only precision but also recall, enabling an objective evaluation of model performance.
Traditional methods of collecting neologisms are still very effective. The Database of Lithuanian Neologisms (DLN) relies on two traditional collection methods: editors regularly record neologisms encountered while reading news portals and social media; members of the general public submit neologism candidates they have identified via a form on the DLN website. As of 2026, the DLN contains more than 12,000 detailed neologism entries with definitions and usage examples. Because individual entries may have several examples, the database currently includes more than 43,000 usage examples in total. Many of these examples consist of multiple sentences, thus making the DLN a source for a very neologism-rich corpus comprising several million words. An additional advantage of manual collection is that neologisms can be documented from sources that are difficult to access for automated tools: social media posts, comments on YouTube and news portals, radio and TV recordings.
A dataset of this size and quality provides a strong basis for developing AI tools for automated neologism extraction from web texts and definition generation.
This presentation discusses prompt engineering experiments aimed at identifying an optimal prompt for the task of neologism detection, incorporating usage examples from the DLN. The experiments focus on the component of the prompt known as the instruction (or task description). Prompt engineering guidelines (Liu et al., 2024) typically describe a prompt as consisting of several elements: (1) the instruction, which tells the model what task to perform; (2) a role specification defining the model’s assumed role; (3) examples with corresponding correct responses; (4) formatting, style and length requirements for the output; (5) the input material that the model is asked to process.
Two alternative instruction strategies for neologism detection were tested. The first strategy instructs the model to process a given text and return only the neologisms it identifies (extraction-style). The second strategy instructs the model to examine each word in the text individually and determine whether it constitutes a neologism (classification-style). Using an open-weight language model, a series of experiments was conducted to evaluate the performance of both prompting strategies. Although the difference between these formulations may appear relatively minor, it results in substantial differences both in output quality and in the computational resources required to complete the task.
The experiments demonstrate that both prompting strategies achieve satisfactory performance in neologism detection. The classification-style prompt shows a slight advantage in the total number of correctly identified neologisms, whereas the extraction-style prompt demonstrates a significant advantage in terms of lower noise, i.e. fewer words incorrectly identified as neologisms. An even greater difference between the two strategies emerges from the perspective of practical application. When processing the same text on identical hardware, the extraction-style prompt required nearly twenty times less processing time than the classification-style prompt.
Both prompting strategies exhibit similar weaknesses. Despite explicit instructions, the model may return words containing introduced misspellings, provide grammatically altered forms instead of the original forms occurring in the analyzed text, omit relevant words from the output, reorder words relative to their order in the source text, or introduce words that do not appear in the analyzed text at all (hallucinations).
Further research should include a lexicographic evaluation of how the two prompting strategies perform in detecting specific categories of neologisms, such as derivatives, compounds, and borrowings. From a technical perspective, methods for mitigating model-induced distortions of analyzed words should also be explored.
The experiments indicate that current large language models can provide substantial assistance to lexicographers in the collection of neologisms. Nevertheless, professional linguistic expertise remains necessary for the final evaluation of candidate neologisms and the removal of false positives.
Speaker: Marius Glebus (Institute of the Lithuanian Language) -
44
Designing an Idiomatic Dictionary
This paper presents the principles and structure of Colidioms, an online idiomatic dictionary designed to provide a fine-grained description of idioms for their full understanding and production across typologically different language families, including French, Japanese, Korean, and Chinese. Grounded in cognitive and corpus linguistics and in recent work in phraseology and phraseography, the tool offers a unified model for both source and target languages, describing each idiom on at least three levels: formal structure, meaning(s), and use. The methodology combines corpus-based analysis, introspection, and linguistic experiments with native and non-native speakers. The paper focuses on the microstructure of entries, namely the canonical form and combinatorial properties, corpus-based definition(s), motivation, lexical, semantic, and syntactic co-occurrence constraints, register labels, variants, semantic links to related idioms (synonyms, equivalents, conversives, antonyms, false friends, hypernyms), notions, and pedagogical examples. The resource is both source- and target-oriented and enables multidirectional search for equivalents. We argue that Colidioms serves both practical aims (decoding and encoding idioms) and theoretical ones (contributing to the description of phraseological systems in typologically different languages and clarifying the theoretical status of idioms).
Speaker: Elena Berthemet (Centre de Linguistique en Sorbonne) -
45
Development of a Web Application for a Multilingual Terminological Dictionary
Most national terminological dictionaries in Ukraine are not accessible for automatic data processing, which makes it impossible to use them in digital information systems. The objective of the research is to develop a technology for converting specialized dictionary text into a website with a developed user interface.
The object of the study was “Dictionary of Ukrainian biological terminology” (7,342 entries and about 26,000 terms in Ukrainian, Russian and English) (Grodzinskiy, 2012), that contains definitions, terms polysemy, synonymy, stresses for Slavic languages, and grammatical information.
Since the dictionary text was available in digital publishing format (PDF), no prior digitization was required. Our approach is to step-by-step transform the linear text of a dictionary into a website. The basic steps are as follows:
-
Dictionary text normalization: restoration of the text line that represents the dictionary entry, stress marking, font markers fixation, correction of inevitable publishing errors in the dictionary entry structure, etc. This was the most time-consuming step, and it required manual processing. The text was converted into .doc format. MS Word text processor was used for processing, the result was text in .txt format, in which HTML tags were used to mark substrings, presented in bold and italic.
-
Designing a dictionary lexicographic system model (Shyrokov, 2021; Kupriianov, 2020). This model serves as a basis for building a parsing algorithm, designing a database schema and interface elements. The model was designed based on an analysis of the printed version of dictionary entries markup. Lexicographic systems model methodology allows us to identify all structural elements that can be identified automatically, and to establish connections between them. Each dictionary entry is assigned one universal structure, i.e. any dictionary entry is considered as a derivative of one “template” entry.
-
Construction of an XML schema based on the conceptual lexicographic model.
-
Automatic conversion of dictionary text (.txt format) into an XML document, allowing to explicate all defined structural elements and the connections between them. To automatically mark the dictionary text with XML tags, a program was developed that highlights the elements of the dictionary entry structure. We consider an XML document as a stand-alone product that effectively represents lexicographic data forfurther use for various purposes.
-
Lexicographic database creation. NoSQL (document-oriented databases) was chosen for this (Belkadi, 2021). In the case of relational databases, data is stored as a set of multiple tables and links between them. Working with individual tables as a single object requires a powerful software infrastructure. Moreover, the evolutionary potential of such a digital object is limited by the opacity of the database. Since dictionary entries are the basic elements of a lexicographic system with a strictly defined structure, it is logical to represent them as classes in object-oriented programming languages with subsequent processing, editing and storage in explicit form. The main advantage of NoSQL databases for our project is their ability to store explicitly lexicographic objects without changing their internal structure, which opens direct access to each element of the lexicographic object and significantly simplifies the possibility of editing and modifying (extending) it.
-
Converting XML file to database. This was performed automatically.
-
Designing of interface schemes and creation of a website (currently in progress).
Speakers: Iryna Ostapova (Ukrainian Lingua-Information Foundation of NASU), Mykyta Yablochkov (Ukrainian Lingua-Information Foundation of NASU), Alona Dorozhynska (Ukrainian Lingua-Information Foundation of NASU), Iuliia Verbynenko (Ukrainian Lingua-Information Foundation of NASU) -
-
46
Dictionaries of Engineering in the Age of AI
As cultural objects, dictionaries perform several types of cultural work. For example, they have the ability to influence knowledge legitimation: they exercise power as they assign value to knowledge by dictating what knowledge is valuable and trustworthy (the knowledge found in dictionaries) and what is not (the knowledge omitted by dictionaries) (Menagarishvili, 2020). Dictionaries also create what Anderson (2006, p. 6) calls “imagined communities” that can be defined as communities consisting of people who “will never know most of their fellow-members, meet them, or even hear of them, yet in the minds of each lives the image of their communion.” Such communities are present in our lives even though most of the time we are unaware of them, and this invisibility, perhaps, makes imagined communities even more powerful. Finally, each dictionary is created within a certain cultural context that affects that dictionary’s content, design, and usage.
As artifacts of technical communication, dictionaries are a genre with many subgenres that can be used to communicate about technology. Each dictionary has an audience, a certain purpose, and a structure and is used in a certain way, which are all aspects of documents discussed by scientific and technical communication (Menagarishvili, 2023). Moreover, definitions are a typical technical communication genre that is taught in every technical communication class. Further, many dictionaries combine text and visuals, which is another technical communication topic taught and studied by technical communication instructors/researchers.
In this poster presentation, I will discuss dictionaries of civil and environmental engineering as cultural objects and as artifacts of technical communication. More specifically, I will compare the following dictionaries describing civil and environmental engineering discourse: The McGraw-Hill Dictionary of Scientific and Technical Terms; McGraw-Hill’s Dictionary of Engineering; A Dictionary of Civil and Environmental Engineering; A Dictionary of Construction, Surveying, and Civil Engineering; and Environmental Engineering Dictionary and Directory. The field of civil and environmental engineering was chosen for analysis because, as an Engineering Communication Program Director in a School of Civil and Environmental Engineering, I observed that dictionaries describing this vibrant field have not been discussed in the literature. These specific dictionaries were selected for analysis because they all describe civil and/or environmental engineering terminology but differ in their level of field specificity. They are listed above in order of increasing field specificity: technical terms–engineering terms–civil and environmental engineering terms–civil engineering-only terms / environmental engineering-only terms. This part of the study will include dictionary macrostructure and microstructure analysis.
In the second part of the study, I will explore what happens from the cultural studies and technical communication perspectives when civil and environmental engineering terms make the jump from traditional dictionaries to GenAI. Based on consultations with six civil and environmental engineering experts at a top-ranked technological university, I will select five widely used civil and environmental engineering terms from each of the six subfields of civil and environmental engineering (30 terms total). The six subfields include construction and infrastructure engineering; geosystems engineering; structural engineering, mechanics, and materials; transportation systems engineering; water resources engineering; and environmental engineering. I will then conduct a cultural and lexicographic analysis (1) of dictionary articles from the dictionaries listed above and (2) of definitions provided by five popular GenAI tools likely to be used by civil and environmental engineering students at a prestigious technological university in the US in the summer/fall 2026.
This poster presentation might be useful for lexicographers, instructors using dictionaries and/or GenAI in their classes, and the general public interested in dictionaries and/or GenAI.
Speaker: Olga Menagarishvili (Georgia Institute of Technology) -
47
Dictionaries, Glossaries and Terminology Databases in the Age of Artificial Intelligence: Rethinking Boundaries and Functions
The increasing integration of large language models (LLMs) into (multilingual) specialized communication is challenging traditional distinctions between dictionaries, glossaries and terminology databases. While these resources differ in structure, purpose, theoretical foundations and user groups, LLMs increasingly access their underlying data as machine-readable ‘knowledge’. This paper examines how approaches ranging from prompt-based glossary integration and retrieval-augmented generation (RAG) to terminology-augmented generation (TAG), API-based access and knowledge-graph-based systems are affecting the relationship between language resources and AI. It argues that reducing rich lexicographic and terminological resources to simple equivalence lists risks losing contextual and conceptual information, whereas structured access can preserve and exploit their descriptive richness. At the same time, resource typologies remain relevant: resource type and metadata provide information about a resource’s theoretical basis, purpose, intended users and epistemic status, particularly when LLMs combine multiple resources that may provide competing information. Therefore, resource type, underlying data and resource function should be considered together. This perspective highlights the continuing importance of lexicographic and terminological resources in AI environments, suggesting an increasing role for lexicographers and terminologists in creating language data and organizing knowledge for users and machines alike.
Speaker: Barbara Heinisch (Eurac Research) -
48
Digital Lexicography for STEM Education: Designing a Multilingual Digital Glossary of Scientific Terminology in Greek Sign Language
Despite increasing efforts toward inclusive education, Deaf and hard-of-hearing students continue to encounter barriers in STEM education, mainly due to the limited availability of systematically documented scientific terminology in Greek Sign Language (GSL). This lack of accessible terminological resources restricts conceptual understanding, classroom participation and engagement with scientific knowledge. At the same time, specialised STEM educational resources require multilingual terminological alignment, concept-oriented organisation and accessible multimodal representation (Cabré, 1999; Pavel & Nolet, 2001).
In this study, STEM is used as an umbrella term covering Science, Technology, Engineering and Mathematics, while Natural Sciences are treated as a subcategory within the Sciences. The glossary focuses primarily on terminology from the Natural Sciences included in secondary education STEM curricula.
The paper presents the methodological design and development of a multilingual digital glossary of STEM terminology in GSL. The glossary integrates Greek, English and German scientific terminology with multimodal GSL representations and accessibility-oriented digital lexicographic design, drawing on principles from digital lexicography, terminology studies and inclusive educational design (Jackson, 2013; Gavriilidou & Mavromatidou, 2016).
Greek functions as the language of schooling, English as the dominant language of international scientific communication and German as an additional scientific language with historical importance in European scientific terminology, especially in the Natural and Earth Sciences (Howarth, 2020; Sfoini, 2021).
The glossary adopts a multilingual and multimodal structure in which verbal terminological equivalents are connected with visual and linguistic representations in GSL. Scientific concepts are represented through written terminology, visual support and GSL video representations, enabling accessibility across different linguistic and sensory modalities. These representations include recorded GSL signs, visual conceptual support and pedagogically adapted explanatory material designed to facilitate conceptual accessibility for Deaf learners (Charamba, 2019, 2020).
Methodologically, the glossary development followed a five-phase protocol grounded in contemporary approaches to digital lexicography and terminology management (Klosa, 2013). First, a specialised school corpus was compiled from Greek secondary education Geology-Geography textbooks and analysed using AntConc, processing 18,206 lexical units. Through term extraction and evaluation procedures, 448 core scientific terms and 155 supplementary terms were identified.
The terminology set subsequently underwent empirical multi-criteria evaluation according to six parameters: curricular relevance, scientific validity, corpus frequency, conceptual transparency for learners, representability in GSL and pedagogical accessibility. Quantitative evaluation was supplemented by qualitative feedback from experts in STEM education, GSL and lexicography, allowing iterative refinement of the glossary entries.
The fourth phase involved the lexicographic and terminological structuring of the resource. The macrostructure combines thematic and alphabetical organisation to support efficient retrieval and conceptual navigation. At the microstructural level, each entry integrates multilingual terminological equivalents, educational definitions, conceptual cross-references and multimodal GSL representations informed by the linguistic properties of GSL as a natural language (Sapountzaki, 2015; Sapountzaki et al., 2025).
The final phase focused on accessibility-oriented digital implementation. The platform incorporates high-contrast visual design, adaptable typography, linear navigation and compatibility with assistive technologies, aiming to ensure usability for learners with diverse sensory profiles, including deafblind users. Future validation procedures involving Deaf learners, educators and GSL specialists are also planned in order to evaluate the comprehensibility and pedagogical effectiveness of the glossary.
Overall, the paper contributes to current discussions on inclusive STEM education, multilingual terminology management and digital lexicography by proposing a systematically designed educational glossary integrating multilingual scientific terminology with multimodal sign language representation. The study offers a methodological framework that may inform the development of similar accessible terminological resources in other sign language and educational contexts.
Speakers: Efi Tzelepi (Democritus University of Thrace), Zoe Gavriilidou (Democritus University of Thrace) -
49
e-DRAE 1884: Towards a Critical Digital and Hypertextual Edition of a Nineteenth-Century Dictionary
This paper presents the project “Model for a Digital and Hypertextual Edition of the DRAE 1884 (e-DRAE 1884)”, carried out at the Autonomous University of Barcelona and the Spanish National Research Council (CSIC). Its main goal is to produce an online critical edition of the Diccionario de la lengua castellana published by the Real Academia Española in 1884 (hereafter, DRAE 1884). The 1884 edition of the DRAE was chosen because of its innovative character within the lexicographical tradition of the Real Academia Española and because it was the first edition to take into account the lexical proposals of the newly established American academies. The digital and hypertextual edition offers the full text in an interoperable format, https://edrae1884.uab.cat/edicion/, enriched with structured annotations that provide the information needed for the historiographical, lexicographical, and lexicological interpretation of the dictionary's contents.
Speakers: Margarita Freixas (Universitat Autònoma de Barcelona), Esther Hernández (Consejo Superior de Investigaciones Científicas), Marija Žarković Eriksson (Universitat Autònoma de Barcelona) -
50
English Adjectives in German: Processes of Integration Since the 17th Century and Lexicographic Implications
This paper presents the current state of a PhD project on the integration of English adjectives into German. While English nominal and verbal loanwords have undergone thorough investigation, adjectives were often only covered marginally. Some researchers have approached English adjectives more systematically and have identified integration processes and problems, but they mostly examined small data sets or specific genres (e.g. newspapers) at a certain point in time. For this PhD project, adjectives were collected from 1650 up to 2025 to create a substantial, diachronic basis for analysis. This large time span was further divided into three time periods, based on major turning points in 1870 and 1945, to render diachronic comparisons possible. The sources were scientific works as well as various types of dictionaries (e.g. general, foreign words, Germanising). First results show an increase of adjectives between the time periods, with a total of roughly 1000 adjectives in the collection. Currently, the adjectives are being categorised according to their morphological, phonological and orthographical properties, with the assumption being that adjectives with the same properties behave in the same way in regards to inflection and syntactic usage. These groups will later be verified through further (corpus) analyses.
Speaker: Luisa Cimander (Leibniz Institute for the German Language) -
51
From Detection to Standardization: Practices and Criteria for Terminological Neology
In terminology planning and terminology work more broadly, the creation and establishment of new terms within specialized domains may appear to be a well-governed practice, guided by widely accepted terminology standards and established procedures across many languages. However, rapid technological development and globalized communication increasingly challenge this model. The continuous creation of English terms, particularly in fast-moving domains, leads to their swift diffusion into other languages, often without adaptation. When adaptation does occur, it frequently follows English morphological patterns and preserves the original, often metaphorical motivation.
These developments raise important questions regarding the identification of emerging terms in texts and other communicative settings, e.g., through corpus-based and AI-assisted term extraction. Criteria for identifying terminological neologisms therefore differ substantially from those applied in lexicography. Whereas lexicographic practice typically relies on frequency as a key indicator distinguishing neologism from individual lexical creativity, in terminology the emergence of a new concept within a specialized field justifies in itself the documentation of a new term, even if its textual frequency remains limited. Consequently, the detection of new terms must combine textual evidence with conceptual and domain-specific validation.
This paper examines current European guidelines and institutional practices concerning neology in terminology, as reflected in relevant research and in the work of terminology institutions. Drawing on data from the ENEOLI survey (Kallas et al., 2025), literature describing national approaches (e.g., Vaus & Janson, 2025; Zanola & Resi, 2025), the recommendations developed within the framework of ISO/TC 37/SC 3, as well as on selected national and institutional materials related to terminology planning (e.g., Prys & Jones, 2007), the study analyses how emerging concepts and their term variants are identified, evaluated, and formalized. Two main research questions are addressed: which criteria are used to determine if a candidate neoterm should be documented in a terminological resource, and what methodological guidelines can be proposed for the detection, description, validation and eventually standardization of neoterms?
The analysis of the responses collected in the ENEOLI survey shows that terminology institutions have not yet developed dedicated policies or formal workflows specifically addressing terminological or lexical innovation. Despite the lack of formalized procedures, several workflows used for the identification and description of new terms have been described, which largely implement the already established principles of terminology work as suggested in ISO standards (ISO 704, 2022). The most frequently mentioned criteria for inclusion of terminological neologisms in terminological resources included domain specificity, conceptual relevance, the frequency of use, corpora attestations, communicative necessity, and the ability to fill lexical or semantic gaps.
In the ISO TC/37 standards, the inclusion of neoterms in terminological resources is also not described by an explicit set of inclusion criteria. Instead, the relevant criteria are distributed across standards on terminology principles, concept analysis, term formation, harmonization, structured terminological representation, and resource modelling. In a similar vein, the majority of publicly available documents on terminology planning and terminology work that are published by different terminology institutions in Europe largely implement the criteria from ISO standards.
On the basis of the materials collected and analysed, methodological guidelines for the systematic detection and description of terminological neologisms are proposed. The management of terminological neology is understood as an iterative workflow consisting of detection, description, validation, harmonization, dissemination, and revision. These phases form a cycle rather than a one-directional sequence, and they include the application of contemporary corpus-based terminology work alongside generative AI systems, as well as concept analysis and expert validation as the foundations of terminology work.
Speakers: Ana Ostroški Anić (Institute for the Croatian Language), Federica Vezzani (University of Padua), Kris Heylen (Dutch Language Institute (INT)), Ilan Kernerman (Lexicala by K Dictionaries) -
52
From Field Recordings to an Online Dictionary: A Model for Processing Štokavian Dialect Material
This poster introduces a multi-modular model for developing an online Croatian dialect dictionary. Although Croatian dialect lexicography advanced significantly in the 20th century (Lisac 2006) and continues to develop in the 21st century, online Croatian dialect dictionaries remain uncommon. Most available resources are digital reproductions of print dictionaries or amateur compilations that lack adherence to established scientific principles of dialectological analysis.
The Croatian language exhibits a threefold dialectal division. However, no existing dictionary encompasses material from all three Croatian dialects, nor does any cover more than one dialect (Magaš 2026). Additionally, there is no Croatian dialect dictionary with a multi-modular structure similar to that developed for the Croatian standard language (Hudeček & Mihaljević 2024). Despite the predominance of Štokavian speakers in Croatia, dictionaries dedicated to individual Štokavian local varieties are limited.
To address this gap, the project focuses on Štokavian dialects and aims to develop a dialect dictionary that integrates multiple corpora, various levels of linguistic processing, and diverse target age groups. This dictionary is intended to serve as a model for processing the lexical material of other Croatian local varieties, groups of varieties, and dialects, especially those recognized as intangible cultural heritage.
Within the project Rijekom neretvanskih riječi – od kuće do škole (“Along the River of Neretva Words – From Home to School”), an online dialectological Dictionary of Neretva Dialects is under development. This dictionary includes varieties from two Štokavian dialects: Neo-Štokavian Ijekavian and Neo-Štokavian Ikavian. It represents the first attempt to present material from two Croatian dialects simultaneously and to implement transparent, functional lexicographic solutions. The dictionary features modules designed for pupils from the first grade of primary school through the fourth grade of secondary school. It is also the first Croatian online dialect dictionary to incorporate multiple fields, including headword, linguistic annotation (expandable), meaning, sentence example, idioms, audio recording, video recording, drawing, additional notes, and references.
Because dialect dictionaries frequently omit forms essential for precise phonological and morphological description (Kapović 2008), this project emphasizes the recording of grammatical forms relevant to determining the accentual paradigms of individual lexemes. Dictionary entries are compiled using TshwaneLex software, with three distinct modules developed:
- a module for grades 1–4 of primary school, which includes accented lemmas, meanings, drawings, additional notes, audio recordings, and video recordings;
- a module for grades 5–8 of primary school, which includes accented lemmas, grammatical forms, meanings, and audio recordings;
- a secondary-school module, which includes accented lemmas, grammatical forms, extended meanings, idioms, accented sentence examples, and audio recordings.
The core corpus of the dictionary comprises lexical items collected through a dialect survey questionnaire that covers a range of thematic domains, including family relations, fruits and vegetables, agricultural work, life by the river, children's games, recipes, and customs. To prevent restriction to predefined thematic fields, the corpus is systematically expanded with lexical items extracted from recordings of informants’ spontaneous speech. This thematic approach to data collection is particularly appropriate for primary school pupils in grades 1–4 and serves as a basis for further corpus expansion. Pupils in grades 5–8 actively participate in collecting and processing additional material, thereby contributing to both the educational and research aspects of the project.
Because the corpus includes material from two dialects, the dictionary records different phonological and morphological variants of lexical items. Each variant is accompanied by an abbreviation of the settlement where it was documented, enabling precise representation of local variation and systematic lexicographic treatment.
Secondary school students collect phraseological material using a targeted questionnaire developed from previous research on Neo-Štokavian dialects and through the investigation of selected concepts, such as strength, beauty, and stupidity. These concepts reveal conventional perceptions, stereotypes, and value systems characteristic of the speakers of a particular community. This approach enables systematic documentation of phraseological units and provides insights into the cultural and cognitive dimensions of dialectal language use.
By presenting these dictionary structures, the project offers insight into a contemporary and standardized approach to processing Croatian dialect material, ensuring both accessibility and searchability.
Speaker: Perina Vukša Nahod (Institute for the Croatian Language) -
53
From User Feedback to Better Definitions: Applying LLMs for Pedagogical Lexicography
In the presentation we will introduce the research on applying LLMs to help create definitions for learners of Estonian as a second language. For Estonian, there have been experiments conducted to create general language dictionary definitions with LLMs (see Tuulik et al., 2025). For English, LLM-s are tested for generating learner’s definitions as well (see for example Ide et al., 2026; Rees & Lew, 2024). Our research is based on a survey we conducted with 55 learners representing different proficiency levels (A1–C1). The aim of the survey was to investigate how learners of Estonian comprehend dictionary definitions in a learner-oriented digital resource and identify linguistic features that support or hinder pedagogical effectiveness. Participants evaluated two alternative definitions – an original and a pedagogically adapted version – for 20 headwords from a dictionary portal Learner’s Sõnaveeb. The analysis focused on lexical, morphological, and syntactic parameters influencing comprehension.
The results showed that learners consistently prefer pedagogically simplified definitions that are compact, clearly structured, and free of semantically empty pronouns, excessive conjunctions, and abstract modifiers. At lower proficiency levels (A2–B1), learners tend to favour familiar and internationally transparent loanwords, while this preference diminishes at higher levels. Morphological analysis revealed that the impersonal voice does not significantly impede comprehension, whereas non-finite constructions present notable difficulties. From a syntactic perspective, comprehension is enhanced when the headword appears at the beginning of the definition and when sufficient, relevant information is provided, even at the expense of brevity.
We are applying the results of the survey to adjust the existing definitions in Learner’s Sõnaveeb and to guide LLMs to create definitions for new B2-level headwords. Based on expert evaluations, the best performing model is selected for the task. We prompt LLMs using new principles and exemplary definitions to explore how to automatically get best quality draft definitions for lexicographers to edit. Using drafts provided by LLMs would significantly speed up the compilation process and thus allow for a faster expansion of the headword set in Learner’s Sõnaveeb.
Speakers: Maria Tuulik (Institute of the Estonian Language), Kristina Koppel (Institute of the Estonian Language), Ene Vainik (Institute of the Estonian Language), Margit Langemets (Institute of the Estonian Language), Lydia Risberg (Institute of the Estonian Language), Hanna Maask (Institute of the Estonian Language), Esta Prangel (Institute of the Estonian Language), Liina Lutsepp (University of Tallinn) -
54
Pedagogical Lexicography in Practice: Raising Dictionary Use Strategies through Card-Based Instruction
This paper presents the pedagogical rationale, theoretical foundations, and instructional design of the book 100 Practice Cards for Dictionary Use and Vocabulary Development (Gavriilidou & Konstantinidou, 2025), a classroom-oriented educational resource developed to promote strategic dictionary use among upper-primary and lower-secondary learners. Grounded in contemporary pedagogical lexicography, strategy-based instruction, and vocabulary acquisition research, the material aims to transform dictionary consultation from a mechanical reference activity into a meaningful learning process that supports vocabulary development, metacognitive awareness, learner autonomy, and language proficiency. The study highlights the broader educational value of the material, especially its contribution to inclusive education, strategic dictionary use, soft skill cultivation and autonomous language learning.
Speakers: Zoe Gavriilidou (Democritus University of Thrace), Evanthia Konstantinidou (Democritus University of Thrace) -
55
Pedagogy as the Core of Education: A Domain-Analytic Route to a Corpus-Based Dictionary
Mastering academic language remains a central challenge across university disciplines, as learners are expected to handle increasingly complex topics, genres, and stance-driven patterns, often with higher stakes for those working in an additional language. In response, corpus linguistics has reshaped lexicography (Rees, 2021) by evidencing language use not only as single lexical units, e.g., academic wordlists (Coxhead, 2016), but also as recurrent phraseological combinations, where corpus methods have been especially transformative (Paquot, 2015). Projects such as the Louvain EAP Dictionary (LEAP) demonstrate that online, needs-driven academic writing aids can now be structured around discipline-specific phraseological evidence (Granger & Paquot, 2015).
Despite growing interest in discipline-sensitive EAP resources (Rees, 2021; Ang & Tan, 2018), Education remains completely unexplored as a target domain for discipline-specific EAP lexicography. Education was not formally recognised as an academic discipline within universities until the twentieth century (Wyse, 2020), and its lexicon is distinctive: high-stakes yet conceptually diffuse, multidisciplinary, institutionally embedded, and strongly value-laden, with terminology that is often contested, relabelled, and unevenly stabilised across subfields. The practical need is substantial: beyond researchers and postgraduate students writing about pedagogy in English as an additional language, the global English-language teaching industry produces a large population of trainee teachers (e.g., CELTA, TESOL programmes) who must master pedagogical terminology in English. Against the United Nations SDG 4 (Quality Education), accessible and precise pedagogical language becomes more acute still, particularly in multilingual academic communities (Adipat & Chotikapanich, 2022).
Methodologically, the poster presents a needs-driven corpus design (Nesi, 2015; Flowerdew, 2013) in which the target domain was not imposed from a pre-existing taxonomy but derived empirically from the community itself. Semi-structured interviews with twelve members of the Pedagogy in Education Studies (PES) community of practice were analysed using reflexive thematic analysis (Braun & Clarke, 2021), yielding fifteen inductive thematic clusters. These were consolidated into five macro-themes, pedagogy and assessment; inclusion and diversity; teachers; curriculum and technology; and PES as discipline, which served as the topical architecture of the corpus and as seed-term generators for web-based corpus construction via BootCat (Baroni & Bernardini, 2004). The resulting PES corpus was then evaluated retrospectively against Egbert, Biber and Gray’s (2022) representativeness framework, following Kemp’s (2024) concept of a Representativeness Argument: a structured post-hoc validation addressing both target-domain coverage and linguistic stability (Gray, Egbert & Biber, 2017). The poster visualises this full workflow, from community needs analysis through corpus construction to domain evaluation and the resulting pathway for extracting a pedagogy-centred lexicon that captures both high-frequency academic vocabulary and the phraseology through which pedagogical knowledge is routinely articulated by PES scholars.
Finally, the study is motivated by a commitment to strengthening Education’s standing as an academic discipline: scholars who are driven by an ethical purpose to make a difference in the lives of learners and society constitute, in Furlong’s (2013) terms, a hugely important force in defining the field. Lexicographic resources that clarify pedagogy’s core lexis and phraseology can function as enabling infrastructure for this purpose, supporting educational participation and knowledge exchange beyond institutionalised borders (Liguori et al., 2023). The poster will also address how the impact of such a dictionary might best be measured, including uptake metrics, user studies, and integration with existing EAP writing tools.
Speaker: Maria Ammari (AMU, DUTh) -
56
Sublexical Systems of Colour in Biblical Georgian and Hungarian: Findings of the Pentateuch
As a culturally embedded linguistic category, colour varies across translations of the same source text. This paper examines how colour is encoded diachronically in biblical Hungarian and Georgian, addressing two questions: how colour terms are realised across parallel translations, and what cross-linguistic patterns of lexical and syntactic variation emerge. The analysis draws on a parallel corpus of Old and Modern Georgian and Hungarian Bible translations. Colour terms are defined as lexical units denoting chromatic properties, including both basic colour adjectives and material-based or metonymic expressions. The findings reveal a systematic divergence between the two cultural-linguistic traditions. Georgian shows a diachronic shift from materially grounded, periphrastic expressions towards more abstract and generalised colour terminology, often accompanied by increased morphological transparency. Hungarian, by contrast, demonstrates greater lexical continuity, with variation primarily realised through syntactic explicitation and the redistribution of semantic content. In both cultural-linguistic traditions, colour is encoded via material reference, emerging in descriptions of textiles and ritual objects. The findings demonstrate that colour encoding reflects broader processes of semantic restructuring. Hungarian serves as a comparative baseline, highlighting a trajectory in which syntactic rather than lexical change predominates.
Speakers: Tibor M. Pintér (Károli Gáspár University of the Reformed Church in Hungary), Manana Ruseishvili-Cartledge (Ivane Javakhishvili Tbilisi State University), Dóra Pődör (Károli Gáspár University of the Reformed Church in Hungary), Katalin P. Márkus (Károli Gáspár University of the Reformed Church in Hungary), Marine Makhatadze (Ivane Javakhishvili Tbilisi State University) -
57
The Role of LLMs in the Automatic Distribution of Meanings from Knapiusz's Polish–Latin–Greek Dictionary of 1643: A Pilot Study
Introduction
Grzegorz Knapiusz’s Thesaurus polonolatinograecus (second, expanded edition 1643; hereafter Kn) is the largest seventeenth-century dictionary that also records Polish vocabulary. It is one of the source materials for the Electronic Dictionary of Seventeenth- and Eighteenth-Century Polish (hereafter e-SXVII). One challenge for e-SXVII editors is assigning the content of Knapiusz’s entries—which contain, in addition to Polish, extensive Latin and Greek material provided as equivalents of Polish items—to the senses distinguished in e-SXVII entries. Although the Thesaurus was intended primarily as an extensive synonym dictionary supporting Latin learning, it also contains numerous observations on Polish vocabulary. It is regarded as the first work in Polish lexicography to address the semantics of Polish words. Knapiusz does not explicitly divide meanings within entries; instead, the Latin and Greek equivalents, or associated lexical items, of a Polish headword shift fluidly from one meaning to another, without clearly marked boundaries.Aims
This study aimed to test whether LLMs can support the segregation and disambiguation of meanings in Kn entries, as well as their assignment to corresponding senses in e-SXVII. We designed and evaluated a procedure for the semi-automatic distribution of Kn material within e-SXVII entries using GPT-5.2 Thinking, Gemini 3 Pro Thinking, Perplexity Sonar Pro, and Claude Sonnet 4.5.Method
We tested ten polysemous nominal entries, including mieczyk ‘short sword; gladiolus; yellow iris; bulbous iris; stinking iris; swordfish; Prussian coin’, dowód ‘proof; document; evidence; demonstration; reason’, laska ‘staff; mace; wand; rod; ridge; office; shepherd’s purse; unit of measure’, wyrok ‘oracle; decree; sentence; fate’, zapis ‘record; deed; bequest; legal act; treaty; statute; register; address; lodging assignment’, gwałt ‘assault; force; outrage; rape; corvée labour; violence; necessity; tumult’, wynalazek ‘idea; invention; discovery; decision; finding; result; decree; compromise’, bukowina ‘beech forest; woodland; beech trees; beech wood; beech grove; forest settlement’, motłoch ‘rabble; mob; worthless mass; commoners; riffraff’, and jabłko ‘apple; apple tree; apple-like fruit; pomegranate; truffle; pomander; heraldic charge’.The trials differed in the amount and type of information provided to the models: (1) a basic instruction to translate, segregate and disambiguate Latin and Greek equivalents in Kn entries; (2) the same instruction with reference to the sense structure and definitions of model e-SXVII entries; (3) the instruction from trial 2 supplemented with quotations and established word combinations; (4) the instruction from trial 2 supplemented with the e-SXVII principles for defining senses and distinguishing polysemy from homonymy; and (5) the instruction from trial 4 supplemented with illustrative material from e-SXVII.
To reduce the influence of conversational context, each Kn entry was queried separately, in a new conversation. In evaluating the results, we took into account inter-entry variability resulting, among other factors, from different degrees of lexical polysemy and different entry structures. Outputs were compared with model e-SXVII entries and evaluated by lexicographers and classical philologists according to an agreed scale covering the correctness of sense division, the assignment of equivalents to senses, and agreement with e-SXVII definitions.
Results and conclusions
Perplexity Sonar Pro was excluded because its responses were not consistently generated by the same model. GPT-5.2 Thinking, Gemini 3 Pro Thinking, and Claude Sonnet 4.5 produced outputs of highly variable accuracy. Each model performed flawlessly on some entries but made significant errors in others. This variability appears to be related to differences in polysemy and entry structure. Increasing prompt specificity did not produce a straightforward improvement in accuracy, suggesting that output instability cannot be resolved merely by enlarging the sample without changing the validation procedures.The trials also examined the minimum amount of information needed for an AI model to identify the semantic structure of a Kn entry correctly. For entries with highly typical mechanisms of polysemy, such as bukowina ‘a forest dominated by beech trees’ → ‘individual beech trees’ → ‘beech wood’, satisfactory results were obtained as early as the first trial. However, in comparable cases, such as jabłko ‘apple fruit’ → ‘apple tree’ or ‘round fruit’ → ‘round object’, the results were unsatisfactory both in the first trial and after additional context had been supplied. For some entries, subsequent trials even led to a deterioration in the disambiguation of Kn meanings, regardless of the AI model used.
The pilot study revealed clear limitations of the tested procedure. Verification of the method showed that any attempt to automatically divide meanings within loosely structured historical dictionary entries must take into account both their content and their intended function. In the case of Knapiusz’s Thesaurus, most illustrative material was taken from ancient authors and may therefore represent meanings not attested in Polish. Moreover, the study indicates that the current functioning of the selected models does not allow for standardised results or fully automatic support for lexicographic work. Nevertheless, under appropriate methodological assumptions, AI models with a “thinking” mode may offer support to lexicographers by generating hypotheses about senses, identifying problematic equivalents, and comparing Kn material with e-SXVII definitions
Speakers: Magdalena Majdak (Polish Academy of Sciences), Ewa Rodek (Polish Academy of Sciences), Katarzyna Kryńska (Polish Academy of Sciences), Jagoda Marszałek (Polish Academy of Sciences) -
58
Tracing Neologisms in Neo-Discourses: Diachronic and Cross-Domain Perspectives
Documenting the surge of new lexical items during the COVID-19 pandemic presented a challenge worldwide (Klosa-Kückelhaus & Kernermann, 2024). German Neologisms emerging between 2020 and 2023 were collected (Klosa-Kückelhaus, 2021) and incorporated into a dedicated section of the Neologismenwörterbuch (2006ff.), initially as a simple word list. Each entry includes a brief paraphrase and illustrative usage examples, often drawn from online sources, as validated corpus data were not yet available at the time of compilation. This section now comprises about 2,500 lemmas (Neuer Wortschatz rund um die Coronapandemie, 2020ff.). Taken together, these document the rise of a neo-discourse marked by an influx of new words, structural complexity, lexical competition as well as shifting narrative and topical emphases (Storjohann & Cimander, 2022, Mattfeld et al., 2025).
Consequently, the need arose to enhance user access to these entries. This initially involved assigning ontological categories through a triple-blind procedure, thereby establishing domains and semantic groupings for the lexical items. These were taken from the existing domain classes of the Neologismenwörterbuch (2006ff.) used for grouping headwords between 1990 and 2020. The aim was to create the kind of semantic and thematic connectivity that is crucial for online dictionaries (Dornseiff, 2020). In addition to general domains such as ECONOMY, EDUCATION, HEALTH/MEDICINE, POLITICS (cf. Ilić, 2021), a domain specific to Covid discourse was identified: MEASURES encompasses terms that refer to regulatory interventions and restrictions during the pandemic.
A representation of Covid-related vocabulary is currently being integrated into the new platform IDS Neo2020+, as it offers a range of entry types, e.g. single entries, discourse entries and discourse-morphological entries (Storjohann, in prep.). These enable the documentation of neologisms in single headword entries, in clusters, by concept, by domain and across temporal stages. This structure allows users to explore individual lexical items, their diachronic development, functions in the respective neo-discourse and semantic relations, especially among synonyms. Consequently, users may consult either individual COVID-related lexical items or, alternatively, a discourse entry that aggregates COVID-19 neologisms according to various criteria and presents them in relation to temporal, thematic, or structural dimensions.
While implementing new formats, a series of experiments will assess the extent to which AI can support or complement lexicographic work. These include AI-based categorisation of terms compared to manual classification, the identification of synonyms and the determination of first attestations (Storjohann & Cimander, 2022). AI may also facilitate the migration of the concise entries into the revised and more complex architecture of new headword entries. This step is necessary to construct a comprehensive discourse entry, “COVID,” by incorporating and elaborating upon details that are integral to individual COVID-related terms. Given the large number of items, enrichment through AI-assisted methods appears promising, e.g. by adding reliable citations, encyclopaedic or historical context, information on word-formation productivity, collocation profiles. It will also be examined whether such information can not only be identified by AI but also automatically assigned to designated elements at specific positions within larger XML-based entries. Positive outcomes could substantially influence the scope and depth of the documentation of Covid neologisms and, more broadly, the modelling of this neo-discourse within the new dictionary.
Speakers: Luisa Cimander (Leibniz Institute for the German Language), Petra Storjohann (Leibniz Institute for the German Language) -
59
Unlocking a Nineteenth Century Dictionary: A Closer Look at Konráð Gíslason’s Dönsk orðabók (1851)
This article discusses an important Icelandic lexicographic work published in the nineteenth century. Dönsk orðabók með íslenzkum þýðingum was compiled and edited by Konráð Gíslason (1808-1891), an Icelandic scholar and language purist. The printed dictionary is described in some detail, with particular attention to its structural characteristics and the ways in which purist ideals shaped many of its entries. Loanwords were systematically avoided and explanatory paraphrases were preferred whenever no suitable native Icelandic equivalent was available. The discussion then turns to recent efforts to produce a digital edition of the dictionary. We account for the workflow implemented in order to turn this complicated printed book into a functioning searchable edition where the content has been marked-up and XML-encoded. The process illustrates some of the issues involved and addresses challenges encountered when producing the new digital edition of the dictionary.
Speakers: Ellert Thor Johannsson (The Árni Magnússon Institute for Icelandic Studies), Jóhannes B. Sigtryggsson (The Árni Magnússon Institute for Icelandic Studies), Thordis Ulfarsdottir (The Árni Magnússon Institute for Icelandic Studies)
-
35
-
Software Demonstrations 01 | Sitzungssaal (01 | Sitzungssaal, OeAW Main Seat, 1st floor)
01 | Sitzungssaal
01 | Sitzungssaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna-
60
DEC Facile: Semantic-Web Authoring for Combinatorial Explanatory Dictionaries
This paper presents DEC Facile (Bellandi et al., 2025), a tool for creating Explanatory Combinatorial Dictionaries (ECDs) in a Semantic Web environment. Grounded in Explanatory Combinatorial Lexicology (Melʹčuk et al., 1995) and adopted in both general and specialized lexicography, ECDs provide a highly structured representation of lexical meaning and combinatorial properties.
Because of their expressiveness, however, ECDs are complex and time-consuming to build. Earlier ECD editors (Sérasset, 1996; Gader et al., 2012) supported the formalization of lexicographic data; DEC Facile extends this line of work by enabling the creation of machine-actionable and interoperable ECDs compatible with Semantic Web standards. Its design is based on three principles: ease of use, especially for humanists unfamiliar with RDF/OWL; support for collaborative lexicographic work; and the production of FAIR lexical resources (Wilkinson et al., 2016).
The system is built on LexO-server (Bellandi, 2023), a Java-based backend exposing RESTful services for dictionary and sense management. It orchestrates RDF models based on OntoLex-Lemon and its extensions, notably Lexicog, SynSem, LexFom (Fonseca et al., 2016), and SKOS. Lexical data are stored in GraphDB and accessed through SPARQL queries. The graphical interface is a responsive single-page web application developed in TypeScript with Angular, PrimeNG, Tailwind CSS, and PrimeIcons. It supports structured editing, dynamic language switching, and real-time interaction with backend services.
The main interface (Figure 1*) offers an integrated workspace for navigating, creating, and managing lexical resources through a modular three-panel layout that hides the complexity of Semantic Web technologies from non-technical users. The left-hand panel provides access to the dictionary inventory, with functions for creating, importing, exporting, searching, and filtering dictionaries. The central panel organizes dictionary settings into thematic sections, while the right-hand panel contains the dictionary creation form, where users define core metadata such as label, language, and description. This design supports multilingual projects, collaborative workflows, and immediate synchronization with RDF data stored in GraphDB.
Dictionary entries are listed with essential information (lemma, part of speech, and language), and can be expanded to reveal morphological forms and lexical senses. For each entry, a set of action icons allows users to switch between consultation and editing modes, manage editorial status (e.g. working, completed, reviewed), or perform structural operations such as deletion. The interface also allows lexicographers to create new dictionary entries directly from this panel, facilitating continuous lexicographic work.
When a dictionary entry is selected from the left-hand panel (Figure 2*), the central workspace provides section-based access to its main components, i.e. definitions, government pattern, lexical functions, and usage examples.
DEC Facile supports the linking of lexical senses within definitions: during editing, lexicographers can select an existing sense from a controlled autocompleted list and insert it directly into the definition. As shown in Figure 2*, the definition of the entry venčiavonė contains cross-references to other lexical senses already encoded in the dictionary; in the RDF model, these links are represented through rdfs:seeAlso. This mechanism reduces inconsistencies and encoding errors, while enabling the computational reuse of semantic relations and providing immediate visual feedback through a definition preview.
Lexical functions are instantiated by linking target lexical senses already defined in the dictionary, rather than by manually entering strings. Lexical relations are therefore represented at sense level, improving semantic coherence and interoperability. A dedicated sense-management interface (Figure 3*) supports the hierarchical organization of senses according to graded semantic distance, in line with ECD conventions. Lexicographers can assign levels, adjust ordering, and visualize the structure in real time.
DEC Facile also provides a dictionary-style read-only view (Figure 4*), which reproduces the layout of traditional printed ECDs and displays senses, definitions, lexical functions, and examples. The resulting ECDs are represented in RDF within the OntoLex-Lemon ecosystem, stored in GraphDB, and exportable to LLOD endpoints. The tool is currently used in Old words for a New World and the RUT project.
- Figures see Book of Abstracts
Speakers: Abdelouahab Masbah (Istituto di Linguistica Computazionale “A. Zampolli”), Andrea Bellandi (Istituto di Linguistica Computazionale “A. Zampolli”), Silvia Piccini (Istituto di Linguistica Computazionale “A. Zampolli”) -
61
From Text to Terms: TermScout, an AI Module for Early Terminology Processing
This paper presents TermScout, a GPT-based tool developed to support the initial stage of terminology management by identifying candidate terms in specialised texts. The study evaluates TermScout using a corpus focused on cochlear implantation, assessing both English term extraction and the generation of Croatian equivalents. TermScout identified 100 English candidate terms from a specialised textbook, while Croatian equivalents were generated either directly by TermScout or through a subsequent process utilising TermAI. An independent term validation tool was used to assess all results. The tool demonstrated strong performance in English term extraction, achieving a validity rate of 92%. In contrast, generating Croatian equivalents proved more challenging, with validity rates of 22% for TermScout and 29% for TermAI. These findings indicate that concept identification and bilingual term generation are distinct tasks that require distinct linguistic and terminological expertise. The study underscores the potential of AI-assisted workflows for terminology management, while emphasising the continued need for expert validation.
Speakers: Bruno Nahod (Institute for the Croatian Language), Marijana Tuta Dujmović (Polyclinic for Hearing and Speech Rehabilitation SUVAG) -
62
Inflected-Form Search in Ancient Greek Dictionaries: A Form-to-Lemma Access Model
Dictionary consultation in Ancient Greek presupposes the ability to reconstruct canonical lemmas from highly inflected word forms, a task that can be challenging for users at any level. The lemma-based access model was a material constraint in printed dictionaries, but it can be addressed in digital dictionaries by exploiting the possibilities offered by computational morphology. The aim of this prototype is to link inflected Ancient Greek word forms to their corresponding dictionary lemmas, allowing users to search directly from inflected forms rather than exclusively from canonical lemmas. Although presented as a prototype, the model is designed to be reusable and applicable to any digital Ancient Greek dictionary.
Tools enabling morphological analysis of Ancient Greek forms already exist, most notably the Greek Word Study Tool of Perseus Digital Library and the Morpho function on Logeion. However, in both Logeion and the Perseus Digital Library, dictionary consultation and morphological analysis remain functionally separate. In Logeion, the dictionary interface allows access only to forms attested in its database, while morphological analysis is provided through a separate tool (Morpho). Similarly, in the Perseus Digital Library, the Greek Word Study Tool provides automatic lemmatization and analysis, while dictionary resources such as the Liddell-Scott-Jones Greek-English Lexicon are accessed independently, without systematic interaction between the two components. Although the output of morphological analysers typically provides a lemma linked to a dictionary entry, dictionary access and morphological analysis are presented through separate interfaces and rely on independent datasets: dictionary access is based exclusively on the headword inventory, while the morphological analyser operates independently of this inventory.
The model addresses this limitation by enabling dictionary access through inflected forms rather than exclusively through canonical lemmas. When an inflected word form is entered, the system first searches it within a curated dictionary inventory, which includes both canonical lemmas and selected inflected forms. If the input is found, the system directly returns the corresponding lemma. Only if no match is found is the form processed through an automatic morphological analyser.
At the core of the prototype lies a curated dataset linking inflected forms to dictionary-approved lemmas and their morphological analysis. In its current implementation, the dataset does not aim to represent a full dictionary inventory, but focuses on Ancient Greek polystemic (multi-stem) verbs, a particularly challenging category due to stem alternation and irregular morphology. This restricted scope allows the prototype to address a well-defined morphological problem while testing the underlying lexicographic model.
From the user’s perspective, the prototype processes each query through a sequential workflow. When an inflected word form is entered, the system first attempts to match it against the curated dataset. If a match is found, the corresponding lemma and morphological analysis are returned directly. If no match is found, the input is passed to an automatic morphological analyser, which generates one or more candidate lemmas that can be linked to dictionary entries.
When the user enters a form, the prototype provides three possible types of results. A GOLD result is returned when the input form is attested in the curated dataset. An ACCEPTABLE result is returned when the form is not present in the dataset but the automatic analysis proposes a lemma that is included in the controlled inventory. An EXTERNAL result is returned when the automatic analysis proposes a lemma that is not attested in the dataset. These latter results are explicitly marked as automatically generated. This classification makes the origin of the result explicit, allowing users to distinguish between data derived from the curated dataset and data generated through automatic analysis, and thus to navigate dictionary access more transparently.
From a technical point of view, the prototype consists of two main components: a curated dataset of inflected forms linked to their lemmas and their morphological analysis, and an automatic morphological analyser, specifically Morpheus, used as a fallback. The dataset is stored as a structured table, while the analyser is queried only when no match is found in the dataset. The prototype has been implemented in Python, with Unicode normalization for polytonic Greek and a simple web-based interface developed in HTML.
This prototype demonstrates how a form-to-lemma access model can be applied to Ancient Greek, combining curated lexical data and automatic analysis to enable inflected-form search while preserving a clear distinction between different types of results. It is currently implemented as a local system and is not yet integrated into an online dictionary environment. Future work will focus on integrating the model into a fully functional digital dictionary and testing its application on a larger lexical inventory.
Speaker: Elena Castaldi (EMLex) -
63
Multifunctional Function Words in Upper Sorbian–Czech Dictionary Mudra 2.0
This paper focuses on the lexicographic treatment of multifunctional function words in the emerging Upper Sorbian Czech dictionary Mudra 2.0. As Upper Sorbian lexicography remains strongly dependent on German and lacks a monolingual explanatory dictionary, the compilation of bilingual Upper Sorbian “non German” dictionaries poses specific methodological challenges. These are particularly evident in the treatment of expressions with predominantly grammatical or pragmatic functions, such as conjunctions, modal words, and intensifiers. These expressions are often characterised by polysemy and multifunctionality resulting from grammaticalization processes, and their meanings are highly context dependent. The treatment of these expressions in monolingual dictionaries has been inconsistent so far, whereas in bilingual dictionaries it is typically limited to providing translation equivalents. In this paper, we discuss the various possibilities and limitations of lexicographic descriptions of function words in the context of the minority language. Through a case study of the Upper Sorbian expressions docyła (‘completely’) and cyle (‘quite’), we analyse their semantic and pragmatic properties in comparison to their equivalents in Czech and German. Based on this analysis, we present the lexicographic approach adopted in the dictionary, including sample entries.
Speakers: Katja Brankačkec (Czech Academy of Sciences), Barbora Štěpánková (Charles University) -
64
User-Oriented Specialized Digital Lexicography and Terminological Accessibility: The LexiC Platform Applied to the Glossary of Brazilian Migration Legislation
The technical and legal language of Brazilian migration legislation constitutes a significant barrier to information access for migrants and refugees in situations of social, linguistic, and legal vulnerability. This article presents the LexiC Platform, a web-based computational infrastructure for the collaborative management of terminological data, applied to the Computerized Terminological Glossary of Brazilian Migration Legislation. The research question guiding the proposal is how terminological data extracted through well-established corpus analysis tools can be systematized in a collaborative digital environment so as to preserve methodological consistency, traceability, and the integrity of conceptual relations across a distributed terminographic workflow. The proposal draws on the communicative approach to Terminology (Cabré, 1999, 2005), user-oriented functional lexicography (Tarp, 2008, 2012), and Corpus Linguistics methodologies (Biber, Conrad & Reppen, 1998). The Platform organizes terminographic work into interdependent workflows, covering structured data entry, methodological validation, collaborative review, and automated generation of terminographic products, operationalized through computerized records with more than 60 standardized fields. The specialized corpus comprises approximately 150,000 words, covers laws and decrees published between 2017 and 2025, and includes around 270 terms at different stages of description. Results indicate that terminological rigor, collaborative infrastructure, and user orientation produce resources that are accessible, and socially relevant.
Speakers: Flávia de Oliveira Maia-Pires (Universidade de Brasília), Silvio Licurgo Pires (Universidade de Brasília)
-
60
-
65
Social programme Meeting Points: see website
Meeting Points: see website
Only for registered participants.
If you want to join a group and haven't registered, please get in contact with the local organisers.
More information on our conference website
-
-
-
Lexical Semantics, Neology & Phraseology 01 | Sitzungssaal (01 | Sitzungssaal, OeAW Main Seat, 1st floor)
01 | Sitzungssaal
01 | Sitzungssaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Krzysztof Nowak (Institute of Polish Language PAS)-
66
The Interactional Lexicon of Spoken German: A Functional, Audio-Based Lexicographic Resource
This paper introduces The interactional lexicon, an online lexicographic resource for lexical units in German talk-in-interaction. The resource’s core contribution is a function-first entry model that complements form-based access. Grounded in Interactional Linguistics, all entries are derived from empirical, peer-reviewed studies and document units that are frequent or distinctively used in spoken German, with transcripts from naturally occurring conversations and access to audio/video recordings (FOLK-corpus, IDS Mannheim). To accommodate the organizational specificities of interaction, we propose a compact entry schema: headword and function; position in the turn; sequential position; a corpus-linked example and an analysis thereof. Further categories for prosody and activity framework are optional. A dual access structure supports lookup by form (alphabetical headwords) and by function (e.g., request, response, repair-initiation), based on a controlled functional taxonomy aligned with core concepts of Conversation Analysis. The online resource responds to a documentation gap for German. Lexical units in spoken language, particularly in social interaction, have not yet been brought together in a single resource that serves users with diverse linguistic backgrounds – from linguists to L2-German teachers and advanced students at universities worldwide. The resource will be published online with a restricted number of entries in early 2027, following a user study.
Speaker: Sophia Fiedler (Leibniz Institute for the German Language) -
67
Beyond Binary Metaphor Detection: Dataset-Driven Evaluation of Conceptual Metaphor Type Recognition in Croatian and Slovene
Large language models (LLMs) continue to perform suboptimally in tasks involving metaphor detection and explanation (Devlin et al., 2019; Liu et al., 2022; Klemen & Robnik Šikonja, 2023, Kim et al., 2023; He et al., 2023; Despot, Ostroški Anić, & Veale, 2023; Tong et al., 2024; Puraivan et al., 2024; Yang et al., 2025; Chen & Wang, 2025). Core challenges lie in the inherent semantic complexity of metaphorical expressions, and in the dearth of structured metaphor corpora. In response to this limitation, we introduce CroSloMet, a novel structured metaphor dataset for two South Slavic languages, Croatian and Slovene, designed to support both metaphor identification and explanation tasks (Ge et al., 2023, Despot et al., 2025). CroSloMet comprises over 1,120 metaphorical sentences and 1,120 matched literal sentences for each language, annotated with semantic metadata. Each data point is aligned with corresponding conceptual metaphors, multi-word expressions, canonical forms, and literal usages, providing a multi-layered annotation scheme that goes beyond surface metaphor tagging to encode conceptual and linguistic structure. This resource is grounded in the MetaNet.HR framework (Despot et al., 2019) enabling researchers to explore generality and specificity in metaphoric expressions. The dataset supports metaphor identification tasks—the classification of sentences as metaphorical vs. literal through annotated pairs that preserve lexical overlap between figurative and non-figurative contexts. Moreover, it facilitates metaphor explanation (conceptual metaphor type recognition, where the conceptual metaphor name is used as a proxy for explanation), allowing evaluation of model output not merely on binary classification but on the degree to which generated conceptual metaphor type capture intended conceptual mappings. CroSloMet’s dual annotation quality makes it suitable for corpus-based analysis, cognitive linguistics studies, and the development of metaphor understanding modules in NLP systems. To demonstrate the dataset’s utility, the original work reports preliminary experiments using a fine-tuned CroSloEngual BERT model for metaphor classification, achieving an accuracy of 88.5%, and an evaluation of LLaMA 3-8B for metaphor detection. While classification results were promising, strict exact-match evaluation for explanation generation (conceptual metaphor type recognition) yielded low scores, revealing a gap between model outputs and human interpretive expectations despite the semantic validity of many generated texts. This discrepancy pointed to the need for improved evaluation metrics that capture semantic similarity and interpretive nuance.
Building on these insights, the current paper proposes to extend CroSloMet with a multi-level validation framework based on the comparison of manual annotation to natural language inference (NLI), semantic similarity scoring, and LLM-based judgment to assess metaphor explanations more holistically. By integrating human expert judgments with automatic evaluation signals, the proposed research seeks to close the gap between computational performance and cognitive plausibility in metaphor interpretation. This study leverages this dataset and evaluation framework to address several core research questions: How can structured metaphor annotations be exploited to improve cross-lingual metaphor understanding? What evaluation strategies best capture the quality of metaphor explanations in LLM-generated text? And how can conceptual metaphors be operationalized to support both symbolic and distributional approaches to figurative language modelling? To answer these questions, the paper outlines an experimental pipeline that integrates CroSloMet with advanced representation learning techniques, such as transformer architectures and semantic embedding spaces calibrated for figurative meaning. A key aspect of this research is the incorporation of semantic generalization layers within explanation models that can abstract over lexical variation to capture deep conceptual mappings. We take advantage of the hierarchical organization of conceptual metaphors in the MetaNet.HR database (Despot et al., 2019) to refine the evaluation of model-generated metaphor explanations. Instead of relying solely on exact matches between predicted and gold-standard conceptual metaphors, we implement a graded evaluation scheme that accounts for metaphor hierarchy and allows us to calculate conceptual distance or similarity between predicted and reference metaphors even if its formulation is more abstract or differently phrased. By combining frame-based representations with contextualized embeddings, the study aims to evaluate metaphor understanding at the level of conceptual pattern recognition – a capability that bridges computational modelling and cognitive theories of metaphor (Lakoff and Johnson, 1980; 1999 – for an overview of different and more recent approaches see Dancygier & Sweetser, 2014 and Despot, 2024). The dataset’s parallel nature further enables cross-lingual transfer experiments, where models trained on one language can be tested on another. To ensure that CroSloMet can serve as a shared resource for language technology research, supporting tasks such as frame mapping across languages and constructive evaluation of metaphor explanation models aligned with human interpretive criteria, CroSloMet is aligned with standardized semantic tagging schemes, and can be exported in interoperable formats providing benchmarks for corpus-based metrics.
This paper presents CroSloMet as both a robust dataset for metaphor research in under-resourced languages and an impetus for future method development in metaphor interpretation. By articulating a clear research agenda centered on structured annotations and nuanced evaluation, the work addresses persistent gaps in metaphor understanding with computational models. The proposed multi-level validation framework and cross-lingual experimental design position CroSloMet as a contribution to lexicography, computational linguistics, and cognitive semantics, making it relevant for figurative language modelling.
Speakers: Kristina Š. Despot (Institute for the Croatian Language), Ana Ostroški Anić (Institute for the Croatian Language), Matej Klemen (University of Ljubljana) -
68
(UG)TPEX: A (Usage-Guided) Typical Example Score
The aim of this paper is to explore methods to utilize the linguistic knowledge accumulated in language models to evaluate corpus sentences as potential examples in dictionary entries. We examine the TPEX (Typical Example) measure in three variants in the task of assessing the typicality of a sentence based on how probable the appearance of the headword in the given context is, as per the selected language model. Then we present an extension, the UGTPEX (Usage-Guided Typical Example) measure, in six variants, where the contextual typicality is supplemented by data reflecting the similarity of the headword’s contextualized embedding to representations seen in known-good (dictionary-derived, sense-specific) examples. The combination of the two components, i.e. the contextual typicality and the embedding-based similarity, results in a score that is sensitive to both frequent usage patterns and the usage categories (sense categories) established by the lexicographer. We will examine the applicability of these measures in case studies involving five headwords using dictionary example sentences and corpus samples. The proposed measures are for ordering concordance lines, in combination with other ranking methods, as part of the lexicographic workflow.
Speaker: Ágoston Tóth (University of Debrecen)
-
66
-
Multilingualism, Language Varieties & Contact Lexicography 01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)
01 | Johannessaal
01 | Johannessaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Lars Trap-Jensen (Society for Danish Language and Literature)-
69
Reframing Wan (完), Hao (好), and Dao (到) in Chinese–Korean Bilingual Lexicography: A Corpus-Based Cognitive Analysis
Introduction
Chinese resultative complements such as Wan (完), Hao (好), and Dao (到) are frequently treated in bilingual dictionaries as generalized markers of completion or result. However, existing lexicographic descriptions often rely on simplified lexical glosses such as ‘finish’, ‘good’, or ‘arrive’, obscuring important cognitive-functional distinctions among these complements. In actual usage, the three markers exhibit differentiated event-structural patterns, semantic preferences, and translational correspondences. This study reexamines Wan, Hao, and Dao through a Chinese–Korean parallel corpus and argues for a cognitively informed, corpus-based approach to bilingual lexicography, with implications for learner-oriented and AI-assisted dictionary design.Theoretical Framework
Following corpus-driven approaches to bilingual lexicography (Atkins & Rundell, 2008), this study adopts a cognitive-functional perspective combining cognitive scanning and image-schema theory. Langacker’s notion of cognitive scanning explains how speakers construe events (Langacker, 1987, pp. 244–249), while Johnson’s image schemas, such as CONTAINMENT and PATH (Johnson, 1987, pp. 21–30), provide a framework for modeling boundedness, resultant configuration, and goal attainment.Figure 1* presents the cognitive-image-schematic profiles of the three complements. Wan profiles a bounded process brought to terminal exhaustion, Hao profiles the transformation of a disordered or incomplete state into an acceptable configuration, and Dao profiles movement or access toward a spatial, perceptual, or abstract target.
Methodology
The dataset consists of 871 valid instances of Wan, Hao, and Dao extracted from the AI Hub Chinese–Korean parallel corpus of broadcast discourse. Each example was manually annotated for verb class, verb supercategory, object type, object-role structure, functional group, aspectual subtype, and Korean translational correspondence. Frequency distributions, percentage comparisons, and χ² tests were used to identify statistically differentiated semantic-functional tendencies.Findings
The corpus analysis reveals that Wan (完), Hao (好), and Dao (到) form statistically differentiated cognitive-functional profiles across verb classes, object preferences, functional distributions, and Korean translational correspondences.First, the three complements exhibit significantly different verb-class distributions (χ², p < .001). Wan preferentially co-occurs with consumption, communication, and creation-related verbs, reflecting exhaustive completion and resource depletion. Patterns such as 吃完, 用完, and 做完 profile bounded completion and sequential event closure. In contrast, Hao frequently occurs with action, arrangement, and cognition-related verbs, indicating successful configuration or readiness, as in 准备好 and 安排好. Dao shows strong associations with perception, cognition, and motion verbs, including 看到, 听到, and 找到, reflecting attainment toward perceptual, cognitive, or goal-oriented endpoints.
Second, functional-group distributions also differ significantly (χ², p < .001). Wan is associated with SEQUENTIAL and CONFIGURATION functions, and also appears in MODALITY-related contexts involving obligation release, such as 做完就不用…. Hao is strongly associated with DIRECTIVE and STATUS_NOTIFICATION functions, signaling preparation, readiness, or resultant-state confirmation. Dao is linked to EVALUATION and STATUS functions, highlighting attainment and realization of abstract or perceptual targets.
Third, object-type distributions reinforce these distinctions. Wan tends to co-occur with bounded eventive and discourse-related objects, including tasks, speech content, and information. Hao predominantly combines with manipulable concrete entities, reflecting state-configuration processes. Dao frequently occurs with abstract, discourse-related, and locational objects, emphasizing endpoint-oriented attainment.
Finally, Korean translational correspondences also reflect these profiles. Wan is often translated as ‘-다’, ‘끝내다’, or ‘다 해버리다’; Hao as ‘준비되다’, ‘잘’, or ‘-해 놓다’; and Dao as ‘도달하다’, ‘찾다’, ‘보다’, or ‘듣다’. These patterns show that the semantic distinctions among Wan, Hao, and Dao are reflected in both Chinese corpus distributions and bilingual translational behavior.
Based on these distributional patterns, the study proposes an integrated [V-RC-O] selection-constraint model in which the three complements interact differently with verb semantics, object boundedness, and image-schematic patterns.Lexicographic Implications and Conclusion
These findings suggest that bilingual dictionaries should not reduce Wan, Hao, and Dao to isolated glosses such as ‘finish’, ‘good’, or ‘arrive’. Instead, entries should reflect their distinct cognitive-functional profiles: Wan as exhaustive completion, Hao as readiness-oriented configuration, and Dao as attainment-oriented realization. By integrating corpus distributions, functional tendencies, and Korean translational mappings, this study proposes a cognitively informed, corpus-based model for Chinese–Korean bilingual lexicography.
* see Book of Abstracts
Speaker: Seulki Lee (Chungbuk National University) -
70
From Lexical Theory to AI Translation: Procomplement Verbs and ChatGPT
Procomplement verbs (PCVs) are a highly frequent yet understudied class of semantically idiosyncratic verb–clitic constructions found across Romance languages. Despite their relevance for everyday communication, they remain only marginally represented in grammars and dictionaries. Building on recent theoretical work on Italian PCVs and drawing on Generative Lexicon theory, this study investigates the relationship between the syntactic–semantic properties of PCVs and their treatment by large language models. A dataset of 160 highly frequent Italian PCVs, primarily extracted from GRADIT and comprising 256 usage examples, was submitted to ChatGPT 5.2 for translation from Italian into English. The analysis focuses on the model’s ability to account for the semantic, pragmatic, and phraseological complexity of these constructions. The results show that translation quality is influenced less by the specific subclass of PCVs than by transversal properties such as connotation and the pragmatic functions encoded by individual utterances. While ChatGPT generally succeeds in translating strongly conventionalised and highly idiomatic expressions, it frequently encounters difficulties with polysemous PCVs, context-sensitive pragmatic meanings, and morphologically complex verb–clitic combinations. The findings highlight the value of a fine-grained syntactic–semantic classification of PCVs for bilingual lexicography and suggest potential applications for the improvement of AI-assisted translation systems.
Speakers: Valeria Caruso (Università degli Studi di Napoli L'Orientale), Geraint Paul Rees (Pompeu Fabra University) -
71
Bilingual Lexicography Under Resource Asymmetry: Designing Yoruba–German Dictionary Articles in Dialogue with Galician–German
This paper addresses the challenge of designing bilingual dictionary articles under resource asymmetry by focusing on the Yoruba–German language pair and drawing methodological insights from the more developed Galician–German ecosystem. It proposes a resource-aware model of dictionary article composition that explicitly accounts for limited evidence by making uncertainty visible, using heterogeneous sources (corpora, pivot resources, expert input). Our approach involves dictionary-critical analyses of both Yoruba–German and Galician–German lexicographic resources and the transfer of microstructural design strategies from Galician to Yoruba. Case studies include Yoruba–German prototype articles (e.g. the verb ṣe and the nouns ayé and akọ́nimọ̀ọ́gbá) developed from these findings, incorporating sense divisions, translation equivalents, examples, grammatical/constructional notes and cross-references. The study also demonstrates a human-led use of AI tools to enrich example sentences and collocations in the prototypes. Main contributions are the concept of a resource-aware dictionary article and a cross-pair methodological dialogue, operationalised in prototype articles. These contributions provide practical guidelines for lexicographers of under-resourced languages and suggest scalable, AI-assisted workflows for bilingual lexicography under unequal resource conditions.
Speakers: Luke Omoyemi Akinremi (Universidade de Santiago de Compostela), Elena Martín-Cancela (Universidade da Coruña)
-
69
-
Specialised & Terminological Lexicography 00 | Anton Zeilinger Salon (00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor)
00 | Anton Zeilinger Salon
00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Alexander Geyken (Berlin-Brandenburg Academy of Sciences and Humanities)-
72
Automatic Drafting of Terminological Dictionaries with One-Click Dictionary for 34 Languages
We present an evaluation of a system for automatic drafting of terminological dictionaries using the One-Click Dictionary feature of Sketch Engine. We bootstrapped five small corpora in Arabic, Czech, English, German and Spanish and evaluated the performance of the methodology on these five corpora yielding five terminological dictionaries for five different terminological domains; the system currently supports 34 languages in total.
Speakers: Marek Blahuš (Lexical Computing), Miloš Jakubíček (Lexical Computing), Katarína Petreková (Lexical Computing), Yingtian Wu (Lexical Computing) -
73
User-Oriented Specialized Lexicography: Corpus-Driven Models for Targeted Dictionary Design
Lexicographic practice has often balanced comprehensive coverage with systematic ordering, privileging alphabetical arrangement and encyclopedic completeness over clearly defined user needs. This paper argues that user-oriented specialized dictionaries, supported by corpus-driven and AI-augmented methods, can better serve users who require targeted lexical knowledge rather than exhaustive inventories. The study uses Śabdakalpadruma as a Sanskrit lexicographic case study and reports its conversion into a structured digital corpus of 42,824 headword entries through OCR, post-OCR cleaning, Excel inspection, SQLite structuring, JSON representation, and searchable HTML implementation. Building on this corpus, the paper develops a pilot rule-based recommender model for ritual vocabulary. Twenty ritual query terms were searched in headword and definition fields, and the extracted entries were automatically annotated through semantic tagging, co-occurrence detection, relation scoring, confidence labelling, and relevance classification. A stratified validation sample of 100 entries, with five entries from each query term, was manually evaluated. The results show that 67% of annotations were correct, 19% usable with correction, and 14% wrong. The study demonstrates a staged pathway from digitized lexicographic corpus to rule-based recommendation, human-validated training data, and future AI-assisted semantic prediction.
Speaker: Deepak Solanki (Jawaharlal Nehru University, New Delhi)
-
72
-
Corpus Linguistics & Corpus-Based Resources 01 | Sitzungssaal (01 | Sitzungssaal, OeAW Main Seat, 1st floor)
01 | Sitzungssaal
01 | Sitzungssaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Krzysztof Nowak (Institute of Polish Language PAS)-
74
The Vocabulary of Mental Health: Exploitation of Specialized Diachronic Corpora with Artificial Intelligence for Qualitative Semantic Analysis
This article explores the diachronic semantic evolution of mental health vocabulary, focusing specifically on the concept of hysteria across Spanish and German specialised corpora. Drawing on Diachronic Prototype Semantics, the paper examines how semantic changes reflect shifts in radial, hierarchical networks where meanings are updated according to communicative needs. The key methodological aspect is the use of the Gemini large language model to assist in the qualitative analysis of the MentEs and DeMens corpora contexts. This AI-supported approach enables the large-scale categorisation of contexts into six thematic areas: aetiology, clinical manifestations, differential diagnosis, treatment, conceptual debate, and non-clinical uses. The findings indicate that historical scientific discourse was overwhelmingly preoccupied with conceptual delimitation and identification of hysteria (representing 86–94% of occurrences) rather than therapeutic measures. Beyond the lexical and semantic analysis of hysteria, the validity of the methodology is presented as the primary finding. The author concludes that AI-assisted processing facilitates a comprehensive bridge between quantitative data and qualitative interpretation, allowing researchers to document prototypical modulation, for instance, how specific meanings gain or lose prominence depending on the historical and social context.
Speaker: Pol Garriga Martínez (Pompeu Fabra University) -
75
Collaborative Lexicography for Heritage Language Learning
This study explores the potential of collaborative dictionary compilation to promote vocabulary retention and contributes to research in Chinese heritage language education. A pilot intervention was conducted with 22 second-generation Chinese students attending a complementary school in Italy. The participants studied literary vocabulary from Lu Xun’s Village Opera and collaboratively compiled dictionary entries using Lexonomy. To assess the broader effectiveness of the protocol and compensate for the absence of a control group of heritage learners in Italy, the framework was replicated with two age-matched groups in China: an experimental group of 50 students engaged in dictionary-compilation activities and a control group of 53 students receiving traditional teacher-led instruction. Vocabulary knowledge was assessed through a pre-test and a delayed post-test administered three months after the intervention. A mixed-effects logistic regression analysis showed that the experimental group was approximately twice as likely as the control group to improve its retention of the literary vocabulary studied. The findings provide preliminary evidence that collaborative lexicography can support medium-term literary vocabulary acquisition and culturally oriented language learning, although larger-scale studies involving students from complementary schools are needed.
Speakers: Zongyuan Wang (Università degli Studi di Napoli L'Orientale), Valeria Caruso (Università degli Studi di Napoli L'Orientale), Anna De Meo (Università degli Studi di Napoli L'Orientale) -
76
Corpus-Based and AI-Assisted Design of a Chinese–English Business Dictionary
Recent research in lexicography has shown that the integration of generative AI into dictionary-making can work well for bilingual tasks (Chen, 2025; de Schryver, 2025; Han & de Schryver, 2025). The present research presents a corpus-based study of Chinese–English parallel business news and explores its implications for bilingual business lexicography in the age of AI. The study is based on a purpose-built Chinese–English parallel corpus drawn from the “Bilingual News” section of China Daily, comprising over 36,000 Chinese tokens and 42,000 English tokens. Although the corpus is relatively small and restricted to a single source, it includes all available bilingual business news published by this major official outlet in 2024. It can therefore be seen as a complete dataset within a clearly defined domain.
The English texts are typically original reports aimed at international audiences, while the Chinese texts are translations for domestic readers. Using the Sketch Engine, the analysis combines single-word keyword extraction and multi-word term extraction, with comparisons against large general-news reference corpora in both languages. A minimum frequency threshold of 5 is applied, and items with very low keyness scores are excluded, as they do not show domain-specific salience. This approach makes it possible to identify lexically salient items that are characteristic of contemporary business and business-related policy discourse.
The findings show clear cross-linguistic asymmetries. Chinese business texts tend to be data-driven and domestically oriented, favouring compact noun compounds, numerical indicators, and institution-specific terminology that compress complex economic relations into dense nominal structures. English texts, by contrast, rely more heavily on abstract policy framing, analytic syntactic constructions, and standardised acronyms and compound expressions, with repeated use of established policy phrases. These asymmetries should be interpreted with caution. As the English texts are produced within Chinese official media contexts, they may reflect patterns of China English and may differ from the expressions typically used in media from English-speaking countries, rather than purely language-internal variation.
Across both languages, multi-word units show markedly higher keyness values than single-word items. This suggests that meaning in this domain is primarily conveyed through fixed expressions, collocations, and policy phrases. While some high-frequency Chinese–English multi-word pairs have stabilised as standard equivalents, many Chinese business terms do not have direct lexical counterparts in English. These items are often culture- or institution-specific and therefore call for a paraphrastic or definition-based treatment rather than a word-for-word translation.
Building on these findings, the study proposes several lexicographic implications. First, it argues for prioritising multi-word units as dictionary headwords, as these units capture recurrent usage and encode domain-specific concepts more directly. Second, it supports a morpheme-plus-collocation entry structure. Under this structure, core lexical elements are systematically linked to their most frequent and productive multi-word patterns. Third, it advances a dual-track dictionary model. This model is designed to accommodate both the production-oriented needs of Chinese users and the reception-oriented needs of foreign users, allowing lexicographic content to be organised according to different user situations and functions.
To explore the role of genAI in contemporary lexicographic practice, the study further includes an AI chatbot, DeepSeek-V3, to generate preliminary dictionary entries. These are generated entirely by AI, with prompts iteratively refined to improve the quality and structure of the output. The prompt is shown in Addendum A. DeepSeek-V3 was provided with compilation notes, as seen in Addendum B. An example of generated entries is shown in Addendum C.
The study compares AI-generated entries with human-produced lexical resources. When accessed via dictionary apps, reference works such as the New Century Chinese–English / English–Chinese Dictionary (2026) and the Youdao Dictionary (2026) mainly provide equivalents, but do not offer full definitions nor usage examples. Overall, then, the present study demonstrates how corpus evidence and genAI chatbots can jointly inform the next-generation bilingual specialised dictionary design. By integrating corpus-based insights and AI assistance, the paper contributes to ongoing discussions on dictionary-making processes in the age of AI.
Speakers: Wanjing Han (Ghent University), Gilles-Maurice de Schryver (Ghent University)
-
74
-
Digital Lexicography: Design, Data Modelling & Theory 00 | Anton Zeilinger Salon (00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor)
00 | Anton Zeilinger Salon
00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Alexander Geyken (Berlin-Brandenburg Academy of Sciences and Humanities)-
77
Management of Large Lexicographic Projects with Lexonomy
Modern lexicographic workflows increasingly combine automatic dictionary drafting from large corpora with extensive human post-editing and validation. Advances in corpus linguistics and natural language processing make it possible to generate comprehensive dictionary drafts automatically, but the quality of the final product still depends on systematic human editorial work. As a result, contemporary dictionary projects are characterized by multiple iterative cycles in which editors assess, correct, enrich, and approve automatically generated entries. While this approach significantly accelerates dictionary production, it also introduces new challenges in organizing editorial work, maintaining consistency, and supervising progress across large editorial teams.
In this paper we present Lexonomy, a system designed to support the management of large lexicographic projects following this draft-and-post-edit paradigm. The system is tailored for dictionaries whose editors working in parallel. It supports the division of editorial labour into clearly defined tasks where each editor or team focuses on a concrete entry element according to their expertise and lexicographic experience. Lexonomy enables editors to focus on their lexicographic decisions while ensuring that individual contributions fit coherently into the overall editorial plan.
A key objective of the system is to provide the editor-in-chief with a clear and continuously updated overview of the dictionary creation process. By integrating support for collaborative editing with tools for project supervision, Lexonomy facilitates efficient management of large-scale dictionary projects and helps ensure that automatically generated drafts are systematically transformed into high-quality, human-validated lexicographic resources.
Speakers: Ondřej Matuška (Lexical Computing), Miloš Jakubíček (Lexical Computing), Vojtěch Kovář (Lexical Computing) -
78
Toward a Narrative-Oriented Dictionary: Computational Modelling of Youth Language in Socially Vulnerable Contexts
This paper presents a computational narrative dictionary of youth language in contexts of high socio-economic vulnerability, where “narrative” is understood, following Bruner’s distinction between paradigmatic and narrative modes of meaning-making (Bruner, 1986), as meaning grounded in experience, examples, values, and self-positioning rather than only in categorisation. The resource is designed for social and community workers, educators, and psychologists who work with young people in fragile settings and need to interpret not only what words denote, but also the experiences, values, and expectations that speakers attach to them. Its aim is therefore both descriptive and practical: to document situated meanings and make them usable in social and educational intervention. The dictionary does not replace standard lexicographical description; rather, it complements it with information that is usually excluded from dictionary entries but is central to intervention work: emotions, examples, social roles, expectations, and forms of self-positioning.
The dictionary is being developed within the project “Futuri (im)possibili: diagnosing the present and imagining the future” with the youth of Caivano , carried out in the metropolitan area of Naples, Italy. The project adopts an Action-Research approach (Elliot, 1993; Lewin 1946), combining empirical investigation with community-based social intervention. In line with user-oriented lexicography, the dictionary addresses concrete needs: it supports professionals in recognising semantic shifts and resemantisations, and in designing interventions that can help young people imagine alternative and more positive futures.
Building on Bruner’s distinction, the resource focuses on socially sensitive Italian terms connected with power, deviance, gender, violence, and identity, such as ‘aborto’, ‘criminale’, ‘femmina’, ‘infame’, and ‘soldi’. While standard dictionaries usually formalise meanings through paradigmatic definitions based on categorisation, synonymy, and semantic inclusion, the proposed dictionary also records associative networks, lived experience, and identity-related narratives. For instance, the word ‘femmina’ may activate biological or dictionary-like meanings, while also eliciting narratives of strength and care, but also inferiority, sexualisation, and violence. Meaning is thus treated as a cognitive and cultural construct shaped by shared representations, stereotypes, and lived experience.
The empirical basis consists of 58 semi-structured interviews (Adams, 2015), corresponding to 61 recordings and over 30,000 tokens of transcribed speech, collected mainly in street contexts in Caivano by outreach volunteers. Participants were aged 12 to 30, with a smaller group over 30, and were mostly residents of the area. Interviews were audio/video recorded, digitally transcribed, anonymised to ensure privacy protection, and organised in a structured corpus with metadata and unique identifiers. Automated transcription was combined with manual revision to guarantee accuracy. At the time of writing, the dictionary contains 16 lexical entries with both paradigmatic and narrative meanings.
The dictionary has been provided with a computational representation following OntoLex-Lemon , the de facto standard for lexical resources in the Semantic Web. This choice supports FAIR principles (Wilkinson et al., 2016) and facilitates interoperability with resources in the Linguistic Linked Data cloud. Each entry is modelled as a lexicog:Entry linked to an ontolex:LexicalEntry, while meanings are represented as ontolex:LexicalSense. Paradigmatic and narrative senses are encoded as SKOS concepts in a dedicated skos:ConceptScheme and assigned through dc:type. In the case of ‘femmina’ (Figure 1*), the paradigmatic sense is linked to the concept WOMAN, whereas narrative senses are connected to concepts such as STRENGTH and POWER. Evaluative polarity is modelled through the MARL Opinion Ontology by reifying the ontolex:isLexicalizedSenseOf relation.
The model is evaluated through competency questions that reflect expected use cases. These include, for example, retrieving concepts associated with narrative senses for a given entry, listing concepts connected with negative or positive polarity, and identifying all narrative senses related to a specific concept. Rather than treating polarity as a simple label attached to an entry, the model relates it to specific senses and concepts, since the same word may carry different, even conflicting, evaluations depending on the speaker’s narrative frame and social experience.
Future work will connect senses to textual occurrences in the corpus through OntoLex-FrAC and enrich interview metadata using standard vocabularies such as Dublin Core, FOAF, and PROV-O. Integrating lexical, conceptual, textual, and metadata layers will enable cross-layer queries over the dataset, for example to examine which narrative senses and associations emerge in relation to speakers’ age and gender. This can provide evidence-based insights for professionals working with youth in socially fragile contexts.
* see Book of Abstracts
Speakers: Michela Bandini (Istituto di Linguistica Computazionale “A. Zampolli”), Silvia Piccini (Istituto di Linguistica Computazionale “A. Zampolli”), Andrea Bellandi (Istituto di Linguistica Computazionale “A. Zampolli”), Emiliano Giovannetti (Istituto di Linguistica Computazionale “A. Zampolli”) -
79
Optimising Interface Design in Online Lexicographic Products
This paper focuses on certain aspects of interface design in online dictionaries. A discussion of a number of examples of interface design in existing online dictionaries is presented. It is emphasised that successful interface design in transformative lexicographic products should be characterised by a human-centred approach that enables intuitive access to the required data. Dictionary articles should offer users an overview of the available contents. The user should then be able to zoom in on specific data types whilst others could be filtered out so that information on demand can be retrieved. An interdisciplinary approach is needed to optimally utilise the possibilities offered by the online environment to design a dynamic article structure with an interface that gives the users rapid access to specific data. Lexicographic theory should offer an exposition of innovative procedures already prevailing in existing dictionaries. Lexicographic theory should also make proposals that can be applied in the practice to enhance the interface design in order to assist users in experiencing successful dictionary consultation endeavours.
Speakers: Rufus H. Gouws (Stellenbosch University), Theo J.D. Bothma (University of Pretoria)
-
77
-
Historical & Diachronic Lexicography 01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)
01 | Johannessaal
01 | Johannessaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Lars Trap-Jensen (Society for Danish Language and Literature)-
80
Digital Metalexicography as a Tool for Historical (Contact) Linguistics
This contribution presents a data collection on ‘Slavic in Viennese’, i.e. (glottonyms for) Slavic languages, (ethnonyms for) their speakers as well as single linguistic phenomena in German varieties typical for Vienna, that comprises 6,406 articles on 268 lemmas from 30 dialect dictionaries published between 1800 and 2020. In the contribution’s first part, corpus compilation and the underlying data structure are described and results of existing studies on this data collection are discussed. So far, these have allowed for the identification of reception networks between the dictionaries and revealed the reconstruction of stereotypes regarding speakers of South Slavic languages in them. Inspired by these results, the second part of the contribution presents an approach towards the use of the data collection (and similar ones) for the description of integration, variation and change of (West Slavic) loanwords in ‘Viennese’. By means of exploratory cluster analyses, three main clusters of lemma-meaning pairs are identified: one general that includes forms and meanings used throughout the complete observation period, one of historical forms and meanings that were common in the 19th century, and one of innovations from the second half of the 20th century. Building on this and the exemplary comparison of forms and meanings of three Czech loanwords (Kaluppe, Kolatsche, pomali), the contribution suggests that most similar loanwords must have been borrowed before the 19th century, thereby demonstrating the usability of the approach for historical (contact) linguistics.
Speaker: Agnes Kim (University of Ostrava) -
81
Digitization of Adam Patačić's Manuscript Dictionary
Adam Patačić is a Croatian lexicographer (1716–1784), Archbishop of Kalocsa and President of the Royal Council of the University of Buda and a promoter of music with his orchestra and choir. His extensive Dictionarium latino-illyricum et germanicum remained in manuscript, written after the example of Nomenclator omnium rerum by Hadrian Junius (Samardžija, 2019, p. 63).
The mentioned dictionary is a completed manuscript work written in ink, the pages are numbered, and on the inside title page the author is signed as D. ADAMUS L. B. PATACHICH De Zajezda. Although the manuscript does not contain information on the title page indicating the year the work dates from, the data relevant to the dating of the dictionary can be read from the preface, as well as from the title page. According to the data on the title page, the author's name is accompanied by a list of his titles and honors, some of which he achieved only towards the end of his life. In support of this dating, and according to the numbered third page of the dictionary, it can be concluded that the work was bound and completed in the late seventies or early eighties of the 18th century. According to the data from the preface, it can be stated with certainty that the dictionary was created in the period from 1772 to 1779 (Jonke, 1949, pp. 95–96). Patačić's dictionary was written in an era that, together with the previous century, is considered the golden age in the history of the Kajkavian Croatian literary language. It was at this time that major Kajkavian lexicographical works were created – Belostenec's, Jambrešić's and Patačić's dictionaries, as well as the first orthography and the first manuscript Kajkavian grammar (Štebih Golub, 2013, p. 258).
The dictionary is conceptually structured with a macrostructure of 13 thematic areas and several sub-areas, with encyclopedic interpretations in Latin. The headword side is Latin, and the explanations are given in Croatian, i.e. in the Kajkavian literary language, while the equivalent is also given in German. The dictionary contains a rich treasury of Kajkavian literary lexis (for example, kinship, zoological and botanical terminology), as well as the Kajkavian literary language of the 18th century.
The preface contains XIII pages, while the content consists of V pages. The dictionary material covers 1034 numbered pages. The dictionary is divided into four major parts. The first part entitled De Deo Spiritibus, Coelo, Elementi set Homine comprises 13 chapters on 220 pages and provides general information about the extralinguistic reality. The second part is entitled De iis, quae ad Hominis vitam tum recte agendam, tum tuendam pertinent: sive Sacrorum Praesides, Civilia et Militaria, and contains 15 chapters (pp. 221–415) and covers political and military duties, literature, music, and art. The third part is entitled Oeconomica and contains 14 chapters (pp. 416–734) and relates to the economy. The fourth part Technica (pp. 735–1034) consists of 16 chapters and relates to technical sciences.
Although Patačić's dictionary was not used as a corpus in the preparation of the Academy's historical multi-volume Dictionary of the Croatian or Serbian Language (1880–1976), it served as a corpus in the preparation of the Dictionary of the Croatian Kajkavian Literary Language (1984–present, kajkavski.hr) where, in structuring the dictionary entry, confirmations from well-known Kajkavian dictionaries of the 17th and 18th centuries are cited, along with Patačić's manuscript dictionary (Brlobaš and Horvat, 2019, p. 146).
The largest study of the dictionary is Ljudevit Jonke's Dikcionar Adama Patačića (1949) and the more recent one is Marijana Horvat's research on musical terminology in Croatian pre-revival dictionaries (Horvat, 2018, pp. 437–454).
Due to its exceptional importance for the Croatian language in general and the Kajkavian literary language, at the beginning of 2026, at the initiative of the director of the Institute for the Croatian Language, Željko Jozić, and with the help of collaborators on the Dictionary of the Croatian Kajkavian Literary Language project, and with other colleagues from the Department of the History of the Croatian Language, a project was launched to study Patačić's manuscript dictionary. The transliteration of the work has begun with the aim of further studying the extensive and the last historically extremely important Kajkavian dictionary. In the presentation, we will talk about the principles of digitization and the problems we encounter.
We are digitising Adam Patačić’s handwritten Dikcionar, a manuscript exceeding 1,000 pages, by combining philological transcription and handwritten text recognition (HTR) in Transkribus. As initial Ground Truth, we reused the manuscript text transcribed by Ljudevit Jonke in his study ‘Dikcionar’ Adama Patačića: studija iz hrvatske kajkavske leksikografije. This transcription corresponds to 25 manuscript pages of the Dikcionar. The text was imported into Transkribus and manually aligned with the corresponding manuscript images and baselines, yielding a first controlled Ground Truth set. Our very first training attempt—based on an incomplete subset of this Jonke-derived material and without an appropriate base model—produced poor recognition performance (Transkribus-reported CER 58.61%), showing that substantial script-specific adaptation was still needed. We therefore introduced transfer learning and retrained the model using the publicly available base model “Cosimo Bartoli’s Italian Humanistic and Cursive Scripts (1562–1572)” together with the Jonke-based Ground Truth. This markedly improved recognition, lowering the Transkribus CER to 11.52%. Since our research use case tolerates several predictable early-modern graphemic ambiguities, we additionally evaluated performance under a normalisation that ignores case and treats u/v, s/ſ, and æ/ae as equivalent. Under these conditions, an independent CER calculation yielded 6.01%, which we consider an excellent and practically usable starting point for large-scale digitisation. The next phase will proceed iteratively: a team of researchers will expand Ground Truth by correcting model output on newly processed pages, while the model will be periodically retrained to further improve robustness across the manuscript and to support subsequent linguistic and lexicographic analysis of the Dikcionar. This digitisation also marks an exceptional pilot step within the MONOGRAF project (Development and Implementation of a Model for Normalising the Orthography of Old Texts Printed in Latin Script), which otherwise focuses on printed texts but has, on a trial basis, extended its workflow to a handwritten manuscript in this case.
Speakers: Željko Jozić (Institute for the Croatian Language), Marijana Horvat (Institute for the Croatian Language), Martina Horvat (Institute for the Croatian Language) -
82
Canadian English Dictionary: A Progress Report
Canada is long overdue for a new English dictionary. In this presentation, we describe progress to date on the Canadian English Dictionary (CED), a new general dictionary of Canadian English being edited by the recently incorporated Society for Canadian English, with the support of Editors Canada, the University of British Columbia’s Canadian Word Centre, and Queen’s University’s Strathy Language Unit. A brief examination of the last generation of Canadian English dictionaries (Gage Canadian Dictionary [1996], ITP Nelson Dictionary [1997], Canadian Oxford Dictionary, 2nd edition [2004]) reveals the need for an up-to-date resource. Oxford’s second (and last) edition, published in 2004, remains the most recent resource for Canadian dictionary users, indicating that over twenty years of language change remains undocumented.
Currently, Canadian English speakers must turn either to American or British English dictionaries for up-to-date lexicographical information. We need look no further than a recent spelling controversy that flared up in the media to remind ourselves that Canada’s variety of English is distinct, and that it matters to the population: our prime minister’s use of British -ise spellings in government documents has been met with criticism by Editors Canada, who remind readers that Canadians use -ize spellings (see Yousif 2025; ‘Linguistic experts’ 2025) and that Canadian English spelling conforms neither to American English nor British English (e.g. Canadian tire centre vs. British tyre centre vs. American tire center). Such cases reveal the pressing need for a current resource that lays out and cements our own variety of spellings, pronunciations, usages, and unique lexical items, as distinct from the British and American varieties. Additionally, during this time of geopolitical uncertainty, where Canada’s identity and even sovereignty feel under threat, Canadians want and need a resource to call their own.
In 2022, Editors Canada recruited John Chew as editor-in-chief of the project. The core team quickly expanded, bringing in linguists, editors, lexicographers, fundraisers, and social media professionals, then set about documenting what would be involved in creating our new dictionary, its style guides, its procedures, its business model, its governance, and its public engagement. We chose as a starting point and proof-of-concept, fascicle Q. In Canadian English, Q is a significant letter because it demonstrates well our need for an up-to-date, inclusive dictionary: the definientia for queen bee, quantum computing, and quarantine, for example, have gone through major transformations since 2004 and need addressing; neologisms (e.g., quiet quitting, quadrex) need documenting; now offensive terms need labelling (e.g., quadroon, quarter-blood Indian); Canadian legal and military terms that carry some history from either American and British systems, and yet correspond fully with neither (e.g., Question Period, question of law, and quartermaster), need distinguishing; and the important work of decolonizing Indigenous language spellings and definientia (e.g., Q'omoχws, Queen Charlotte Island caribou) is urgent as Canada pursues reconciliation with Indigenous Peoples.
With these goals in mind, we discuss our establishment of lexicographical principles using several entries of interest from within the Q fascicle. We reveal how our editors construct all the aspects of an entry, with attention paid to issues of definition, spelling, pronunciation, etymology, labels, and presentation. We provide an overview of methodology, including the compilation of the word list, editorial procedures, drafting of editorial guides, and corpus construction.We also discuss our social media and outreach efforts, including the recent establishment of a Canadian Word of the Year, with which we hope to bring attention to Canadian English lexica amongst other major British and American dictionaries and their respective Word of the Year traditions. The very first Canadian Word of the Year, maplewash (the deceptive practice of making things seem more Canadian than they actually are) gained popularity in 2025 as Canadians embraced local products in the face of US tariffs. Its selection as Word of Year in a public poll indicates Canadians’ desires to establish our own ideological and linguistic identity on the global stage.
The initial fascicle Q, consisting of 500+ entries, is now available for online review, and work is continuing alphabetically, currently in fascicle R. If our ongoing funding efforts are successful, we plan to offer a free app by 2028, with premium and institutional subscriptions to ensure the sustainability of the CED. We look forward to introducing our project to international lexicographical experts and we welcome informed feedback on our methodologies and draft entries.
Speakers: John J. Chew III (Canadian English Dictionary), Emma Ferrett (Queen's University), Anastasia Riehl (Queen's University)
-
80
-
11:00
Coffee Break 00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna -
Keynote: Polysemy: Operationalization, Emergence, Perception, and the Human in the Loop? 01 | Festsaal (Festive Hall), OeAW Main Seat, 1st floor
01 | Festsaal (Festive Hall), OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Andreas Baumann (University of Vienna)-
83
Polysemy: Operationalization, Emergence, Perception, and the Human in the Loop
In lexica of natural languages, polysemy is the rule rather than the exception. Understanding why it is so prevalent—given that the presence of multiple senses for a single form, from a purely semiotic point of view, constitutes ambiguity—and how it emerges is central to lexicology. Addressing this question, of course, requires proper operationalization of polysemy in the first place. A robust, straight forward, and commonly used way of measuring the polysemy of a word is simply counting the number of sense entries in a lexicographically curated resource. Such an approach, however, neglects the relative frequencies of a word’s senses as well as how similar senses are to each other. By integrating lexicographic and corpus data, I will discuss frequency and similarity aware measures of polysemy, originating from quantitative ecology (Leinster & Cobbold, 2012), and how they relate to ratings of subjectively perceived polysemy, gathered through crowdsourcing efforts. I will then discuss the roles that frequency and acquisition play in the emergence of polysemy (Baumann & Hartmann, 2026). Finally, I will zoom in on one specific dimension of meaning: lexical sentiment. I show that variation in human sentiment annotations correlates well with how emotional polysemy is subjectively perceived, but that purely NLP based approaches in fact struggle with appropriately capturing subjective emotional ambiguity. I take this to stress the relevance of ‘the human in the loop’ also in lexical and lexicographic research.
Speaker: Andreas Baumann (University of Vienna)
-
83
-
12:30
Lunch Break 00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna -
Computational & AI-Based Lexicography 01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)
01 | Johannessaal
01 | Johannessaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Iztok Kosem (University of Ljubljana & Jožef Stefan Institute)-
84
From Content to Code and Back: Evaluating Agentic Coding in LLMs for Building R Shiny Dictionary Interfaces
While recent scholarship highlights the generative capabilities of LLMs for drafting lexicographic entries, their “black box” nature poses challenges regarding training corpora, linguistic representativity, and semantic hallu-cination. We propose that this opacity is less problematic when LLMs are deployed not as linguistic experts, but as software engineers. By leveraging LLMs as agentic coding tools, lexicographers can bridge the gap between linguistic data and digital publication, democratizing access to custom lexicographic tools. This study investigates how LLMs conceptualize “user-friendliness” when tasked with designing lexicographic interfaces. We evaluate three leading models in agentic coding – Claude Opus 4.6, ChatGPT-5.5, and Gemini 3.1 Pro – using a zero-shot prompting strategy within a case study on noun valency. Each model generates an R Shiny application from curated CSV data, accompanied by explicit justifications for its design choices. The generated interfaces are assessed against criteria fundamental to e-lexicography: access structures, visual design, information architec-ture, and alignment with user-centred principles for language learners. Results reveal substantial differences across models in how they conceptualize and operationalize usability, with each model emphasizing fundamen-tally different design aspects. Crucially, not all models prove equally reliable as agentic coders for generating lexicographic interfaces that meet professional standards, highlighting the need for critical evaluation before adoption.
Speakers: Iván Arias-Arias (Universidade de Santiago de Compostela), María José Domínguez Vázquez (Universidade de Santiago de Compostela), María Teresa Sanmarco Bande (Universidade de Santiago de Compostela), Carlos Valcárcel Riveiro (Universidade de Vigo) -
85
Linking Dictionaries Through Meaning: A Sense-Level Evaluation of Taxonomy-Guided LLM Classification in Lexicography
Despite the growing availability of digital dictionaries, semantic interoperability often remains limited to headwords, metadata, or article structures. Dictionary senses remain difficult to compare because they are shaped by different editorial traditions, segmentation practices, and levels of semantic granularity. This paper asks whether taxonomy-guided LLM classification can generate reliable, reviewable candidate links between heterogeneous dictionary senses. We test this approach on sense-level units from the letter range N in four German dialect dictionaries: Mecklenburgisches Wörterbuch, Hessen-Nassauisches Wörterbuch, Pfälzisches Wörterbuch, and Schweizerisches Idiotikon. Using GPT-5 via the KISSKI interface, each unit is assigned ten times to Post’s onomasiological taxonomy, which serves as a shared conceptual reference model. Evaluation combines exact agreement with reference labels, hierarchical proximity, run-to-run reproducibility, and majority aggregation. Exact correctness ranges from 74.91% to 80.88%; top-level taxonomic agreement exceeds 91%; and Fleiss’ κ values between 0.78 and 0.83 indicate substantial reproducibility. Majority aggregation raises exact correctness to 78.61–84.28%. The results show that taxonomy-guided LLM classification can support sense-level interoperability by producing interpretable candidate links and reliability signals for targeted lexicographic review.
Speakers: Hanna Fischer (Research Center Deutscher Sprachatlas), Alfred Lameli (Research Center Deutscher Sprachatlas), Nathalie Mederake (Research Center Deutscher Sprachatlas), Nico Urbach (Research Center Deutscher Sprachatlas) -
86
Historical Lexicography in the Age of AI: A Progress Report from the OED
The Oxford English Dictionary (OED) has long been recognized for its rigorous scholarship, yet it is equally notable for its readiness to adopt new technologies. As the volume and availability of linguistic evidence has grown in the digital era, so too have the challenges of efficiently revising a dictionary of this size and scope: in response, the OED has been undertaking a systematic investigation into the potential role of AI in supporting historical lexicography, aspects of which have been reported at previous conferences (Hawkes, Nicholson, and Rogers 2025; Wild 2024). In this paper we present findings to date on successful and less successful experiments, and describe the OED’s approach to scaling up the successes.
Much has been written about the capabilities of generative AI to draft dictionary definitions; in this paper, we report on two experiments that instead explore how LLMs can support the adaptation or updating of existing definitions for different purposes or audiences. The first experiment involved adapting definitions from the OED (a large historical dictionary of English) into a style suitable for the Oxford Dictionary of English (a general-purpose dictionary of current English). This requires removing obsolete or minor senses, making technical or formal language more accessible, simplifying long definitions, and similar transformations. We tested whether prompting an LLM to perform this task, followed by editorial review, could yield consistent and time-saving results; however, the output proved too variable to be implemented at scale.
The second experiment (initially reported in Wild 2024) focused on modernizing unrevised OED definitions, many of which were written in the 19th century, into a style appropriate for contemporary readers. With over half a million entries and more than 800,000 senses, the OED is one of the largest dictionaries in existence. This presents significant challenges in determining which entries and senses are most in need of updating. We first discuss methods of identifying outdated definitions, including traditional computational methods (such as finding definitions containing words which have declined in frequency in a corpus), and AI-supported methods (including diachronic embeddings and LLM verification). We then present the results of an experiment using an LLM to 'translate' archaic wording into modern English, when supplied with the unrevised OED definition, its associated illustrative quotations, and a prompt with examples. The results showed that AI-modernized definitions can serve as an aid for editors, along with other sources, although editorial input remains essential.
We have also investigated the capacity for AI to assist with retrieving and analyzing quotation evidence. OED's rich historical sense inventory is supported by typical, illustrative examples of a word or meaning, including the earliest use of a given sense and a recent example if the sense is still in current use. Research activity on the OED in this area involves laborious sifting of a large quantity of data from primary source databases, and the manual nature of this activity has limited the efficiency of how we work, and the volume of entries we can update. We evaluate the extent to which AI-assisted word sense disambiguation can support this task, firstly in a project using LLMs to retrieve quotations from a corpus of historical English and match them to OED senses to create a ‘Quotations Finder’ tool that aims to support the work of OED editors, and supplement OED quotation paragraphs for external users (initially reported in Hawkes, Nicholson & Rogers 2025). As part of this project, we explored the performance of 7 different LLM models and their comparative ability to carry out word sense disambiguation, and we will describe the relative merits of the model selected for the production of this tool. We also discuss a similar project for modern English which matches OED senses with examples from large corpora of 21st-century English. Our analysis shows that AI-assisted word sense disambiguation (in both historical and modern texts) has a relatively high accuracy rate for certain types of entry and sense, but that there are limitations and areas where editorial intervention is essential in selecting suitable quotations.
Taken together, these projects demonstrate both the promise and the limits of applying AI to historical lexicography. While current models cannot replace the nuanced judgement required for OED revision, they can meaningfully augment editorial work by accelerating manual tasks, and improve the discoverability of evidence. Our findings suggest that, when deployed selectively and with rigorous oversight, AI can help the OED scale its research and modernize its content, while preserving the scholarly standards that define the dictionary.
Speakers: Phoebe Nicholson (Oxford University Press), Will Rogers (Oxford University Press), Kate Wild (Oxford University Press)
-
84
-
Historical & Diachronic Lexicography 02 | Museumszimmer (02 | Museumszimmer, OeAW Main Seat, 2nd floor)
02 | Museumszimmer
02 | Museumszimmer, OeAW Main Seat, 2nd floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Carolina Flinz (University of Pavia)-
87
Vol. IX of the Etymological Dictionary of Old High German [Etymologisches Wörterbuch des Althochdeutschen (EWA)]: History Structure, Contents
The paper presents vol. IX of the Etymological Dictionary of Old High German [Etymologisches Wörterbuch des Althochdeutschen], the most encompassing etymological dictionary of any (Old) Germanic language. The volume will have roughly 800 pages, contain a total of 2,950 entries, of which 234 are main entries (the number is only about half the number of the main entries of vol. VIII) coming with a presentation of the full etymology and history of the word. Of these 234 main entries over 40 turn out to be loanwords, mostly attained from Medieval Latin (though originally sometimes from elsewhere). About 25 % of all entries (about 60 % of the main entries and almost 22 % of the fillers) are continued in modern Standard German or in German dialects. With these percentages the volume takes a middle position among the last volumes.
Speaker: Harald Bichlmeier (Saxon Academy of Sciences and Humanities in Leipzig) -
88
Legal Lexicon and the e-DRAE 1884: A Diachronic Study of Specialised Vocabulary in a General Dictionary Using a Digital Critical Edition
This paper presents a methodology for the diachronic study of legal lexical units related to criminal law in a nineteenth-century general-language dictionary, the twelfth edition of the Diccionario de la lengua castellana of the by the Real Academia Española (DRAE 1884), as analyzed through its critical and digital edition (e-DRAE 1884). The study combines a lexicographical examination of how these units are defined, structured and interconnected with their contextualization within the legal, social and ideological framework of nineteenth-century Spain, drawing on the sociocognitive theory of ideology and on the notion of the dictionary as discourse. Selected case studies from the criminal lexicon illustrate continuities and ruptures between the inherited juridical tradition and the emerging liberal order. The paper finally discusses the methodological advantages that the e-DRAE 1884 environment offers for research on specialized lexis in historical dictionaries.
Speaker: Marija Žarković Eriksson (Universitat Autònoma de Barcelona) -
89
From Ḋamma to Ġabra: Four Centuries of Maltese Lexicography
This paper presents a historical and typological overview of Maltese lexicography, tracing the development of lexical resources from the earliest manuscript dictionaries to recent digital initiatives. It surveys major multilingual, bilingual, monolingual and specialised dictionaries while discussing the methodologies underlying Maltese lexicography over the years. Emphasis is given to the interaction between the hybrid lexical structure of Maltese and lexicographic practice, including issues of lemmatisation, orthography and dictionary structure. The paper also discusses recent digital initiatives, particularly Ġabra and Dizzjunarju.mt, which illustrate the growing role of computational methods in Maltese lexicography.
Speakers: Michael Spagnol (University of Malta), John J. Camilleri (Chalmers University of Technology and the University of Gothenburg), Dwayne Ellul (University of Malta)
-
87
-
Lexical Semantics, Neology & Phraseology 01 | Sitzungssaal (01 | Sitzungssaal, OeAW Main Seat, 1st floor)
01 | Sitzungssaal
01 | Sitzungssaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Andrea Abel (Free University of Bolzano/Bozen & Eurac Research)-
90
Modelling Phraseological Semantic Structures in DMLex: Aligning a Croatian Lexonomy-Based Idiom Dictionary with a Thematic Index
Phraseological resources contain complex semantic structures that are often difficult to represent in interoperable lexical models. Although digital phraseological dictionaries increasingly use hyperlinks, thematic indices, and semantic groupings to improve navigation, many of these relations remain implicit and are primarily intended for human interpretation. This study explores the expressive capacity of the DMLex model for representing conceptual and semantic relations in a Croatian phraseological resource originally developed in Lexonomy. The resource consists of the Online Dictionary of Croatian Idioms and a manually curated thematic index grouping phraseological units into semantic domains and subdomains. The study adopts a scenario-based modelling approach focused on recurring phraseological structures observed in the resource. Three modelling scenarios are analysed: (1) variant and constructional relation networks, (2) polysemy and sense-level thematic classification, and (3) multiple thematic affiliation. The analysis demonstrates how lexical variation, constructional relations, synonymy, and thematic organisation can be represented through explicitly encoded semantic relations within DMLex. The results show that DMLex can accommodate much of the semantic complexity of phraseological resources while also revealing challenges related to overlapping classifications and conceptual hierarchies. The proposed approach contributes to the interoperable modelling of phraseological data for lexicographic and NLP-oriented applications.
Speakers: Ivana Filipović Petrović (Croatian Academy of Sciences and Arts), Slobodan Beliga (University of Rijeka, FIDIT) -
91
The Interplay Between Individual New Lexical Items and Multi-Layered Discourse Representations of Neologisms
This paper examines how discourse lexicography can be extended to the study of contemporary neologisms, particularly those emerging from major societal crises such as COVID-19, climate change, energy shortages, migration, inflation and the war in Ukraine. Traditional discourse lexicography has mainly focused on historically controversial terms and has typically relied on selected textual sources rather than large-scale corpora. In contrast, the paper argues for a corpus-assisted discourse-analytic approach that systematically captures the semantic, evaluative and ideological dimensions of newly coined and borrowed lexical items. Particular attention is paid to features such as semantic prosody, affective meaning, stance-taking, politicisation and power relations, which often become entrenched in neologisms and shape their use in public discourse. The paper introduces the German neologism resource IDS $Neo^{2020+}$, which combines two complementary levels of consultation: detailed lexical entries for individual neologisms and broader discourse or cluster entries that trace thematic developments and discourse networks over time. Both entry types are interconnected. This dual structure allows users to explore both the meanings of specific new words/expressions and their wider discursive functions and relations with other new terms within public communication. The paper also outlines the corpus-linguistic and computational methods used in the project, including word embeddings.
Speaker: Petra Storjohann (Leibniz Institute for the German Language) -
92
Trend Word Detection for DWDS, a Lexical Information System for Contemporary German
This paper addresses the challenge of prioritizing dictionary revisions in corpus-based lexicography, specifically for the Digital Dictionary of German (DWDS), which relies heavily on legacy content. Given the massive data volume in modern corpora, manual selection of relevant words for updating is no longer feasible. We investigate two methods for identifying “trend words”—terms showing an overproportionally high present-day frequency relative to past intervals. The first method applies established linear regression to large monitor corpora, highlighting limitations due to reliance on the chosen time period. To mitigate this, the second method employs a dispersion-oriented keyness metric on a specialized “discourse corpus” of current online text. This metric scores lemmas based on their spread across unique sources, effectively prioritizing words relevant to contemporary discourse. Evaluation of the ranked candidate list confirms that the high-score group contains a significantly higher density of relevant trend words compared to lower-ranked groups. This system allows lexicographers to systematically prioritize legacy entries for revision, such as those that are outdated or missing a current sense, thereby optimizing the lexicographical workflow and ensuring the dictionary reflects current language use.
Speakers: Alexander Geyken (Berlin-Brandenburg Academy of Sciences and Humanities), Gregor Middell (Berlin-Brandenburg Academy of Sciences and Humanities)
-
90
-
Morphology, Word Formation & Grammatical Lexicography 00 | Anton Zeilinger Salon (00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor)
00 | Anton Zeilinger Salon
00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Tomasz Michta (University of Bialystok)-
93
NomVallex: Valency of Czech Nouns and Adjectives in Networks of Derivationally Related Words
NomVallex is a valency lexicon of Czech nouns and adjectives, aimed primarily at describing the valency of individual noun and adjectival derivational categories in response to their insufficient theoretical description and coverage in lexical resources. Annotation has proceeded from the more extensively studied derivational categories (deverbal nouns and adjectives) through deadjectival nouns to the least researched denominal derivatives and compounds. Since valency is often affected by word-formation relations, NomVallex has implemented features for capturing derivational relations between derivationally related lexical units, and for comparing their valency. Although the annotation in NomVallex often encompasses more than one derivative of a particular base word, only a limited portion of the word-formation relations in Czech is covered there. The lexicon therefore newly provides links to DeriNet, where users can view a given word in the broader context of its derivational relations. Moreover, to display derivational relations and changes in valency associated with the word-formation processes in whole networks of derivationally related words recorded in NomVallex, version 2.7 has introduced a DeriNet-like graphical representation of derivational relations, which is available in a user-friendly format on the lexicon’s website, facilitating further linguistic research into the influence of word formation on the valency of derived words.
Speakers: Veronika Kolářová (Charles University, UFAL MFF), Václava Kettnerová (Charles University, UFAL MFF), Michal Olbrich (Charles University, UFAL MFF), Jiří Mírovský (Charles University, UFAL MFF) -
94
Word-Formation Chains in Slovenian: Evaluating the Effectiveness of AI
This study investigates word-formation chains, i.e. sequences of derivationally related words in which each derivative serves as the base for the next. Due to their rich morphemic structure, such chains are particularly complex in Slavic languages. Their analysis is relevant for understanding word-formation and semantic potential, morphotactics, and lexicographic criteria for vocabulary inclusion. The research builds on the project Formant Combinatorics in Slovenian (2022–2025), which applied automatic extraction methods to data from the Word-Family Dictionary of the Slovene Language (BSSJ). Abstract suffix-chain patterns derived from this resource were extended to larger lexical datasets. Preliminary findings showed limited success of automatic extraction methods (24% for the most frequent suffix chain), whereas AI achieved higher accuracy (59%). However, performance declined markedly with increasing chain complexity. The present study aims to evaluate the effectiveness of ChatGPT 5.5 in identifying word-formation chains with nominal, verbal and adjectival bases. It focuses on the impact of chain complexity, the role of part-of-speech within chains, and problematic suffixes and zero-formant derivation patterns. Particular challenges include derivatives without overt suffixes and cases where zero-formant nouns are misclassified as unmotivated words. The analysis examines suffix chains of varying complexity and compare AI performance with previous automatic extraction approaches, with the goal of improving both extraction methods and AI prompting strategies.
Speaker: Boris Kern (ZRC SAZU & University of Nova Gorica) -
95
A Graph-Based Lexicographic Workflow for Japanese Adjectives and Their Lexical Relations in an AI Tutor
This study presents a graph-based workflow for extending the lexicographic description of Japanese adjectives with large language model (LLM) relation candidates and expert review. JMdict provides the principal dictionary evidence. Lexical-item nodes carry written form, reading, grammatical category, English glosses, learner information, and provenance; relations record a proposed type together with semantic, register, domain, scalar, provenance, and review information. Written form plus hiragana reading serves as an operational lookup key that distinguishes homographs such as 辛い (karai ‘spicy’) and 辛い (tsurai ‘painful’). The pilot stored 1,960 proposal observations from gpt-4.1-mini and 2,890 from gpt-5.4-mini across 400 source records representing 305 distinct items; the latter total includes a 101-observation preliminary run. An exploratory run added 76 observations. Three linguists supplied 420 judgments on 214 proposals: 314 judgments combined acceptance with an overall score of at least four, and 180 proposals received at least one such judgment. The results describe the reviewed subset and establish the feasibility of preserving dictionary evidence, model observations, target-resolution outcomes, and expert judgments as distinct components of a lexicographically interpretable resource.
Speakers: Benedikt Perak (University of Rijeka), Dragana Špica (Juraj Dobrila University of Pula), Irena Srdanović (Juraj Dobrila University of Pula)
-
93
-
15:30
Coffee Break 00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna -
Computational & AI-Based Lexicography 01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)
01 | Johannessaal
01 | Johannessaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Kristina Š. Despot (Institute for the Croatian Language)-
96
Can Lexicographic Entries Embed Didactic Modules? A Case Study on LLM-Generated Outputs for a Dictionary Interface
Digital dictionaries can become more than consultation tools when their internal data are sufficiently structured to be reused and transformed into embedded learning components. This paper evaluates whether large language models (LLMs) can convert PORTLEX entries, represented as reviewed JSON files, into interactive HTML didactic modules integrable into their corresponding dictionary entries. PORTLEX is a multilingual valency-oriented dictionary of the nominal phrase, and its microstructure offers a controlled source for activities on formal links, ordering, complement function and semantic compatibility. The experiment compares a GPT-based custom system and a Gemini-based Gem across three languages, four semantically aligned nouns per language and two model outputs per entry, producing 24 modules and 48 expert evaluations. Results show that both systems generated broadly usable modules and generally preserved the linguistic-valency information encoded in the JSON input. The main limitations did not concern basic HTML functionality, but pedagogical regulation: feedback, correction, scaffolding, distractor quality and ambiguity management. Overall scores did not differ significantly; only feedback and correction favoured GPT after statistical adjustment. Language-based patterns were diagnostic rather than conclusive. The results support treating LLM-generated modules as controlled prototypes within a documented, reusable workflow, not as ready-made pedagogical products.
Speakers: Carlos Valcárcel Riveiro (Universidade de Vigo), María Teresa Sanmarco Bande (Universidade de Santiago de Compostela), Iván Arias-Arias (Universidade de Santiago de Compostela), María José Domínguez Vázquez (Universidade de Santiago de Compostela) -
97
Detecting Semantic Variation between German Standard Varieties through Word Embedding Comparison
From a pluricentric perspective, German possesses several standard varieties, including the one used in South Tyrol (Italy). While automated lexicography has long focussed on differences of lexical forms across these varieties, meaning differences in formally identical words are more difficult to identify systematically. This study addresses such diatopic semantic variation by comparing usage patterns in corpora from Germany and South Tyrol. Using distributional evidence from word embeddings, it shows how corpus-based comparison can support lexicographic analysis by highlighting shared words whose meanings or typical contexts diverge across varieties. The focus lies on the linguistic interpretation of these patterns and on their relevance for describing variety-specific meanings. The results confirm existing lexicographic descriptions while also identifying additional candidates for future verification.
Speakers: Andrea Abel (Free University of Bolzano/Bozen & Eurac Research), Luca Ducceschi (Eurac Research), Federico Villa (Independent Researcher) -
98
The Ninth Art in Lexicography: The Power of Comic Strips
This paper presents the background to a book on the visual history of academic lexicography, being a sequence of hundreds of one-page comic strips, created through conversations with LLM chatbots and GenAI image generators, each page representing a major publication in the field. Of course, I knew that this submission would be a gamble: Would a comic version of a conference paper be taken seriously? Well, one extremely positive adjudicator gave it a ‘Strong Accept’, and a more critical one gave it an ‘Accept’. Adjudicator 1 summarised it head-on, writing that it is “an extremely original contribution to the field, as it addresses the impact and outreach of scientific research both within and beyond academia”, while Adjudicator 2 pointed out that it is “untraditional but entertaining, and comic strips could very well be part of lexicography”. In what follows, my original extended abstract is presented first, after which I interact with the feedback left by Adjudicators 1 and 2. No doubt, and again, a very uncommon approach, but times have changed in the era of generative AI.
Speaker: Gilles-Maurice de Schryver (Ghent University)
-
96
-
Corpus Linguistics & Corpus-Based Resources 02 | Museumszimmer (02 | Museumszimmer, OeAW Main Seat, 2nd floor)
02 | Museumszimmer
02 | Museumszimmer, OeAW Main Seat, 2nd floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Dominika Kováříková (Charles University)-
99
Camfranglais in the Age of Artificial Intelligence: Lexicographic Challenges, Opportunities and a Human-AI Approach to Documenting Linguistic Hybridity in Multilingual Cameroon
In postcolonial societies, language often functions as both a site of power and resistance, a reality profoundly exemplified by Camfranglais in multilingual Cameroon. As Foucault (1978, pp. 95–96) argues, power inevitably generates resistance, and in Cameroon this resistance manifests linguistically through a hybrid sociolect that disrupts inherited colonial hierarchies. Emerging in the 1970s within a “communicative vacuum” (Kießling, 2015, p. 9) created by exoglossic language policies, Camfranglais combines French, English, Cameroonian Pidgin English, and numerous indigenous languages. Despite its widespread use and cultural significance, it remains substantially under-documented. Existing lexicographic resources resemble glossaries or linguistic guides constrained by static print formats, limited corpora, and the absence of systematic phonetic transcription. They also reveal a metalanguage mismatch: Kamdem Fonkoua (2015) provides English definitions with untranslated Camfranglais examples, creating barriers for Francophone users. Essome Bouti (2024) and Ndongo (2015) offer French metalanguage, which serves Francophone speakers but excludes Anglophone Cameroonians. Neither orientation serves speakers from the less-educated variety identified by Ebongue and Fonkoua (2010), who may have acquired French or English informally rather than through formal schooling. Consequently, a significant gap persists between the sociolinguistic reality of Camfranglais and its lexicographic representation.
The rise of artificial intelligence opens new possibilities for low-resource language documentation through large-scale data processing, pattern recognition, and phonetic prediction (Zhong et al., 2024; Wang, 2024; Midigo 2025). However, these opportunities are constrained by structural biases in current NLP systems, which are predominantly trained on high-resource and typologically stable languages. Ebrahimi et al. (2022, p. 6279) show that multilingual models achieve only 38.48% average zero-shot accuracy on genuinely low-resource languages. To address this tension, this paper proposes a Human–AI collaborative framework for documenting Camfranglais, grounded in the Function Theory of Lexicography (Bergenholtz & Tarp, 2002, 2003; Tarp, 2008a, 2008b). This framework informs the design of a bilingual, bidirectional digital dictionary with dual metalanguage access (French–Camfranglais and English–Camfranglais), intended to serve both Francophone and Anglophone users. Critical Discourse Analysis (CDA) is also applied to capture how Camfranglais expresses identity, ideology, and power relations in context (Fairclough, 1989; Blommaert & Bulcaen, 2000; Hidalgo Tenorio, 2011; Discourse Analyzer, 2024).
The study is further informed by an exploratory survey of forty Cameroonian speakers from diverse sociolinguistic backgrounds. Results confirm widespread informal use of Camfranglais and strong support for its documentation as a marker of national identity, while also revealing concerns that excessive standardisation could undermine its creativity and fluidity. In response, AI is positioned as a “copilot” subordinated to human expertise, ensuring linguistic sovereignty and aligning with principles of anti-extractivism (Riofrancos, 2020).
Methodologically, the study relies on a deliberately constructed multi-source corpus capturing both historical written forms and contemporary oral-digital manifestations of Camfranglais. The written component includes the Grioo Forum (2004–2006), documenting early diasporic digital interaction, and the Bonaberi Forum (2008–2026), reflecting spontaneous written exchanges on everyday topics. The oral-digital component is drawn from the YouTube channel Warman du Terre à Terre (2013–2026). From this source, 1,583 video URLs were extracted to map lexical diversity, and 100 videos were fully transcribed to capture spontaneous speech, prosody, and emerging vocabulary. Together, these sources provide a heterogeneous corpus documenting the evolution of Camfranglais from street speech to digital discourse.
The analytical workflow operationalises a three-stage Human–AI feedback loop. First, computational tools identify candidate entries, frequency patterns, and collocations. Python was used for statistical extraction, while Sketch Engine (Kilgarriff et al., 2014) enabled deeper contextual analysis through concordances and Key Word In Context (KWIC) functions, particularly useful for identifying hybrid forms often missed by automated methods. Second, the Large Language Model Claude generated contextual definitions and etymological notes, while ChatGPT produced candidate phonetic transcriptions. Only standard French items were excluded; Cameroonian French forms with semantic or phonological deviation, as well as English items phonologically adapted by Camfranglophones, were retained as authentic lemmas. This procedure yielded approximately 2,850 lexical items, including multi-word expressions.
Finally, all outputs were validated by fourteen native Camfranglais speakers through independent five-point Likert-scale evaluations followed by consensus sessions. Results highlight both the strengths and limitations of AI-assisted lexicography. While AI accelerates extraction and contextual analysis, it struggles with tonal minimal pairs, pragmatic inversion, neologisms, hybrid morphology, and distinctions between Cameroonian and standard French. Nevertheless, validation achieved substantial inter-rater agreement (Fleiss’ kappa = 0.74) (Landis & Koch, 1977), with high phonetic concordance: 94.2% perceptual agreement; 96.8% acoustic agreement using Praat (Boersma & Weenink, 2026).
This paper argues that a principled Human–AI approach is not merely a technical convenience but an epistemic necessity for documenting Camfranglais. By combining computational efficiency with native-speaker expertise, the study produces a pilot electronic dictionary that is phonologically informed, sociolinguistically grounded, and adaptable to ongoing lexical change. More broadly, it offers a transferable model for documenting hybrid urban varieties and demonstrates how artificial intelligence can support marginalized speech communities without reproducing extractive dynamics.
Speaker: Emmanuel Sylvain Fomat (University of Hildesheim) -
100
Towards a French–Serbian Dictionary of Sports Terms: Corpus-Based or AI-Assisted Methodology?
Earlier multilingual dictionaries of sports terminology that include French and Serbian were largely based on simple translations from foreign sources, without systematic terminological analysis whereas more recent works do not include French. To be able to address this gap and compile a reference dictionary, we first need to identify the optimal methodology for the compilation of this type of dictionary entries. This paper presents the evaluation of two different possible approaches — corpus-based and AI-assisted generation. First, we selected ten polysemous French sports terms with multiple Serbian equivalents across different contexts and generated dictionary entries using both corpus analysis and AI models. The corpus used in the process is a multilingual parallel corpus ParCoLab containing aligned English, French, and Serbian texts from official sports rulebooks of 14 sports, providing a solid basis for terminological extraction, whereas the AI models evaluated were GPT-5.5, Sonnet 4.6 and Gemini 3.6. The generated entries were then evaluated by human experts for accuracy, consistency, and practical usefulness. The results show that while AI significantly accelerates the compilation process, corpus-based methods remain more reliable. Combining both approaches, thus, proves to be most effective: corpus ensures data quality, while AI supports efficiency. However, final validation by a human terminologist remains essential for refining meanings, standardizing terminology, and ensuring long-term usability.
Speakers: Dušica Terzić (University of Belgrade - Faculty of Philology), Dejana Mirković-Birtašić (University of Belgrade - Faculty of Philology), Saša Marjanović (University of Belgrade - Faculty of Philology) -
101
Towards a Repository of Similes in Contemporary Serbian
This paper presents the methodological framework for building the Similes Repository for Contemporary Serbian (Simili-SR), a corpus-based linguistic resource designed for the collection, annotation, and analysis of similes in contemporary Serbian. The study presents a multi-stage extraction and validation pipeline based on the genre-balanced corpus SrpKor2021+ of the Reference Corpus of Contemporary Serbian, combining CQL queries, local grammars implemented in UNITEX, and LLM-assisted semantic annotation. The initial extraction phase produced 574,195 candidate sentences across six major simile patterns, confirming both the productivity and structural diversity of simile constructions in Serbian. Subsequent rule-based filtering and LLM-assisted validation substantially reduced noise while preserving a large pool of linguistically relevant figurative comparisons. Manual evaluation performed on a stratified sample additionally revealed that simile interpretation frequently involves distributed semantic structure, and complex interactions between grammatical realization and figurative meaning, and highlighted the importance of linguistic post-processing, as well as several challenges related to semantic annotation in morphologically rich languages such as Serbian. By combining rule-based methods, LLMs, and detailed linguistic analysis, the paper contributes to both the development of computational resources for Serbian and broader research on figurative language annotation and representation.
Speakers: Aleksandra Marković (Institute for the Serbian Language), Cvetana Krstev (JeRTeh Society), Milica Dinić Marinković (University of Belgrade - Faculty of Philology), Ranka Stanković (University of Belgrade, Faculty of Mining and Geology)
-
99
-
Multilingualism, Language Varieties & Contact Lexicography 01 | Sitzungssaal (01 | Sitzungssaal, OeAW Main Seat, 1st floor)
01 | Sitzungssaal
01 | Sitzungssaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Veronique De Tier (Dutch Language Institute)-
102
Beyond the Single Standard: Addressing Variation and User Needs in an Online Frisian Dictionary
This paper examines the challenges of representing linguistic variation in Frisian lexicography, focusing on the Dutch-Frisian online dictionary. While existing lexicographic resources already accommodate variation to some extent, discrepancies remain between dictionary representation and actual language use. Survey data and usage examples reveal complex patterns in which standard forms coexist with contact-induced variants, and in some cases no fully equivalent Frisian alternatives exist, leading speakers to adopt Dutch forms in Frisian discourse. These findings highlight a tension between normative guidance and descriptive adequacy, particularly in contexts where speakers vary in competence and where written Frisian differs from spoken usage. The current label system, although informative, does not always provide sufficiently systematic or user-oriented support. Building on insights from user-oriented lexicographic theory, this paper proposes a more flexible model in which dictionary information is accessed through multiple modes. These include a normative mode offering clear guidance, a comparative mode making variation explicit, and a learning mode tailored to users with less experience. By structuring variation in a transparent and user-sensitive way, the proposed approach aims to align lexicographic practice more closely with actual language use while supporting both standardization and linguistic diversity.
Speakers: Nika Stefan (Fryske Akademy), Johan van der Zwaag (Fryske Akademy), Hindrik Sijens (Fryske Akademy) -
103
Synonyms That Never Met: On the Compilation of the Dictionary of the Dubrovnik Idiom
This article outlines the Dictionary of the Dubrovnik Idiom project at the Institute for the Croatian Language. Given its unconventional conceptual framework, this dictionary adopts a diachronic approach, incorporating headwords from diverse historical periods (spanning from the 16th to the 21st century) and across a broad spectrum of linguistic registers: from the vernacular to high-prestige literature. From the perspective of a fixed point on the diachronic axis (e.g. present-day usage), some synonyms may, depending on their temporal distance, be classified as obsolete or archaic. However, when arranged along both the synchronic and diachronic dimensions, synonym sets can be further differentiated according to functional-stylistic and register-based distribution. The paper presents a typology of the synonymy encountered during the lexicographic processing. Particular attention is paid to cases in which the dictionary records synonyms that never met, that is, synonyms for which corpus evidence shows no temporal overlap.
Speakers: Ivana Lovrić Jović (Institute for the Croatian Language), Lana Hudeček (Institute for the Craoatian Language) -
104
From Plurilingual Lexicography and Terminography to Generative AI Translation: Potential, Limits, and the Enduring Role of Dictionaries
Recent advances in generative artificial intelligence (AI) have profoundly reshaped the landscape of translation. Large Language Models (LLMs) such OpenAI’s GPT series are now capable of producing fluent, coherent and context-aware translations that go beyond the word-for-word logic traditionally associated with earlier forms of Machine Translation (MT) systems (Kenny, 2022). Despite their apparent accuracy, however, these tools demonstrate uneven translation performance with texts belonging to different textual genres and degrees of specialization.
This paper investigates the potential and limits of current leading MT paradigms through the qualitative analysis of translations of texts situated along the continuum between general and specialized language (Cabré, 2007), with particular attention to semi-specialized discourse. The study draws on empirical data from a comparable corpus of travel literature and compares translations of French texts related to artistic and cultural heritage into English and Italian, produced by two Neural Machine Translation (NMT) systems (DeepL and Google Translate) and three generative AI models (ChatGPT-5.5, Claude Sonnet 4.6 and Mistral Medium 3.5).
The findings show that, while generative AI models often outperform traditional NMT systems in terms of fluency and overall textual coherence (Jiao et al., 2023; Hendy et al., 2023), they continue to display significant weaknesses in areas that are crucial for specialized and semi-specialized translation. These include inadequate handling of multi-word terms (e.g. tableau d’autel) and specialized lexical combinations (L’Homme, 1998, 2017) (e.g. colonne ronde), difficulties in semantic disambiguation (e.g. tombeau), and instances of so-called “hallucinations”. Such limitations can be attributed to the probabilistic, data-driven nature of LLMs, which lack genuine semantic understanding and are trained predominantly on large but largely generic and unevenly distributed corpora (Bender et al., 2021; Ciotti, 2023; Raus, 2024; Tekwa, 2025).
Against this backdrop, the paper revisits a question often considered outdated in the “Machine Translation Era” (Asscher, 2025, p. 4): are traditional dictionaries still worth using? The analysis strongly supports an affirmative answer. General bilingual and monolingual dictionaries remain essential for managing polysemy, stabilizing semantic correspondences, and preventing uncontrolled lexical variation – tasks that generative models perform implicitly and inconsistently (Mattioda, 2024). Even more crucial is the role of specialized lexicographic and terminological resources, which provide explicit conceptual structuring, domain-specific distinctions, and historically grounded knowledge that AI systems are currently unable to guarantee.
Rather than being rendered obsolete by generative AI, monolingual and plurilingual dictionaries emerge as complementary resources supporting human pre- and post-editing processes, guiding prompt design, and constituting high-quality supervised datasets for domain adaptation in neural and generative systems (Chu & Wang, 2018; Hansen, 2024). In this perspective, translation is increasingly configured as a hybrid activity in which human intelligence, lexicographic expertise, and AI interact complementarily rather than competitively.
The paper concludes that the future of translation does not lie in the replacement of dictionaries, but in their renewed centrality as linguistic, cognitive, and epistemological anchors within an evolving ecosystem of generative AI tools.
Speakers: Valeria Zotti (University of Bologna), Rita Gramellini (University of Bologna)
-
102
-
Specialised & Terminological Lexicography 00 | Anton Zeilinger Salon (00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor)
00 | Anton Zeilinger Salon
00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Rufus H. Gouws (Stellenbosch University)-
105
Lexical and Conceptual Tools for Learning Mathematics
Studies with first-semester mathematics students indicate that after leaving highschool, they understand the subject primarily as a set of calculation rules, i.e. as a ‘toolbox’ (Törner & Grigutsch, 1994; Girnat & Hascher, 2021). However, mathematics as taught at university, as well as any scientific approach to mathematics, should rather be understood as an interconnected system of concepts expressed linguistically in a rich terminology. Learning mathematics therefore involves learning mathematical language and understanding how mathematical concepts are interrelated (cf. Riccomini et al., 2015).
Students are thus in need of lexical/terminological tools that support them in two ways: (a) to understand basic mathematical concepts and their relations (cognitive needs, in terms of the lexicographic function theory, cf. Tarp, 2008; Fuertes-Olivera & Tarp, 2014); and (b) to get acquainted with the phraseology needed for understanding and producing texts about mathematical topics (communicative needs: receptive and productive).
We analysed a number of sample entries from sources students may typically resort to: Wikipedia or other (online) encyclopedias, AI generated summaries of answers to keyword-based questions as well as (online) learning material for first years. We checked, a.o., the German terms 'Gleichung' (equation), 'Äquivalenz' (equivalence), 'Lösung' (solution) and 'Lösungsmenge' (solution set), all of which appear early in the first semester introductory lecture.
Wikipedia and similar tools provide substantial conceptual knowledge, but they require a considerable effort to extract relational concept data from the rather complex article texts (cf. Kruse, 2025, pp. 81-83). The analysed online material either contains wikipedia-like mini-essays or it focuses on computing examples. None of these make phraseology explicit.
We assume that formulating meaning explanations and conceptual relations explicitly may help students to master (and to memorize) mathematical concepts (a sort of 'learning by doing'). In an experimental course (cf. Kruse et al., 2024) administered since 2023, we thus invited participants of the first year introductory lectures to mathematics to restructure elements of the lecture contents by designing their own dictionary entries for the mathematical terms learned in the lecture; we also asked them to produce small ontological networks. To ensure that they would adhere to a general guideline, we gave a lexicographical mini-introduction to the microstructural items we are interested in, e.g. definitions, possibly paraphrases of definitions for school children, synonymy, hyp(er)onymy, related concepts that are part of the subdomain in question as well as collocations. We equally introduced the students to concept maps and ways to represent relations.
An analysis of a subset of the students' submissions (96 submissions) shows that their dictionary entries contain more types of microstructural items than their networks. All dictionary entries and almost 80 % of the networks contain definitions. Synonyms and Hyponyms are present in ca. 80 % of both types of submissions, while collocations only show up in dictionary entries and mathematical notations (symbols etc.) almost exclusively in the network graphs.
Overall, the students are quite successful in the exercise: they provide more conceptual relations than most online learning material. It seems, however, that they should need to be made more aware of mathematical phraseology. A combination of particularized information in the style of dictionary entries and of small network-like representations may be a good way to provide them with their own learning material. In the medium term, a study of the effectiveness of the approach would be needed.
Speakers: Barbara Schmidt-Thieme (University of Hildesheim), Theresa Kruse (University of Hildesheim), Ulrich Heid (University of Hildesheim) -
106
Defining Legal Concepts in the AI Age: An Exploratory Study into Influencer Practices
Definitions play a crucial role in the acquisition, transfer, or translation of specialized knowledge. In the legal profession, precise definitions of complex concepts facilitate seamless knowledge management, application, and interpretation. However, despite growing reliance on artificial intelligence (AI) for this task, the real-world feasibility and user acceptability of AI-generated legal drafting remain largely unexamined. When addressing this gap, it is critical to consider that the use of authoritative legal terminology is central to the validity of any legal definition. Indeed, precise terms facilitate the application and interpretation of core concepts. However, new legal domains lack stable legal terms and conceptually demarcated concepts. Resolving the inherent tension between linguistic evolution and the need for legal certainty make the legal domain a unique use case that necessitates domain-specific task adaptation for general-purpose large language models (LLMs). With that in mind, this paper explores the potential of LLMs to generate legal definitions perceived as clear, precise, and usable by domain users through an exploratory study of a generated dataset of key influencer-related concepts. The results underscore the critical importance of data quality and precise prompt engineering when leveraging LLMs for complex, domain-specific tasks.
Speakers: Martina Bajčić (University of Rijeka, Faculty of Law), Ana Ostroški Anić (Institute for the Croatian Language) -
107
Lexicography in the Age of AI: Training Systems with Human Expertise in Art Terminology
The emergence of artificial intelligence represents the latest in a series of technological challenges that have periodically reshaped the relationship between human expertise and technological tools. This paper takes as its point of departure the conviction that AI cannot exist, function, or evolve independently of human intelligence; lexicography, as a discipline situated at the intersection of linguistic knowledge, cultural competence, and technological practice, is particularly well placed to explore how human expertise can guide AI systems rather than be displaced by them. Drawing on the resources of the Lessico dei Beni Culturali (LBC) project, a multilingual lexicographic initiative at the University of Florence covering Italian cultural heritage terminology across Italian and seven further languages, the paper investigates how human-curated reference materials and expert-designed prompts can systematically improve AI performance on culturally embedded art terms. Three representative lemmas (cartone, grazia, and maniera) were tested across three conditions, namely baseline, expert-prompted, and LBC-constrained, using three major AI systems (Claude Sonnet 4.6, Gemini 3.5 Flash, and GPT-5). The results confirm a consistent gradient: corpus evidence not only improves output quality but, in some cases, reveals lexicographic realities that expert prompting actively conceals, most strikingly the absence of a stable specialised equivalent for maniera in Spanish, which the LBC corpus exposes as a structural lexical gap rather than a straightforward translational choice. The study argues that the value of expert lexicographic knowledge does not diminish in AI-assisted workflows but becomes more concentrated, shifting from the direct production of entries to the design of the knowledge infrastructure within which AI systems can operate effectively.
Speaker: Riccardo Billero (Free University of Bozen)
-
105
-
108
General Assembly 01 | Festsaal (Festive Hall), OeAW Main Seat, 1st floor
01 | Festsaal (Festive Hall), OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaFor EURALEX members only
-
-
-
Computational & AI-Based Lexicography 01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)
01 | Johannessaal
01 | Johannessaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Ana Ostroški Anić (Institute for the Croatian Language)-
109
Chatbots in Pedagogical Lexicography: Evaluating AI-Generated Dictionary Entries from Three LLMs
As early as the eighteenth century, dictionary compilation has been perceived as a labor intensive and time-consuming endeavor, often characterized as meticulous and repetitive work (Lew, 2023). The creation of lexicographic articles, in particular, constitutes a complex process requiring the careful coordination of multiple parameters, including the target user group, users’ native language, the type and purpose of the dictionary, as well as the design of its macrostructure and the organization of its microstructure (Hartmann, 2001). In pedagogical lexicography, this complexity is further intensified, as definitions and supporting information must also account for learners’ age, language proficiency, learning objectives, and the conditions under which the dictionary is consulted, such as reception- or production-oriented use (Tarp, 2011). Against this background, contemporary lexicography has increasingly emphasized the need to reduce the time and cost of dictionary making, with automation widely regarded as a key factor in achieving this goal (Jakubíček & Rundell, 2023; Rundell, 2024). Recent advances in artificial intelligence, particularly over the past five years, have therefore prompted lexicographers to explore new solutions for both dictionary compilation and use. Notably, the emergence of Large Language Models (LLMs) has intensified interest in the potential role of chatbots in supporting—and possibly reshaping—various stages of the lexicographic process, including the generation and mediation of lexicographic articles (Phoodai & Rikk, 2023; Rees & Lew, 2023; Lew et al., 2024; Li & Tarp, 2024; Ptasznik et al., 2024; Ptasznik & Lew, 2025).
Against this backdrop, the present study investigates how three widely used LLMs generate dictionary entries and examines their potential contributions and limitations within pedagogical lexicography. Focusing on AI-generated dictionary entries produced by ChatGPT, Gemini, and DeepSeek, the study adopts a comparative, criteria-based evaluation framework grounded in learner lexicography, pedagogical principles, and AI-assisted language learning. Its primary objective is to assess the extent to which LLMs can simulate the writing practices and decision-making processes of human lexicographers. To this end, the models’ ability to generate definitions, examples, and idiomatic expressions is evaluated using twenty carefully selected entries from Helix, a bilingual illustrated lexicon designed for Greek heritage learners (Chadjipapa & Gavriilidou, 2024). All models were prompted under controlled conditions to produce dictionary entries including definitions, sense distinctions, usage notes, examples, collocations, register labels, and, where relevant, learner-oriented or cross-linguistic explanations.
The chatbot-generated lexicographic articles were evaluated by four expert lexicographers specializing in pedagogical lexicography. The evaluation combined qualitative and quantitative dimensions and was conducted under blind conditions, with assessors unaware of whether entries were produced by human lexicographers or AI systems. Entries were assessed along five main axes: (a) definitional accuracy and semantic precision, (b) internal consistency and sense discrimination, (c) pedagogical usefulness for learners, (d) transparency and explicability of lexical information, and (e) alignment with established learner lexicographic principles, including clarity, economy, controlled defining vocabulary, and relevance to communicative use. Particular attention was paid to whether the models reproduced established lexicographic conventions, extended them, or violated core norms. The results show that all three LLMs are capable of producing superficially well-formed dictionary entries that resemble the structure and tone of learner dictionaries. Nevertheless, systematic differences were observed. ChatGPT demonstrated stronger coherence in extended explanations and pedagogical scaffolding, particularly in usage notes and learner oriented paraphrasing. Gemini excelled in the range of examples and contextual variation, though sometimes at the expense of semantic precision, while DeepSeek offered comparatively limited learner-specific guidance. These findings align with patterns reported in previous research (Lew, 2024). Across models, recurrent weaknesses included inconsistent sense hierarchies, blurred distinctions between core meaning and pragmatic inference, occasional hallucinated usage restrictions, and limited sensitivity to frequency, register, and learner difficulty.
Despite these limitations, the findings point to a significant pedagogical opportunity: when guided by appropriate pedagogical prompts and lexicographic constraints, LLMs can meaningfully support aspects of dictionary compilation. The contribution of this paper is therefore threefold. First, it provides one of the first systematic and comparative evaluations of AI-generated lexicographic articles produced by multiple LLMs, employing a blind assessment procedure that enables a more nuanced evaluation of whether chatbot generated outputs meet established lexicographic standards beyond surface-level fluency. Second, it proposes an evaluative framework grounded in pedagogical lexicography, offering concrete criteria for assessing and improving chatbot-mediated dictionary entries, particularly with respect to learner-oriented definitions, sense structuring, and usage information. Third, the paper advances the methodological integration of chatbots into lexicographic workflows by examining prompt design as a key mediating factor in AI-assisted dictionary making. Focusing on few-shot prompting (Brown et al., 2020), it analyes how chatbots adapt their outputs in response to carefully curated lexicographic examples and identifies parameters necessary for the effective, reliable, and pedagogically sound use of prompts. Taken together, these contributions frame chatbots as potential partners in pedagogical lexicography, capable of extending traditional dictionary functions through interaction, personalization, and AI supported mediation when guided by lexicographic expertise.
Speakers: Elina Chadjipapa (Democritus University of Thrace), Zoe Gavriilidou (Democritus University of Thrace), Maria Mitsiaki (Democritus University of Thrace), Ifigeneia Dosi (Democritus University of Thrace) -
110
The Reversal of a Bilingual Dictionary Using AI
This paper presents a pilot study exploring the use of artificial intelligence in reversing a bilingual dictionary. The Modern Icelandic-English Dictionary is a freely available online resource. The aim was to transform this Icelandic-English dictionary into a functional English-Icelandic dictionary while minimising manual labour. Using the DANTE English lexical database as a reference, the study focused on English words beginning with the letter m. The process was divided into five phases. First, SQL queries were used to generate an initial reverse listing. Second, the Gemini large language model was employed to process the output and combine it with data from the DANTE list. The resulting dataset contained 4,750 potential entries, of which 1,750 lacked an Icelandic equivalent. In the third phase, Gemini supplied missing translations, while in the fourth phase it performed sense disambiguation for polysemous entries. The fifth phase involved post-processing, including the elimination of unnecessary entries and extensive proofreading. The findings suggest that AI can successfully handle three of the five phases. However, substantial post-processing still requires human expertise, although AI might also be used to facilitate parts of that work.
Speakers: Thordis Ulfarsdottir (The Árni Magnússon Institute for Icelandic Studies), Ellert Thor Johannsson (The Árni Magnússon Institute for Icelandic Studies) -
111
Effectiveness of AI for the Lexicography of Low-Resource Languages: The Dictionary of Georgian Neologisms
The purpose of this paper is to present the results of a study investigating the effectiveness of AI in the compilation of the Dictionary of Georgian Neologisms, an ongoing project at the Centre for Lexicography and Language Technologies at Ilia State University. The data for the study were selected from a dataset of 1,700 neologisms identified in previous research that developed a semi-automatic method for detecting neologisms in Modern Georgian (Laluashvili & Margalitadze, 2025). Three major categories of Georgian neologisms were selected: (1) borrowings, including direct loans, morphologically adapted loanwords, and loan translations; (2) word formation, with a special focus on internally created Georgian formations; and (3) semantic evolution. The experiments conducted in this study employed ChatGPT 5.2, Gemini 3.0 Thinking, and Grok 4.1 Thinking in combination to mitigate model-specific bias in handling Georgian linguistic material. These models were applied to the primary task of evaluating the effectiveness of LLMs in identifying the meanings of neologisms and generating lexicographic definitions. The findings demonstrate that LLMs can serve as powerful lexicographic tools capable of effectively processing Georgian linguistic data. However, their performance and accuracy are highly dependent on the prompts designed and employed by lexicographers.
Speakers: Tinatin Margalitadze (Ilia State University), Giorgi Meladze (Ilia State University), Ketevan Mchedlishvili (Ilia State University), Tamar Laluashvili (Ilia State University), Giorgi Okropiridze (Ilia State University)
-
109
-
Learner's Lexicography & Dictionary Use 01 | Sitzungssaal (01 | Sitzungssaal, OeAW Main Seat, 1st floor)
01 | Sitzungssaal
01 | Sitzungssaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Hanna Fischer (Research Center Deutscher Sprachatlas)-
112
Pedagogical Value and Insights from Comparing EFL Dictionaries: Evidence from Apparently, Be on Speaking Terms, and Suffer (from)
Certain aspects of word meaning and usage become visible only when dictionaries are compared directly. This paper examines three cases—apparently, be on speaking terms, and suffer (from)—to show how cross-dictionary analysis reveals patterns that remain hidden in single-dictionary consultation. By comparing definitions, sense divisions, and examples across major EFL dictionaries (CALD, COBUILD, LDOCE, MWALED, and OALD), the study demonstrates the pedagogical value of such comparison. The analysis of apparently shows notable variation: CALD lists three senses, whereas others merge or reduce them. Example sentences further reveal mismatches, such as an OALD example illustrating a sense not explicitly defined there but found in CALD. The expression be on speaking terms illustrates how usage tendencies emerge only through comparison. Dictionaries show that the phrase appears predominantly in the negative, and positive uses typically describe the restoration of communication after conflict. The verb suffer shows distinct transitivity and collocational patterns. Cross-dictionary analysis reveals contrasts among intransitive use, suffer from, and transitive constructions, clarifying semantic tendencies and demonstrating the value of comparative consultation for more precise lexical understanding. Overall, dictionary comparison reveals usage patterns invisible in single sources, offering clearer insights for interpretation and learning.
Speaker: Shigeru Yamada (Waseda University) -
113
Dictionary Use in a Technology-Driven Era: Longitudinal Evidence from EFL and GFL Graduates
This paper presents findings from the third phase of a longitudinal repeated cross‑sectional study examining dictionary‑use patterns among graduates of English as a Foreign Language (EFL) and German as a Foreign Language (GFL) programmes in Hungary. Drawing on questionnaire data collected in 2020, 2023, and 2026, the study investigates how dictionary use is reshaped in an increasingly digital and AI‑mediated reference environment, with particular attention to generative AI‑based tools. The results indicate both continuity and change in reference practices. Online dictionaries and search engines remain widely used, but their dominant position is increasingly complemented by machine translation systems and generative AI tools. Print dictionaries continue to be widely owned, yet their frequency of use shows a gradual decline, pointing to a growing gap between ownership and actual use. The 2026 data provide early evidence that chatbots have become part of graduates’ language‑support repertoires, functioning as multifunctional tools that combine dictionary‑like, translation, and text‑support functions. The findings suggest a redistribution of reference practices within a hybrid digital environment and highlight the need to reconsider dictionary didactics in light of learners’ expanding engagement with AI‑based language technologies.
Speakers: Ida Dringo-Horvath (Karoli Gaspar University), Katalin P. Márkus (Károli Gáspár University of the Reformed Church) -
114
Lexicographic Functions Without Dictionaries: GenAI‑Mediated Lexical Problem‑Solving Among Non‑English‑Major Tertiary EFL Students
With the increasing presence of generative artificial intelligence (GenAI) in learning and teaching practices, English as a Foreign Language (EFL) pedagogy is being fundamentally reshaped. In particular, GenAI has emerged as a core resource influencing students’ approaches to learning and knowledge construction (Johnston et al., 2024; Ravšelj et al., 2025). Recent research indicates that students increasingly rely on GenAI chatbots not only for general writing support but also to resolve lexical problems and vocabulary challenges (Franjić & Božinovski, 2025; Xiao, 2024). Several studies have shown that large language model (LLM)-based chatbots such as ChatGPT can outperform traditional learner dictionaries in receptive and productive lexical tasks, especially in terms of speed and perceived usefulness (Božić Lenard & Šokčević, 2024; De Schryver, 2023; Lew et al., 2024; Ptasznik et al., 2024).
This is not to say that GenAI chatbots offer the same quality of lexicographic data as traditional dictionaries; rather, quality may no longer be a decisive factor for users anymore (de Schryver, 2023). This raises a critical question for lexicography: if dictionaries are no longer viable tools for most users, what role—if any—do lexicographic functions and knowledge still play in GenAI-mediated language use?
While a growing body of research has compared GenAI tools and dictionaries, it has largely focused on English-major students or controlled task-based comparisons (Chen, 2025; de Schryver, 2023; Lew et al., 2024; Mai, 2024; Ptasznik et al., 2024). Much less is known about non-English-major tertiary EFL students—who often lack dictionary literacy and increasingly rely on GenAI as their primary lexical access point. While some recent studies have begun to address GenAI use among ESP/EFL learners outside language-specialist programmes (e.g. Božić Lenard & Šokčević, 2024; Zhong, 2025), research remains limited.
This small-scale, user-centred study investigates lexicographic functions under GenAI mediation rather than dictionary use per se. Conducted in classroom settings in April 2026, it draws on data from approximately 40 non-English-major tertiary EFL students enrolled in English-language courses in Slovenia and Japan, all of whom regularly use GenAI tools and have received no formal dictionary training. The study examines (1) how students self-report their use of GenAI for lexical purposes, (2) how these perceptions align with observed lexical judgement when evaluating short GenAI-like texts containing collocational, patterning, and register-related issues, and (3) whether brief GenAI-mediated vocabulary exploration is associated with immediate uptake of target lexical items in a subsequent paper-based task.
The study adopts a one-session classroom design combining a brief survey (based on Ptasznik & Lew, 2025) with two task-based components (cf. Chen, 2025). First, a lexical judgement task completed without GenAI probes students’ ability to identify and revise near-miss collocations, grammatical patterning errors, and register mismatches, allowing comparison between self-reported confidence and actual lexical evaluation behaviour. This is followed by a time-limited GenAI-for-learning phase, in which students use a GenAI tool of their choice to explore selected lexical expressions, and a subsequent paper-based uptake task assessing their ability to use these items appropriately without tool support. The design enables triangulation of beliefs about GenAI as a lexical resource with observable judgement and short-term uptake, offering insights into the extent to which GenAI functions as a dictionary-like interface in L2 vocabulary learning and use.
The findings are expected to show that, even when dictionaries disappear as explicit tools in everyday writing practices, lexicographic functions remain central but become opaque under GenAI mediation (‘invisible lexicography’). In contemporary EFL contexts, the pedagogical relevance of lexicography may lie in sensitising learners to the lexical dimensions of AI output—such as sense, collocation, patterning, and register—that require human interpretation and evaluation.
Speakers: Biljana Božinovski (University of Maribor), Makimi Kano (Kyoto Sangyo University), Peter Gobel (Kyoto Sangyo University), Mihaela Franjić (University of Maribor)
-
112
-
Morphology, Word Formation & Grammatical Lexicography 00 | Anton Zeilinger Salon (00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor)
00 | Anton Zeilinger Salon
00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Veronika Lipp (ELTE Research Centre for Linguistics)-
115
Using LLMs in Dictionary Compilation: The Case of Classifying Adjectives in the Dictionary of the Slovenian Standard Language eSSKJ
In recent years, several lexicographic studies have addressed the question of which lexicographic processes can effectively make use of LLMs (De Schryver, 2023; Lew, 2024; McKean & Fitzgerald, 2024). Some studies are dedicated specifically to Slovenian lexicography (Gantar, 2024; Arhar Holdt et al., 2025), including Meterc & Jakop (2025) who have examined the use of LLMs in the compilation of the Dictionary of the Slovenian Standard Language eSSKJ (2016–), which is being compiled at the Fran Ramovš Institute of the Slovenian Language, ZRC SAZU, and published on the Fran dictionary portal (Ahačič et al., 2015).
This paper focuses on the performance of ChatGPT 5.5 in compiling eSSKJ dictionary entries for classifying adjectives. According to the editorial principles used in eSSKJ, classifying adjectives tend to have relatively uniform type-based definitions and predictable semantic groups of collocates, which makes them a suitable case study for testing the extent to which LLMs can successfully generate complete dictionary entries. The study tests the LLM’s performance in the compilation of 35 dictionary entries for classifying adjectives derived from nouns denoting sports, academic disciplines, and animals (e.g. košarkarski ‘pertaining to basketball’, geološki ‘pertaining to geology’, kengurujev ‘pertaining to kangaroo’), selected because the corresponding base nouns are often monosemous or exhibit regular polysemy.
The LLM was provided with the following background knowledge prior to prompting: (1) a typology of definitions used for classifying adjectives, (2) a training set of 72 previously published eSSKJ dictionary entries for classifying adjectives, (3) corpus data in the form of word sketches for these adjectives, and (4) up to 1,000 randomly selected concordances for each adjective from the Gigafida 2.0 Slovenian reference corpus (where fewer concordances were available, all examples were provided). Apart from the typology of definitions, which was supplied to the LLM in textual form, all other data were encoded in XML. The same types of data had originally been available to human editors during the compilation of the dictionary entries.
We evaluated the performance of the LLM in the following editorial tasks: (1) sense differentiation, (2) formulation of (type-based) dictionary definitions, (3) selection of typical collocations and their assignment to the appropriate dictionary senses, and (4) illustration of senses with typical dictionary examples drawn from the Gigafida 2.0 corpus. As a control measure, the same dictionary entries were also edited manually in parallel and were not uploaded to the LLM.
Testing the performance of the LLM in its editorial treatment of classifying adjectives in eSSKJ shows that it formulates type-based dictionary definitions in accordance with the established typology, although the ordering of predictable metonymic senses does not always follow the established hierarchy. It also selects appropriate collocates, both in terms of the range of collocates included and their assignment to dictionary senses, and successfully adheres to formatting guidelines for collocations (semantic clusters, formal and alphabetical ordering, etc.). The model also performs well in the selection of dictionary examples. In these respects, the LLM’s editorial treatment is comparable to dictionary entries produced by human editors. The LLM is also successful in distinguishing senses in dictionary entries that contain only predictable, type-based senses; however, it often fails to identify additional metaphorical or derived senses, as well as fixed expressions.
The results suggest that although ChatGPT 5.5 does not yet perform all tasks involved in compiling entries for classifying adjectives in eSSKJ entirely satisfactorily, it nonetheless represents a highly useful tool for lexicographers. At present, it can assist eSSKJ editors particularly with more routine tasks, which is especially valuable because these tend to be very time-consuming.
Speakers: Nina Ledinek (ZRC SAZU), Mija Michelizza (ZRC SAZU), Špela Petric Žižić (ZRC SAZU), Domen Krvina (ZRC SAZU), Janoš Ježovnik (ZRC SAZU), Andrej Perdih (ZRC SAZU) -
116
Learning Like Humans, Parsing Unlike Machines: A Construction-Based Approach to Lexical Annotation
Background and aim
Large language models generate fluent, dictionary-like and analytically plausible text, yet their linguistic knowledge is encoded implicitly in numerical representations rather than as inspectable lexical entries, paradigms, feature structures or constructions. Classical annotation frameworks make categories explicit, but often reduce complex linguistic behaviour to flat tags and relations. This creates a representational gap for lexicography: lexical knowledge must be searchable, correctable, reusable and explainable, not merely recoverable from model output. Drawing on Construction Grammar (Goldberg, 2006; Croft, 2022), this paper presents an experimental model for Albanian that restores the lexicon to the beginning of processing, while redefining it as a constructicon: a structured inventory of lexical and grammatical form-content pairings. The aim is not to replace large language models or Universal Dependencies (UD; de Marneffe et al., 2021), whose purposes differ, but to develop a linguistically accountable architecture for richer lexical resources and annotation.
Model and method
The model represents constructions as transparent nested Python structures that can later be exchanged through JSON, XML or RDF-oriented formats. Each entry contains a compound_set with three components: constituent_set, which records the internal morphosyntactic parts; pos_set, which assigns a category to the complete entry or phrase; and feature_sets, which stores one or more compatible feature bundles. For nouns and noun-associated categories, the current feature inventory comprises gender, case, definiteness and number. A value may be concrete or left open as "ok", allowing another construction to supply it during unification. Two lexicons are presently distinguished: lexical-category dictionaries for nouns, adjectives, pronouns, prepositions, verbs and related classes, and marker dictionaries for inflectional elements. In a deliberate refinement of broader Construction Grammar terminology, form is restricted here to the graphic or phonological matter that activates an entry, whereas lexical-grammatical information is organized within content. Processing matches textual forms to stored constructions, tests compatibility, and builds larger usage-level constructs through licensed unification.
Illustration and preliminary evaluation
The Albanian noun shef 'chief' illustrates the procedure. The lemma is a lexical construction which contains information about (masculine) gender but leaves case, number and definiteness open. Unification with the postposed marker -i produces shef-i 'the chief', analysed internally as a noun plus marker and externally as a noun phrase. The surface form initially licenses more than one feature analysis; in the larger construct shef-i i zbulim-it 'the chief of intelligence', contextual compatibility eliminates the unavailable reading, thus solving the ambiguity. The same mechanism models adjective and prepositional phrases, numerals, and the distinct processing roles of Albanian postposed and preceding articles. Grammatical agreement is thus represented as alignment information distributed across constructions, rather than as a label checked only after parsing. Initial domain-limited trials on nominal, adjectival, prepositional and selected verbal structures exceeded 90% accuracy and rose above 95% after missed constructions and compatibility patterns were corrected. These figures remain provisional because the test material, evaluation unit and error distribution have not yet been standardized.
Contribution and future development
The principal contribution is architectural. Annotation becomes the transparent result of recognizing stored constructions and projecting their information into constructs, rather than a separate layer of labels attached to bare tokens. This relocates POS and feature information within a richer representation and permits ambiguity without premature disambiguation. For Albanian lexicography, the approach supports a transition from digital dictionary to constructicon (Lyngfelt, 2018), encompassing lexical items, markers, phraseological patterns and larger constructions across the lexicon-grammar continuum. The present implementation is nevertheless incomplete: semantic frames, corpus-based frequency and productivity, broader verbal coverage, and external benchmark evaluation remain future tasks. Because its hierarchy is regular and explicit, the resource can also be mapped to linked-data models (Cimiano et al., 2016) and connected with corpus examples, ontologies, semantic frames and cross-linguistic correspondences. It therefore offers a practical basis for an interpretable, extensible and interoperable Albanian lexical resource.
Speaker: Edmond Cane (Luarasi University) -
117
The “Hell Of A” Construction in Casual Written English Online: A Synchronic Functional Account
Describing and classifying colloquial constructions that contribute to non-propositional meaning in discourse poses a persistent challenge for lexicography, particularly where interpretation depends on contextual inference rather than on stable form–meaning correspondences. Drawing on a 165-million-token corpus of Reddit posts and comments representing contemporary casual written English online, this paper explores the uses of the construction a/one hell of a and its orthographically reduced forms helluva and hella (cf. hella, OED, n.d.), with the aim of describing and operationalizing their synchronic functional range, building on ten Wolde’s (2023) and Brůhová and Vašků’s (2025) research into of-binominal phrases.
A grammaticalization cline of the construction proposed by ten Wolde (2023) distinguishes three recurrent uses of a/one hell of a. In the first use, the construction occurs in evaluative binominal noun phrases, where hell is interpreted literally or figuratively as attributing hellish, and by extension extremely bad, properties to the following noun (e.g. This is a hell of a council of war). In the second use, it functions as a broadened evaluative modifier (e.g. They were having a hell of a time), encoding a positive or negative extreme on a scale whose properties and interpretation are determined by the modified noun phrase and context (ten Wolde, 2023, p. 74). In the third use, the construction serves as a binominal intensifier modifying scalar adjectival premodifiers of nouns (e.g. I’ve got a hell of a good bulldog). The construction therefore operates with both implicitly and explicitly expressed scales “on which the speaker locates the degree of properties, which is a matter of speaker assessment and stance” (Ghesquière, 2014, p. 72). The process of grammaticalization of a/one hell of a therefore involves subjectification, with the bleaching of the descriptive meaning and the strengthening of the speaker’s perspective (Jucker, 2010, p. 116).
Combining ten Wolde’s (2023) diachronic account with the functional classification of binominal constructions proposed by Brůhová and Vašků (2025), and adapting both to the specific case of a/one hell of a and its reduced variants, we argue that given the semantic bleaching of hell, maintaining multiple evaluative distinctions or a separate descriptive function is neither necessary nor analytically productive for the analysis of contemporary usage. Instead, we distinguish three recurrent functions: evaluative, intensifying, and quantifying.
The consolidated evaluative function is attested for all three examined forms of the construction. Its defining property is the absence of an overt scalar or gradable element. In terms of encoded meaning, hell contributes only the notion of extremeness, while procedurally the construction prompts addressees to pragmatically infer a contextually appropriate evaluative scale on which this extreme value is to be interpreted, rather than encoding polarity or a specific evaluation independent of context. The category comprises all uses of a/one hell of a preceding nouns without modifiers, as well as nouns with non-gradable, non-scalar modifiers, in which it scopes over the NP as a whole (This whole thread has been a hell of a rabbit hole…). Helluva and hella perform the evaluative function as well, with hella only when preceded by an indefinite article (And nostalgia is a helluva drug. / …you made a hella profit).
The second functional category represents intensifying use of a/one hell of a and its reduced forms. In this category, the construction intensifies an explicitly expressed scalar or gradable property, typically realized by an adjectival premodifier of nouns (That’s a hell of a long drive), The intensifier a/one hell of functions as a booster, meaning “to a high degree” (ten Wolde, 2023, p. 109), and is felicitously paraphrasable by very. The construction frequently co-occurs with a lot, intensifying either quantity or degree, and is in these cases paraphrasable by quite, whether a lot functions as a pronominal quantifier (And young men have a hell of a lot of problems…) or as a degree adverb (He seems a hell of a lot happier than you. / It helps a hell of a lot.). The reduced forms helluva and hella display broader distribution beyond the noun phrase. Both can intensify adjectival phrases (This is helluva impressive. / That’s hella dope!) and hella is additionally attested as an intensifier of adverbial and verb phrases (…wake up hella early to make him breakfast…/ I hella respect you for this response).
Finally, hella, unlike the full construction a/one hell of a and helluva, performs a quantifying function. In this use, it occurs with both countable (I had hella snacks waiting for me) and uncountable nouns (But yea they got hella charisma…) and is paraphrasable as a lot of.
The paper proposes an operationalized, pattern-based classification of the functions performed by a/one hell of a and its reduced forms in contemporary casual written English online. This classification constitutes a first step toward implementing these distinctions as deterministic logic in an automated analysis pipeline, enabling systematic and replicable identification of functions.
Speakers: Gabriela Brůhová (Charles University), Veronika Raušová (Charles University) -
118
Improving Dictionary Headword Lists with Word-Prevalence Metrics: An Examination of Low-Prevalence, High-Frequency Words in Slovenian
Objective
This study evaluates whether word-prevalence data can improve the selection and review of dictionary headword candidates. It focuses on Slovenian lemmata that combine higher corpus frequency with lower prevalence, because this mismatch may signal problems that frequency alone does not reveal.
Background and data
Word prevalence is the percentage of a population that knows a word (Keuleers et al., 2015). Large online lexical-decision studies have produced prevalence norms for Dutch, English, Spanish, Catalan, and Italian, based on tens of thousands of words and large participant samples (Keuleers et al., 2015; Brysbaert et al., 2019; Aguasvivas et al., 2018, 2020; Guasch et al., 2023; Amenta et al., 2025). These studies show that very frequent words are generally widely known, whereas words at lower frequencies vary considerably in prevalence. Lew & Wolfer (2024) represents a rare example of applying word-prevalence data to the field of lexicography.
The ongoing Slovenian megastudy has so far collected one hundred YES/NO responses for each of 35,000 words from 39,063 participants in 48,659 sessions; the complete list contains 79,413 words. In each session, participants classify 120 letter strings as Slovenian words or nonwords. Among the 120 letter strings given in each session, 84 were real Slovenian words and 36 were not, The list was compiled from existing dictionaries (Perdih et al., 2025), which provide verified words (as opposed to potentially unverified corpus lemmata) and to serve as a motivational element at the end of a session – lists of successfully and unsuccessfully recognized words linked to the Fran Slovenian dictionary portal (https://www.fran.si) are provided to participants at the end of each session. Consequently, frequent corpus lemmata absent from the dictionaries are outside the present analysis.
Selection and classification
We selected the 50 highest-frequency lemmata in the Gigafida reference corpus of Slovenian (Krek et al., 2020) whose prevalence values were below zero, meaning that fewer than half of the participants recognized them. The least frequent item, masen ‘pertaining to weight’ occurred 832 times (0.62 per million tokens). We then classified the likely sources of the frequency–prevalence mismatch.
Results
Lemmatization or tokenization errors account for 15 items (30%) and often inflate corpus frequencies. A further 17 items have dictionary forms that are uncommon in actual use: 13 adjectives, 3 verbs, and 1 pronoun. For example, the adjectival headword kaven ‘pertaining to coffee’ was recognized by 45.23% of respondents (prevalence value -0.118) despite a frequency of 4,094 (3.07 per million tokens); its more frequent definite form kavni would probably be easier to recognize. Similarly, the infinitives poiti ‘to run out’ and onemoči ‘to tire, to languish’ are much less common than their inflected forms.
Other mismatches reflect restricted use or ambiguous status. Eight items are terms, including obtežba ‘load, weight’; eight may be interpreted as foreign, although some are Slovenian loanwords, such as go ‘Japanese game’ and song ‘a song genre’, while others are now archaic or obsolete and only occur in the corpus as parts of foreign-language texts (il ‘clay’, but also Italian definite article). The two interjections hi and jah may be difficult to recognize as Slovenian words without context, and hi may also be associated with the English greeting. Two items are homographs of proper names. Some categories overlap: for example, Li may be lemmatized as li, while li may also result from incorrect tokenization of a masculine plural past-participle ending.Implications for headword-list development
Word prevalence is useful as a diagnostic measure for high-frequency headword candidates. When combined with corpus frequency, it helps identify lemmatization and tokenization errors, specialized terms, foreign-language material, and items whose recognition depends strongly on context. These findings can support both the addition of new entries and the review or removal of existing entries: low prevalence may indicate corpus noise, opens a question of headword form user-friendliness, or suggests that a stylistic or terminological label may be required rather than simple exclusion of the headword.
Speakers: Andrej Perdih (ZRC SAZU), Janoš Ježovnik (ZRC SAZU), Dejan Gabrovšek (ZRC SAZU) -
119
The Learner Dictionary of Italian Collocations (DICI-A): A New Lexicographic Resource
We introduce and describe the DICI-A, a learner dictionary of Italian collocations which is the output of a two-year project funded by the Italian Ministry of University and Research. In the Italian lexicographic landscape, three dictionaries of collocations have been published in the last fifteen years (Lo Cascio, 2013; Tiberii, 2012; Urzì, 2009). However, none of the three has two fundamental features that can be found in DICI-A: it is specifically targeted at L2 learners of Italian, and it has been created according to corpus-based criteria. Other features of the DICI-A are:
- it is monolingual;
- it includes over 10,000 collocation entries belonging to six syntactic configurations;
- each collocation is assigned to a specific proficiency label;
- GenAI and human assessment have been integrated for the creation of definitions and examples;
- it is digital and freely available online, at https://dictionary.dici-a.it/.
We relied on a broad definition of collocation: a co-occurrence of two words with a syntactic relation characterised by its conventional meaning, resulting from the number of times it is used in naturally occurring language (frequency), the range of texts where it occurs (dispersion) and the extent to which its components attract each other and are strongly associated (association measures), either adjacently or within a distance.
In addition to a general description, the presentation will focus on three specific features of the dictionary.
1) Identification and selection of the dictionary entries
We have included in the dictionary collocations that fall into six syntactic types: i. Verb + Direct object (vdobj; mantenere una promessa, ‘to keep a promise’); ii. Adjective + Noun/Noun + Adjective, the adjective is a modifier before or after a noun (amod; brutta avventura, ‘bad adventure’; tempo libero, ‘free time’); iii. Verb + Adjective, the adjective functions like an adverb by modifying the verb (advmod1; stare zitto, ‘to stay quiet’); iv. Verb + Adverb, the adverb modifies the verb (advmod2; fare presto, ‘to hurry up’); v. Adverb + Adjective, the adjective is modified by the adverb (advmod3; altamente positive, ‘highly positive’); and vi. Noun + Noun, compounds made of two adjacent nouns (comp; parco divertimenti, ‘amusement park’).
The collocations belonging to these six syntactic configurations were extracted from the PEC24 (Spina et al., 2025), a large reference corpus of written and spoken Italian, by combining pos-tagging and syntactic parsing (Seretan, 2011). These initial 2 million candidate collocations were filtered via a multi-method approach (Spina et al., 2026) involving both automatic stages (based on five different measures: dispersion, frequency, mutual information, log-dice and log-likelihood) and human evaluation, after a comparison between the filtered candidate collocations and two existing non corpus-based Italian collocation dictionaries.
2) Attribution of proficiency labels to each entry
The 10,596 collocations resulting from this selection process were assigned CEFRCV-based proficiency labels (Common European Framework of Reference Companion Volume; Council of Europe, 2020) based on a set of quantitative and qualitative criteria, including corpus frequency and dispersion, semantic transparency, register, and CEFRCV vocabulary range descriptors. This procedure was entirely based on human assessment: each collocation was annotated independently by two annotators; then a third annotator resolved the cases of disagreement. Inter-annotator agreement varied across collocation types; as an example, it reached 80% for the verb-direct object collocations.
3) Creation of definitions and examples with the support of GenAI
We relied on recent literature (Lew, 2023; Lew et al., 2024) showing that GenAI can be effective in speeding up the process of dictionary creation, performing well both in writing definitions and in producing examples. Thus, we asked ChatGPT-4o to assist us in identifying collocation’s meaning(s) and in providing us with definitions and examples for each meaning. In our prompt, we strongly emphasised the need for the output to contain learner-friendly vocabulary suitable for learners of Italian. The 10,596 AI-generated pairs of definitions and examples were validated by human lexicographers and replaced or modified accordingly. After this evaluation process 72% pairs could be accepted as they were in the AI-generated version; 12% required changes in the definition, 12% had to be modified in the example, and only 4% had to be completely rewritten.
Speakers: Stefania Spina (Università per Stranieri di Perugia), Irene Fioravanti (Università per Stranieri di Perugia), Fabio Zanda (Sapienza Università di Roma)
-
115
-
Corpus Linguistics & Corpus-Based Resources 01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)
01 | Johannessaal
01 | Johannessaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Ana Ostroški Anić (Institute for the Croatian Language)-
120
Corpus-Based versus AI-Generated Dictionary Examples: How They Differ and What Learners Prefer
A distinguishing feature of English dictionaries for learners is that most entries include example sentences or phrases that have been copied or adapted from corpora. Their function is to reinforce meaning by showing how words have been used in context, prioritising typical grammar patterns and lexical collocations (Fox 1987; Frankenberg-Garcia 2014; Frankenberg-Garcia, Rees and Lew 2021). While early corpus-based lexicographers had to scan concordances line by line to select good dictionary examples (Krishnamurthy 1987), nowadays tools like Word Sketches (Kilgarriff et al. 2014) and GDEX (Kilgarriff et al. 2008) shorten the time it takes for lexicographers to find suitable corpus examples. Despite these welcome developments, it is now arguably easier and faster to generate dictionary examples using LLMs. However, in early experiments evaluating AI-generated examples, experts found them to be redundant, unimaginative and inauthentic (Lew 2023, Jakubíček & Rundell 2023), although improvements could be seen after some prompt fine-tuning (Lew 2023). In this study, we wanted to find out how language learners rated AI-generated and corpus-based dictionary examples, to come to a better understanding of how they differ, and to consider the practical implications for lexicography.
Over 200 university students pursuing various degrees (language and non-language) participated in an online task involving 15 monosemous lexical items (from to different part-of-speech categories) that were likely to be unknown. For each lexical item, the participants were presented with the definitions from the Reverso English Dictionary (AI-generated) and the Oxford Advanced Learner’s Dictionary (corpus-based) followed by one example from each dictionary. As most entries contained more than one example, the ones presented to each student were randomly selected from the totality of examples available. The participants were asked to rate the 15 AI-generated and the 15 corpus-based examples they saw on a 1-5 scale (poor to excellent). The order of lexical items, definitions and examples shown to each student were all randomized. At the end of the task, the participants were asked to explain what made them give examples high or low ratings. They then responded to a few demographic questions and completed a standardized vocabulary test (LexTALE, Lemhöfer & Broersma 2012).
Prior to the data collection, the authors of this study conducted a data-driven appraisal of all 63 examples in the dataset. By comparing interpretive coding criteria and discussing divergences, we collaboratively negotiated a systematic analytical framework for describing dictionary examples. Using this framework, we noted a number of significant differences between the AI and the corpus-based examples (e.g., in the corpus-based examples, there was significantly more variability in the number of words and syntax used). However, not all differences were found to be significant (e.g., the presence of contextual cues about meaning).
With regard to the student ratings, overall preliminary findings indicate that there was no significant difference between examples in each resource. However, when only complete sentences were taken into consideration, the Oxford examples were rated significantly more positively.
The full results of the study, including triangulation of our analytical framework for describing dictionary examples along with user ratings and student demographics will be presented at the conference. We believe the study will help to shed new light on how corpus-based and AI-generated dictionary examples can be refined in the future.
Speakers: Tomasz Michta (University of Bialystok), Geraint Paul Rees (Pompeu Fabra University), Ana Frankenberg-Garcia (University of Surrey) -
121
AI-Assisted Detection and Mitigation of Lexicographic Bias: Semantic Reduction in Definitions of Five Socio-Legal Terms
When non-specialist users consult general dictionaries to understand legally significant terms, they risk encountering definitions that are fundamentally incomplete. This study introduces a reproducible computational methodology to quantify the level of semantic reduction and explore AI tools to resolve it. Five socio-legal terms were examined from the Cambridge Dictionary, Oxford English Dictionary, Merriam-Webster (general) and Black’s Law Dictionary, Oxford Dictionary of Law, Merriam-Webster’s Dictionary of Law (legal): harassment, stalking, discrimination, abuse, and neglect. General and legal definitions (35 and 15 respectively) were vectorised using Sentence-BERT and compared via cosine similarity. None of the five terms reached the 0.70 similarity threshold. Mean similarity ranges from 0.491 (abuse) to 0.571 (stalking); 92.4% of all pairwise comparisons fall below this threshold. Statistical significance was confirmed by one-sample t-tests (p < 0.01). The Qwen2.5-3B-Instruct model extracted key legal components for each term, which can serve as editorial checklists for lexicographers reviewing dictionary entries. The findings demonstrate that semantic reduction is systematic, measurable, and varies across terms. AI tools can assist lexicographers in detecting and addressing such gaps, though final decisions remain with the human expert.
Speaker: Mariia Kopylova (Roma Tre University) -
122
Mapping Sociolinguistic Theory into Practice: Headword Candidate Selection in the Dictionary of Pluricentric Portuguese
Presently in its preliminary phase at the Research Centre for General and Applied Linguistics at the University of Coimbra (CELGA-ILTEC), the dictionary of pluricentric Portuguese project seeks to develop an online dictionary that documents the usage of Portuguese across diverse global contexts. Considering the socio-historical complexities inherent to the Portuguese language area (Faraco, 2016; Correia, in print), this project’s lexicographic work is grounded in a deeper understanding of language use and its social context, with sociolinguistic theory serving as a core guiding framework. The purpose of this paper is to present how theoretical insights from sociolinguistics translate into practical lexicographic decisions regarding headword candidate selection.
Portuguese is the official language of nine countries (Angola, Brazil, Cabo Verde, Guinea-Bissau, Equatorial Guinea, Mozambique, Portugal, Sao Tome and Principe, and Timor-Leste) and one territory (Macao), although its actual functional status varies significantly in these multilingual regions. Traditionally, Portuguese is considered to have two dominant varieties, Brazilian Portuguese and European Portuguese, with the latter being adopted as the norm in the other countries. The lexicography of the language reflects this bicentrism and dictionary production is largely centred in Brazil and Portugal (Kuhn & Correia, 2025). However, in countries where Portuguese was introduced as a result of colonisation (all countries except Portugal), research has revealed discrepancy between real used norms and the official discourse legislating the language, including teaching materials, reference instruments such as dictionaries, and official documents. In Brazil, for instance, the dominant exonormative view of the language has led to a significant gap between how language is actually used by formally educated people in contexts requiring more language monitoring, i.e., the cultivated standard, and its prescriptive use, or the ideal language standard (Faraco, 2008). Relatedly, studies on the real norms used in the other countries include the description of Mozambican Portuguese by Gonçalves P. (2010) and Firmino (2011), Angolan Portuguese by Adriano (2015) and Inverno (2009), and Sao Tomean Portuguese by Gonçalves R. (2016) and Bouchard (2017). Considering all this in this dictionary project requires, in addition to a robust foundational sociolinguistic element, the compilation of a corpus that reflects the real use of Portuguese, in a variety of contexts and in different territories. Since a corpus with these characteristics is not available, a decision was made to carry out a pilot-dictionary project, using a small corpus of tweets (Rodrigues Gomide, 2022). The purpose of the pilot is to try methodologies, evaluate results, and develop solutions to the problems found, in a smaller scale, alongside the compilation of a larger corpus and further construction of the theoretical framework. For that, we have adopted the Dictionary Express method (Baisa et al., 2019). The planned workflow consists of several steps of iterative collaboration between human lexicographers and automatically extracted data, with headword candidate selection as the first step. This involves headword annotation using a specially created interface, with different lexicographers assigning flags (e.g., not a lemma, non-standard, etc.) to the same set of candidates and final headword list definition resulting from general agreement among lexicographers. Differently from previous projects for other languages (Baisa et al., 2019; Blahuš et al., 2023; Kovařík et al., 2024), in our project this method is used to make a dictionary that describes, at the moment, eight varieties of one language. One additional challenge stems from the sociolinguistic grounding of the project, by which traditional views on standard language are questioned. As a result, some changes have been made, such as the decision to not filter the list of automatically extracted candidate headwords for part of speech. This is because POS-taggers have been developed for Brazilian and European Portuguese, thus words from other varieties or words that are not considered to be standard might not be identified, leading to wrong POS-tagging which, in turn, would affect the frequency count. Another aspect of the headword candidate annotation process that is especially crucial to our project refers to the flag “non-standard”. This type of decision-making requires a previous common-ground perspective on standardisation and standard language ideologies (McLelland, 2021) and how they affect the representation of language in the dictionary. This consideration becomes even more critical in the case of Portuguese, which is spoken across multilingual regions marked by complex socio-political-cultural contexts, particularly considering the impact of colonial history and the enduring legacy of colonial language policies. At the time of writing, this step of the workflow is still ongoing; however, further results are expected for the time of the conference. With the presentation of how the theoretical framing of this project maps into practical lexicographic work regarding headword candidate selection, we hope to share our experience with our peers, as well as contribute to questioning and changing long-established language ideologies and language attitudes prevailing in the Portuguese language area.
Speakers: Tanara Zingano Kuhn (CELGA-ILTEC, University of Coimbra), Vojtěch Kovář (Lexical Computing), Miloš Jakubíček (Lexical Computing)
-
120
-
Digital Lexicography: Design, Data Modelling & Theory 01 | Sitzungssaal (01 | Sitzungssaal, OeAW Main Seat, 1st floor)
01 | Sitzungssaal
01 | Sitzungssaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Hanna Fischer (Research Center Deutscher Sprachatlas)-
123
The Theoretical Foundation of Dictionary Definitions
The place of theory in lexicography has been the matter of intensive debate. There are different views of what a theory is. In line with most approaches in philosophy of science, it is assumed here that a theory has to provide an explanation. It is argued that lexicography should be considered as an applied science, i.e. the same category as medicine. In the analysis of applied science, the identification of problems and the specification of evaluation criteria for solutions are crucial. In lexicography, problems are the reasons why users consult a dictionary. Problem types are distinguished on the basis of the kind of information they need and the background knowledge they have. For certain problem types, definitions can be used as solutions, but the requirements on the definition depend on the type of problem it is meant to solve. Guidelines issued to lexicographers working on a dictionary are not theories, because they do not explain. They only instruct. The theoretical foundations of a theory should identify criteria for distinguishing problem types.
Speaker: Pius ten Hacken (University of Innsbruck) -
124
Large-Scale Semiautomatic Detection of New and Unregistered Words in Modern Ukrainian
We have created and implemented a workflow for large-scale semiautomatic detection of new and previously unregistered words in Modern Ukrainian. Crucially, it involves a diverse corpus of Ukrainian (2 billion tokens), a large electronic morphological dictionary of Ukrainian, and a specially crafted NLP toolkit. The key methodological innovation lies in the dual use of the dictionary. On the one hand, it serves as a resource for lemmatizing and morphological tagging the corpus; on the other, it is cyclically enriched with new words detected in the corpus. The corpus is dynamic, as it is regularly expanded by adding a variety of texts from the 19th century to the present day. The dictionary, currently containing 444,000 lemmas, is dynamically reused with the help of a tailor-made NLP toolkit as an exclusion source (Janssen, 2009) against the growing corpus. The paper provides a detailed step-by-step description of the procedure. This semiautomatic pipeline has yielded thousands of new and unregistered Ukrainian words, laying the foundation for the systematic, large-scale exploration of the Modern Ukrainian lexicon.
Speakers: Vasyl Starko (Ukrainian Catholic University), Andriy Rysin (Independent Researcher) -
125
How Do You Define the (Known) Unknown? Prompt Engineering as a Method for AI-Assisted Meaning Definition in the Pandemictionary
The formulation of definitions is a core task of lexicography, traditionally carried out manually. Large language models now offer new opportunities to automate this workflow. While AI-assisted definition generation has been well studied for well-known vocabulary, this is not yet the case for lesser-known historical vocabulary, which is scarcely represented in training data. This paper therefore investigates how far targeted prompt engineering can improve the quality of AI-generated definitions for historical pandemic vocabulary, using data from the cholera corpus developed within the Pandemictionary project. Eight headwords of varying frequency and semantic complexity are analysed across three prompt configurations, focusing on the role of corpus evidence in generating lexicographical definitions. The results show that AI systems can generate formally correct, lexicographically structured definitions. Prompts without evidence draw on general model knowledge and often produce anachronistic or hallucinated content, whereas evidence-based prompts yield definitions that are more historically accurate and firmly grounded in the corpus. Corpus evidence proves a crucial quality factor, especially for rare or poorly documented lexemes, though challenges remain regarding hallucinations, polysemy, and reducing definitions to their essential elements. The study highlights the potential of prompt engineering for digital lexicography while underscoring the continuing central role of human expertise in quality assurance.
Speakers: Rukiye Smilla Burkart (Trier University), Susanne Kabatnik (Trier University)
-
123
-
11:00
Coffee Break 00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna -
Keynote: Dictionaries as Pipelines: Lexicographic Work at the Lisbon Academy of Sciences 01 | Festsaal (Festive Hall), OeAW Main Seat, 1st floor
01 | Festsaal (Festive Hall), OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Ana Salgado (FLUP – University of Porto)-
126
Dictionaries as Pipelines: Lexicographic Work at the Lisbon Academy of Sciences
This keynote reframes dictionary-making as the design of a living, governed pipeline: a system that turns linguistic evidence into publishable lexical knowledge while keeping editorial decisions traceable and revisable. Drawing on current work at the Lisbon Academy of Sciences across a portfolio of initiatives, it shows how this pipeline supports complementary goals: a continuously updated Dicionário da Língua Portuguesa (DLP), inclusive and participatory lexicographic practices (ACL+), a collaborative lexicographic platform mapping variation across the Community of Portuguese Language Countries (CPLP) via the Atlas Lexicográfico da Língua Portuguesa (ALLP), and the reuse of lexical data in infrastructures and language technologies (PORTULAN).
The pipeline starts by converting legacy materials into structured, machine-actionable data (such as XML) and by applying consistent normalisation and TEI-based annotation. Evidence is then continuously gathered from corpora, query logs, and community contributions, and assessed through a combination of quantitative signals and expert editorial judgement. A central theme is governance: how inclusion and revision decisions are documented, reviewed, and made auditable—especially when balancing pluricentric Portuguese.
Extending the pipeline from dictionary entries to atlas-style knowledge, ALLP illustrates how lexical description can also support cartography of variation, combining curated entries with usage evidence and, where relevant, multimedia documentation. The keynote concludes with practical lessons—what scales well, what slows teams down, and how to monitor quality over time.
Speaker: Ana Salgado (FLUP – University of Porto)
-
126
-
12:30
Lunch Break 00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna -
Computational & AI-Based Lexicography 01 | Sitzungssaal (01 | Sitzungssaal, OeAW Main Seat, 1st floor)
01 | Sitzungssaal
01 | Sitzungssaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Gilles-Maurice de Schryver (Ghent University)-
127
A Hybrid Approach to Web Accessibility Evaluation in Digital Lexicography: Comparing and Integrating Rule-Based Validators and Generative AI
Electronic dictionaries pose particular accessibility challenges for blind screen reader users, yet the field lacks evaluation frameworks that account for the complexity of their information structures. This paper reports on a three-phase study comparing Automated Web Accessibility Evaluation Tools and Large Language Models in evaluating three online dictionary entries from different lexicographical traditions. Phase 1 establishes a baseline through automated WAET scanning and zero-shot LLM auditing. Phase 2 compares their outputs and derives a four-category framework, reframing compliance as a design question. Phase 3 tests two enhanced prompting strategies — literature-augmented and design-oriented. The former uses specific dictionary accessibility literature while the latter combines a structured checklist derived from Phase 1 with persona-based use-type simulation, surfacing a macro-to-micro hierarchy of dictionary-specific barriers. The study argues for an iterative, multi-condition hybrid workflow integrating WAETs, LLMs, and expert human review for e-dictionaries.
Speakers: Jesús Torres del Rey (University of Salamanca), M.ª Teresa Fuentes Morán (University of Salamanca) -
128
Divide and Define with LLMs: Generating Definitions for the Digital Dictionary Database of Slovene
This paper presents an experiment on generating definitions for Slovene headwords using four Large Language Models (LLMs): two commercial (Gemini and GPT) and two open-source (Gemma and the Slovenian model GaMS). Each model was provided with contextual data, including semantic indicators, collocations, and examples. We tested a zero-shot approach and two few-shot approaches using sample definitions from the Digital Dictionary Database of Slovene (DDDS) and the New Dictionary of Standard Slovene. The results show that while all models produce high-quality definitions, Gemini performs best across all word classes, particularly when utilizing DDDS sample definitions. The second part of the study evaluated LLMs as judges of definition quality, testing Gemini, GPT, and Claude-Sonnet. Gemini again achieved the highest agreement with human lexicographers. Notably, even when LLMs selected a different definition than the experts, their choices were usually viable. These results demonstrate the considerable potential of LLMs in lexicography, and we plan to deploy this pipeline to generate definitions for a much larger set of Slovenian headwords.
Speakers: Iztok Kosem (University of Ljubljana & Jožef Stefan Institute), Polona Gantar (University of Ljubljana, Faculty of Arts & Faculty of Computing and Information Science), Simon Krek (Jožef Stefan Institute), Tjaša Arčon (University of Ljubljana, Faculty of Computing and Information Science) -
129
LLM-Assisted Lexicography in Low-Resource Contexts: From Prompt to Dictionary Article
This paper investigates the potential of Large Language Models (LLMs) to support lexicographic work in low-resource contexts, focusing on Austrian Bavarian dialects documented in the Wörterbuch der bairischen Mundarten in Österreich (WBÖ). Using a structured prompt-engineering workflow and a three-stage data pipeline, 100 dictionary articles were generated with LLaMA 4 (Scout) and systematically compared to human-authored “gold” articles. Evaluation combined automated metrics, LLM-based judgement, and assessment by three human lexicographers. Results show that LLMs reliably reproduce formal lexicographic conventions, producing well-structured and coherent dictionary entries. Nevertheless, limitations emerge in semantic completeness and data fidelity: generated articles sometimes omit senses or misrepresent semantic structure and often inadequately integrate empirical evidence, especially regional information and attestations. While automated metrics indicate a relatively high similarity of original and generated senses based on their embeddings, both human evaluators and LLM-as-a-judge approaches reveal persistent weaknesses. However, the strong alignment between human judgements and the sense similarity metric suggests that embedding-based evaluation is a promising proxy for human assessment, whereas holistic LLM-as-a-judge scores appear comparatively strict and less transparent. The study argues for a hybrid evaluation framework and positions LLMs as assistive tools within human-in-the-loop lexicography.
Speakers: Katharina Korecky-Kröll (Austrian Academy of Sciences, ACDH), Philipp Stöckle (Austrian Academy of Sciences, ACDH), Daniel Elsner (Austrian Academy of Sciences, ACDH), Wolfgang Koppensteiner (Austrian Academy of Sciences, ACDH), Elisabeth Eder (Austrian Academy of Sciences, ACDH)
-
127
-
Digital Lexicography: Design, Data Modelling & Theory 01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)
01 | Johannessaal
01 | Johannessaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Margit Langemets (Estonian Language Institute)-
130
Developing a Digital Version of Gidor Bilbao’s Latin-Basque Dictionary
In this article, we describe several steps taken to convert Gidor Bilbao’s Oinarrizko Hiztegia Latina-Euskara (Basic Latin-Basque Dictionary) from a paper-based resource and move toward presenting it as Linked Data. Starting from a document in MS Word format, we attempted to isolate and tag elements of the dictionary’s mac-ro- and microstructures. The goal of the experiments outlined here was to capture the content of those ele-ments in an innovative format based on the Linked Data paradigm, using the Ontolex-Lemon model. Once inte-grated into Wikibase, the data are ready for final structuring or further enrichment.
Speaker: David Lindemann (EHU University of the Basque Country) -
131
An AI-Driven Natural Language Interface for Graph Database Queries in a Lexicographical Resource
Complex digital lexicographical resources often feature rich data structures that are difficult for non-specialist users to query. This paper presents a pilot study on an LLM-driven natural-language interface for the Lehnwortportal Deutsch, a graph-based database of German loanwords in other languages. Rather than translating user questions directly into the database query language, the proposed architecture maps natural-language requests to the JSON configuration format already used by the portal’s visual query builder. This intermediate representation makes it possible for users to inspect and modify query builder configurations generated from their questions. We discuss the advantages of this approach and present a pilot study that evaluated it on a set of 310 German questions that, due to lack of real user questions, had to be AI-generated using several different strategies to cover a large range of usage scenarios as well as available querying options. Using one-shot prompting, a state-of-the-art commercial-grade LLM produced valid JSON in almost all cases; 260 outputs were rated perfect and 23 acceptable by human assessment under a low reasoning budget. A high-budget rerun improved most initially bad cases, though with impractical latency for real-time use. The paper also discusses methodological challenges in generating and evaluating synthetic query sets, handling unsupported user requests, and extending the approach through larger datasets, comparing multiple language models, varying output parameters, and using fine-tuning as an alternative strategy.
Speaker: Peter Meyer (Leibniz Institute for the German Language) -
132
PoliLex: Modelling Heterogeneous Lexical Resources: Lessons from the Dictionary of Polish Dialects
The Institute of the Polish Language (Polish Academy of Sciences) maintains a body of dictionaries and card-file archives assembled over more than a century, whose differences in structure, editorial convention, and descriptive metalan-guage have long kept them difficult to search together. This paper reports on PoliLex, an effort to bring such material under a common, machine-queryable representation without erasing the editorial character of each source. We de-scribe an encoding and publication workflow that moves corrected dictionary text through TEI into an OntoLex-Lemon Linked Data representation, and a layered project ontology whose load-bearing decision is to keep a dictionary, as a published text, separate from the language system it records. We illustrate the approach with the Dictionary of Polish Dialects, the first resource processed through the full pipeline, and show how the same encoded source is served both as browsable full text and as a SPARQL-queryable graph. We then examine what such harmonisation is designed to achieve across resources, what it achieves today on a single resource, and where it stops — notably at cross-dictionary sense alignment.
Speakers: Krzysztof Nowak (Institute of Polish Language PAS), Dorota Mika (Institute of Polish Language PAS)
-
130
-
Historical & Diachronic Lexicography 02 | Museumszimmer (02 | Museumszimmer, OeAW Main Seat, 2nd floor)
02 | Museumszimmer
02 | Museumszimmer, OeAW Main Seat, 2nd floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Alfred Lameli (Research Center Deutscher Sprachatlas)-
133
Towards a Handbook of Persian Lexicography: A Lexicographical-Functions Analysis of 20th–21st Century Dictionaries
This article examines four major Persian dictionaries—Dehkhoda (1940), Mo’in (1972), Sokhan (2003), and Aryanpur (2009)—through Tarp’s Function Theory of Lexicography (2008). Using a descriptive-analytical mixed-methods design, the study evaluates the dictionaries in terms of four lexicographical functions: text reception, text production, translation, and knowledge acquisition. The analysis is based on 12 lexical and grammatical categories, with three sampled lexical items from each category, yielding 144 dictionary-entry observations across the dataset. The findings reveal clear functional asymmetries. Dehkhoda performs most strongly in knowledge acquisition and text reception; Mo’in preserves much of this interpretive orientation in a more compact form; Sokhan demonstrates the strongest monolingual performance in text reception and contemporary communicative usability; and Aryanpur, as a bilingual dictionary, is highly effective for translation but very weak in knowledge acquisition. Overall, the study shows that modern Persian lexicography has moved gradually toward greater user orientation, though without achieving balanced multifunctionality. It concludes by arguing for stronger dialogue between Persian lexicography and international lexicographical scholarship and recommends the compilation of a Handbook of Persian Lexicography.
Speakers: Saghar Sharifi (Karaj Islamic Azad University), Abolfazl Alamdar (Allameh Tabataba'i University, Tehran, Iran) -
134
The Effects of Purism on a Nineteenth-Century Icelandic Bilingual Dictionary
The Icelandic linguist and philologist Konráð Gíslason (1808–1891) was the author of an influential Danish–Icelandic dictionary published in 1851, the first bilingual dictionary of its kind. The significance of this work lay less in its impact on lexicography than in its role in the history of linguistic purism in Iceland. One of its most distinctive features was the near absence of loanwords as Icelandic equivalents for Danish headwords, despite the widespread use of such loanwords in both spoken and written Icelandic at the time. Although the vigorous coinage of neologisms was a hallmark of nineteenth-century linguistic purism in Iceland, Gíslason’s dictionary contains relatively few newly coined terms. Instead, Gíslason relied primarily on existing vocabulary and occasionally revived archaic words attested in medieval texts as equivalents for Danish lemmas. He also frequently paraphrased the meaning of the Danish words through explanatory phrases rather than resorting to loanwords. His uncompromising puristic ideals limited the dictionary’s practical usefulness and made it a less accurate representation of contemporary Icelandic usage.
Speaker: Jóhannes B. Sigtryggsson (The Árni Magnússon Institute for Icelandic Studies) -
135
The Making of Meaning: From the 1980 Communist Albanian Dictionary to the Great Dictionary of 2026
In the Albanian totalitarian political context, linguistic resources were considered an important tool for the installation of the guiding ideology and the materialist worldview. Dictionaries, being reference works, were very favourable and effective for normalizing any ideological perspective, even presenting them as objective linguistic truths. (Fairclough, 1992, p. 87). Our study aims at identifying the techniques of ideological manipulation imposed on the Albanian lexicographic practice of the communist period. We have taken into consideration The Dictionary of the Contemporary Albanian Language (1980) compiled during the communist period and compared it with The Great Dictionary of the Albanian Language (2026) published online in early 2026 with the aim of analysing how the ideological dimension is reflected in two different historical and methodological periods.
The research poses an explicit question regarding the influence of ideological constraints on lexicographic methodology, macrostructure, and microstructure in Dictionary of 1980, and their differences in relation to more descriptive and corpus-based approaches in Great Dictionary of the Albanian Language (2026). It further investigates which lexicographic tools, like lemma selection, definition strategies, illustrative examples, and usage labelling, carried ideological meaning in the period of communism. In its methodological design, the study is grounded in a qualitative and quantitative comparative analysis of the two dictionaries, addressing both macrostructural and microstructural levels of lexicographic analysis, including selection and organization of headwords, criteria of selection and exclusion, definitions, examples, and evaluative labels that reflect ideological interventions in lexicography. Particular attention is given to those examples and labels that express political and ideologically coloured meanings.
In compiling The Dictionary of 1980, Albanian linguists worked in a political environment, where Marxist-Leninist doctrine shaped both the language policies, and the scientific production. Consequently, this dictionary would reflect not only the codification of the linguistic norm, but also ideological interventions that affected both the macrostructure and the microstructure levels. Terminology from the political field, concepts related to religion, etc. were given a strong ideological content, while terms such as socialism, internationalism were often defined through positive explanations, typical of the language of propaganda.
The 1980 Dictionary was based entirely on Russian lexicographic practice, and the Soviet orientation even appears openly. (Lloshi, 2025, p.36). In fact, the explanation of the term socialism is a translation from the Dictionary of the Russian Language (vol. IV, 1961) and ends with the same words: ‘the first phase of communism’. The meaning of the word internationalism is explained in the Dictionary of 1980 using an elevated style, with positive and evaluative elements in antagonism with the explanation for the word nationalism, a clear example of lexicographic megalomania according to communist lessons. In this case too, the meaning is partly taken from the Russian Dictionary definition of 1957- 1961.
Meanwhile, the meaning of the word capitalist goes beyond the definition found in the 1957 Russian dictionary, because in the Albanian Dictionary of 1980 capitalists are described as brutally exploiting wage workers and enriching themselves by appropriating surplus value. Likewise, the illustrative examples used the same language, often glorifying socialist morality and the narratives of the Party. Even for words like Marxism and Leninism, the illustration is given in the same way as in Russian dictionaries.
Meanwhile, at the beginning of 2026, the Great Dictionary of the Albanian Language was launched online. In contrast to the 1980s dictionary, it (the compilation of which included the author of this study) represents a modern lexicographic paradigm grounded in comprehensive documentation and descriptive methodology. This shift becomes more apparent when we analyse political or socially sensitive terms. For example, the word capitalism in the 2026 Dictionary is defined in completely neutral language. The same applies to terms such as bourgeois, nationalist, identity, idea, moral, etc.
These changes are evidence of the transition from ideologically prescriptive lexicography toward descriptive and corpus-based lexicographic methodology. At the same time, they reflect the processes of semantic neutralization, de-ideologization and reduction of the positive framing of the word meanings in contemporary Albanian lexicography.
Speaker: Manjola Zaçellari (Aleksander Moisiu University (UAMD))
-
133
-
Morphology, Word Formation & Grammatical Lexicography 00 | Anton Zeilinger Salon (00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor)
00 | Anton Zeilinger Salon
00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Boris Kern (ZRC SAZU & University of Nova Gorica)-
136
Building the Largest Dictionary for an Indigenous Language of North America
This paper describes the development of the New Lakota Dictionary, a corpus-based dictionary of Lakota that, with 48,000 entries, is, to our knowledge, currently the largest dictionary of an Indigenous language of North America. Drawing on more than three decades of lexicographic work, we discuss the methodological and practical challenges of compiling a comprehensive dictionary for a severely endangered, under-resourced, and highly polysynthetic language. Particular attention is given to corpus construction, lexical documentation, collaboration with native-speaker consultants, and editorial decisions concerning lexical categories, argument structure, orthography, and the treatment of derived forms. We also describe the integration of corpus data into both print and electronic editions, lemmatization, valency marking, inflection charts, corpus search, audio recordings, and monolingual Lakota definitions. Beyond documenting the development of a major lexical resource, the paper illustrates how sustained collaboration between native speakers, linguists, and software developers can support both rigorous lexicographic description and long-term language revitalization.
Speakers: Ben Black Bear, Jr. (Lakota Language Consortium), Richard Two Dogs (Lakota Language Consortium), Jan Ullrich (Lakota Language Consortium) -
137
Semantic Class Assignment in a Croatian Verb Lexicon: Challenges of Applying VerbNet’s Classes
This paper examines the challenges of assigning VerbNet semantic classes to Croatian verb senses, based on data from the Croatian verb lexicon Verbion. The lexicon provides a multi-level description of the 500 most frequent verbs in Croatian, integrating lexical, semantic, syntactic, and morphological information. Verb senses are classified into semantic classes according to WordNet and VerbNet, using English equivalents as an intermediary. Since VerbNet is based on Levin's syntactic-semantic classification of English verbs, applying these classes to Croatian presents certain challenges. The analysis identifies five types of cross-linguistic mismatch. First, differences in lexicalization patterns reflect language-specific conventions governing what semantic components are encoded in the verb itself as opposed to what is expressed syntactically. Second, Croatian reflexive morphology frequently introduces semantic distinctions that have no direct counterpart in VerbNet classes. Third, alternation-based classification cannot always be applied, as such alternations are either restricted or encoded differently in Croatian. Fourth, some VerbNet classes do not adequately capture the syntactic and semantic properties of Croatian verbs. Finally, certain high-frequency verbs are absent from the VerbNet inventory altogether. These findings suggest that VerbNet requires systematic adaptation for Croatian. The study contributes to the development of a Croatian-specific extension while preserving VerbNet’s cross-linguistic comparative potential.
Speaker: Ivana Brač (Institute for the Croatian Language) -
138
Where the Boundary between Derivational and Inflectional Morphology Meets Lexicography: Building the DeriVallex Lexicon
The lexicographic treatment of word-formation categories that lie at the boundary between inflection and derivation remains an unresolved issue. In this paper, we introduce the process of compiling DeriVallex, a valency lexicon containing automatically generated valency frames, which provides information on the valency behavior of Czech noun and adjectival deverbal derivatives being halfway between forms and lexemes. These word-formation categories inherit their valency frames and other valency-related phenomena (as, e.g., reflexivity, reciprocity, and other semantic alternations) from their base verbs. We first outline the automatic generation of valency frames of these noun and adjectival derivatives, which resulted in 17,145 valency frames assigned to lexical units, i.e., individual senses of these derivatives, represented by 11,248 noun and adjectival lemmas. The focus of this paper in on the process of structuring of the obtained data into a valency lexicon, providing rich information on valency and derivational properties of the selected nouns and adjectives, enabling human users to search and filter the data. The resulting data may be integrated into both valency lexicons and explanatory dictionaries, thereby increasing their descriptive comprehensiveness.
Speakers: Václava Kettnerová (Charles University, UFAL MFF), Veronika Kolářová (Charles University, UFAL MFF), Jiří Mírovský (Charles University, UFAL MFF)
-
136
-
15:30
Coffee Break 00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna -
Computational & AI-Based Lexicography 02 | Museumszimmer (02 | Museumszimmer, OeAW Main Seat, 2nd floor)
02 | Museumszimmer
02 | Museumszimmer, OeAW Main Seat, 2nd floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Kris Heylen (Dutch Language Institute (INT))-
139
Affordance Norms in Czech and English: Human-LLM Misalignment in Embodied Lexical Knowledge
This paper investigates whether large language models can capture the embodied dimension of word meaning by comparing LLM-generated affordance norms with human-produced norms in Czech and English. Affordance norms are systematically collected data on the actions speakers associate with concrete objects. Human data were collected from 30 Czech native speakers in a free-production experiment using 50 nouns. The same stimuli were presented to four LLMs (Claude Haiku 4.5, Gemini 3 Flash, GPT-5, and Llama 4 Maverick) in both languages, in Czech at two temperature settings; English outputs were additionally compared against human norms from Maxwell et al. (2024). We report four main findings. First, human affordance norms are more variable in Czech. Second, LLMs consistently produce narrower affordance sets than humans, with lower lexical diversity across both languages. Third, lexical overlap between human and LLM affordances is low (shared vocabulary 15–19%; Jaccard index below 0.3 at N = 10) and does not improve substantially with higher temperature or greater language representation in the training data. Fourth, human-only affordances are more embodied and emotionally grounded, while LLM-only affordances cluster around maintenance and support. The findings suggest that human norming data may offer a valuable resource for enriching lexicographic entries.
Speakers: Martina Vokáčová (Charles University), Anna Marklová (Charles University) -
140
Agree to Disagree: Intra-Annotator Agreement in Semi-Automatic Vocabulary Building
Intra-annotator agreement is the rate at which an annotator makes the same choices in the same situation. This paper examines the annotation process in a semi-automatic dictionary-making project, Czech Dictionary Express. In the vocabulary-building process, its annotators went through 100,000 headwords from a corpus frequency wordlist, marking each as incorrect or (partially) correct. The aim of the Dictionary Express projects is a high dictionary-making speed, so the annotators weren’t given the context of the headwords.
In this paper, we examine the shift in annotators’ opinions on headwords before and after context was provided. We argue against providing no context of the headwords, because, unlike the (partially) accepted words, which can still be cut from the vocabulary at later stages, the rejected words cannot be put back in, and this is the case of some correct words that have not been recognised by the annotators without context. We show that the context-based annotation process doesn’t take significantly longer. We also take a brief look at a study about using LLMs for the same annotation process.
Speakers: František Kovařík (Lexical Computing), Marek Blahuš (Lexical Computing), Miloš Jakubíček (Lexical Computing), Vojtěch Kovář (Lexical Computing)
-
139
-
Corpus Linguistics & Corpus-Based Resources 00 | Anton Zeilinger Salon (00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor)
00 | Anton Zeilinger Salon
00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Milica Dinić Marinković (University of Belgrade - Faculty of Philology)-
141
Encoding Emotivity in Dictionary Definitions: A Corpus-Based Study of Verbs of Understanding
The study identifies the emotional profiles of verbs of understanding and demonstrates their relevance for corpus-driven lexicography, contributing to corpus-based research on emotivity and offering implications for lexicographic description. The paper examines the interaction between understanding and emotion as reflected in the contextual behaviour of English verbs of understanding and provides empirical evidence for models of emotionally grounded meaning representation, supported by the emotive clusters most frequently associated with understanding. From a lexicographic perspective, the findings substantiate the inclusion of the emotive component in dictionary descriptions of the verbs in question. Thus, the research argues for a more refined approach to emotivity in lexicographic description of verbs of understanding considering that the emotional meaning makes the axes of the semantics of verbs of understanding.
Speakers: Yelena Yerznkyan (Yerevan State University), Diana Movsisyan (Armenian State University of Economics) -
142
Indexicality in the Vocabulary of German Communication Verbs: Examining Language Use and Speaker Knowledge as a Basis for Lexicographical Conception
Speakers may need to consciously assess which expressions are appropriate for a communicative situation. One such dimension is indexicality: linguistic expressions may evoke the situational contexts in which they are typically employed. In a usage-based understanding, these indexical potentials emerge from recurrent co-occurrence between linguistic forms and situational features and become part of speaker knowledge through entrenchment and schematization (Schmid, 2020). Corpus analysis is methodologically important because natural language frequency patterns can be treated as a model of the input from which speakers may abstract linguistic knowledge (Stefanowitsch & Flach, 2017; Ellis, 2017). However, existing lexicographical resources do not systematically represent usage-based indexicality, and studies that triangulate corpus-derived patterns with speaker knowledge at the vocabulary level are missing. The project introduced in this paper aims to describe systematically and model usage-based indexicality in German communication verbs (CVs) (Harras et al., 2004), drawing on data from both language use and speaker knowledge. Its research approach can provide a basis for conceptualizing the representation of indexicality in lexicographical resources. At the current stage, the paper focuses on the methodological design of the project (first results will be available when the conference takes place).
The project builds on a preliminary corpus study of spoken German (Meißner, 2025), in which CVs were investigated in the FOLK corpus across private, institutional, and public interaction domains (Kaiser, 2018). Analyses of variance and distinctive collexeme analyses (Gries & Stefanowitsch, 2004) were used to identify patterns of over- and under-representation of CV lexemes across domains (indexical potential types (IPTs)). The project asks to what extent such IPTs correspond to speakers’ associations, both with respect to fields of CVs with similar indexical profiles and in terms of granularity, i.e., whether indexical meaning attaches not only to lexemes but also to specific grammatical forms. To this end, three forms are selected: first-person singular present, third-person singular present, and impersonal passive, representing speaker-centered, reporting, and agent-detached perspectives.
To answer the question, corpus-based evidence is triangulated with experimental data on speaker knowledge. The corpus component is based on twelve subcorpora, each containing 100,000 full-verb occurrences, designed to represent contemporary German across private, institutional, and public communication and across spoken, written, computer-mediated, and selected special types of communication. Existing corpus resources are used where possible and are supplemented with new web data where necessary. After preprocessing, analysis of variance and distinctive collexeme analysis are used to derive IPTs for lexemes and grammatical forms.
The speaker-knowledge component mirrors this corpus-based perspective through online questionnaires. For lexemes, 605 CVs from Harras et al. (2004) are distributed over ten item lists and tested in five questionnaire designs: one general association-strength rating, three domain-specific ratings for private, institutional, and public communication, and one free-association task. The rating tasks use labelled seven-point scales; the free-association task asks which communicative situation first comes to mind. This combines scaled judgment data with less pre-structured elicitation of prototypical contexts. A form-level study mirrors these designs for a 20% subset of 121 CVs for the grammatical forms. The rating data are analyzed with ordinal mixed models, with rating type, form, and their interaction as fixed effects and random intercepts for participants and items.
By comparing corpus-derived IPTs with rating and association profiles, the study identifies convergences, divergences, and complementary patterns between language use and speaker knowledge. It refines a usage-based account of indexicality by examining how frequency-based distributional patterns relate to metalinguistic associations. Thereby, the study provides an empirical basis for conceptualizing a representation of situational-indexical information in lexicographical resources for German CVs. More generally, the project outlines a transferable method for modelling how expressions are associated with communicative situations. A lexicographical representation of such associations can support users in choosing forms appropriate to specific communicative settings.
Speakers: Cordula Meißner (University of Innsbruck), Anna-Lena Randermann (University of Innsbruck), Janina Deilke (University of Innsbruck) -
143
Semantic Clustering and Categorisation of Constructional Collexemes Using LLMs
This paper investigates whether large language models can perform semantic clustering and categorisation of constructional collexemes to support the analysis of constructional meaning and the organisation of collexemes within constructicon entries. As a case study, we examine the collexemes of the Estonian Nominal Quantifier Construction, identified from lexicographic and corpus data. Using OpenAI’s ChatGPT-5.4, we conducted an informed free sorting task and a closed sorting task with five input types, consisting either in bare lemmas or lemmas accompanied by different types of context: corpus phrases, corpus sentences, dictionary definitions, and dictionary examples. Model outputs were evaluated against a human-created gold standard using Adjusted Mutual Information, Adjusted Rand Index, and a label quality rating scheme. A scaled closed sorting experiment was also conducted. The results of the free sorting tasks approached human agreement levels, with dictionary definitions yielding the most similar clustering and corpus sentence input the most acceptable labels. The results of the closed sorting task and the scaling experiment demonstrated that a bare list of lemmas provided sufficient input for assigning collexemes to predefined semantic categories.
Speakers: Ene Vainik (Institute of the Estonian Language), Heete Sahkai (Institute of the Estonian Language), Ahto Kiil (Institute of the Estonian Language), Jelena Kallas (Institute of the Estonian Langauge), Geda Paulsen (Institute of the Estonian Language, Uppsala University)
-
141
-
Historical & Diachronic Lexicography 01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)
01 | Johannessaal
01 | Johannessaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: David Lindemann (EHU University of the Basque Country)-
144
A Semantic Web Model for Diachronic Explanatory Combinatorial Dictionaries: Capturing Conceptual Change and Denominational Variation
This paper presents a computational and Semantic Web–based model for Explanatory Combinatorial Dictionaries (ECDs), well-established lexicographic resources grounded in Explanatory Combinatorial Lexicology and Meaning–Text Theory (Mel’čuk et al., 1995). The model aims to make ECD structures machine-actionable, FAIR-compliant (Wilkinson et al., 2016), and interoperable with the Linguistic Linked Open Data ecosystem.
The proposal extends ECDs in two directions: by introducing diachronic depth into dictionaries usually conceived synchronically, and by modelling the phenomenon described in the literature as “denominational paths” (Aymerich et al., 2008), i.e. the ways in which lexical units foreground salient conceptual traits within and across languages.
The proposal is illustrated through a computational diachronic ECD, currently under development, devoted to the terminology of marriage in ancient Israel. Marriage is a suitable test case because it is a polythetic concept whose conceptualisation varies across cultures and historical periods, including within the Jewish tradition (Piccini et al. 2024).
While Biblical sources (c. 500-100 BCE) represent marriage as a legal institution based on the acquisition of a woman and the transfer of jurisdiction from her pater familias to the husband—who acquires sexual monopoly and legal paternity—Talmudic sources (c. 150-500 CE) introduce a more sacralised conception, reflected in terminology foregrounding consecration. This conceptual polyhedrality is mirrored by several denominations, each lexicalising a salient trait: acquisition, as in לקח laqaḥ ‘to take for oneself or for another’; possession, as in בעל ba‘al ‘to possess’; consecration, as in קידושין qiddushin ‘consecrations’; and virilocality, as in נשא nasa’ ‘to lead’, הושיב hoshiv ‘to cause to dwell’, and שלח shillaḥ ‘to send’. The modelling of nasa’ provides the illustrative example for the RDF/OWL-based representation presented in Figure 1*.
The model combines a lexical and an ontological layer. The lexical layer relies on three vocabularies: OntoLex-Lemon for lexical units and senses, Lexicog for dictionary-like structure, and LexFom (Fonseca et al., 2016) for Lexical Functions. In Figure 1(b), nasa’ is represented as an ECD entry specifying its canonical form, lexical sense and combinatorial information.
The ontological layer models diachronic conceptual change through a perdurantist, or 4D, approach (Welty et al., 2006), in which change is represented through temporal slices rather than by assigning time-dependent properties to the concept as a whole. Figure 1(a) presents the modelling of Hebrew marriage (HEBREW_MARRIAGE) between 500 and 100 BCE. This temporally bounded phase is represented by the class SLICE_1, defined as a subclass of HEBREW_MARRIAGE and associated with a specific time interval. SLICE_1 is characterised by traits such as legal paternity, male sexual monopoly and virilocal residence, each encoded as an existential restriction on a specific property. At the same time, it inherits more general, time-invariant traits from MARRIAGE, such as the establishment of a contractual bond and a socially regulated relation between sexes. Anchoring ECD entries to such slices the ontoleox:reference property turns the resource into a diachronic ECD.
The second objective, i.e. capturing the selective lexicalisation of salient conceptual traits, is implemented at the lexical–conceptual interface through the property ontolex:isLexicalizedSenseOf, which links a lexical sense (in our case nasa’ lexical sense) to the conceptual trait it lexicalises (virilocality). Since this relation requires an ontolex:LexicalConcept as its object, conceptual traits are represented as lexical concepts within a feature-based schema, consisting of a skos:ConceptScheme and a set of features modelled as SKOS concepts (also typed as ontolex:LexicalConcept).
Such a formalisation supports fine-grained semantic and diachronic queries, from identifying the conceptual facet lexicalised by a term to tracing specific traits across historical phases (Figure 2*).
* Figures see Book of Abstracts
Speakers: Silvia Piccini (Istituto di Linguistica Computazionale “A. Zampolli”), Andrea Bellandi (Istituto di Linguistica Computazionale “A. Zampolli”), Giuliana Elizabeth Vilela Ruiz (Istituto di Linguistica Computazionale “A. Zampolli”), Davide Saponaro (Fondazione Rut) -
145
DiaLexPoL: LLM-assisted Sense Assignment for a Diachronic Dictionary of Latin
The DiaLexPoL project is building a diachronic lexical database of Neo-Latin (ca. 1550–1800) from the semi-automatically compiled diachronic corpus of Polish Latin (DiaCorPoL). Its principal bottleneck is sense assignment: referring early-modern attestations to an inventory designed for the Dictionary of Medieval Latin in Poland (DMLP). We ask whether deployable LLMs – cheap commercial mid-tier and self-hostable open-weight models, rather than frontier systems a chronically underfunded project cannot reproduce – can draft this step in a human-in-the-loop workflow. Two experiments address it: reproducing DMLP's own sense assignments on its quotations, and assigning Neo-Latin concordance lines to the DMLP inventory against independent human gold-standard annotation. Self-hostable open-weight models give the strongest drafts, outperforming the single commercial mid-tier model at negligible cost, and the best corpus configuration matches human inter-annotator agreement overall, though not on every lemma. Gains over a most-frequent-sense baseline – always assigning a lemma's commonest sense – are modest and concentrate on genuinely polysemous lemmas, where editors most need help, while on dominant-sense lemmas the models over-split, marking where LLM drafting helps.
Speakers: Krzysztof Nowak (Institute of Polish Language PAS), Iwona Krawczyk (Institute of Polish Language PAS), Jagoda Marszałek (Institute of Polish Language PAS) -
146
Tracking Eight Centuries: The Czech Monitor Corpus as a Resource for Diachronic Lexical Research
This paper demonstrates the central role of richly annotated diachronic corpora in investigating historical lexical change, using the newly compiled Monitor Corpus of Czech as a case study. The corpus covers over eight centuries of the Czech written tradition and is designed as a general-purpose, consistently annotated resource for diachronic research. Unlike earlier diachronic corpora of Czech, the Monitor Corpus aims at systematic genre balance (where historical conditions permit) and at comparability of annotation across periods. To make the corpus maximally exploitable, it is accompanied by a dedicated application, Timeline, which supports queries and visualizations tailored to diachronic analysis.
The design of the Monitor Corpus seeks to represent, for each period, the broadest possible range of text types at the most general level: fiction, non-fiction, and journalism. Of course, this ideal cannot be met uniformly throughout the history of Czech. For some periods, no specialist texts have survived in Czech and journalism emerges only from the 18th century onwards. Nevertheless, in those historical stages where it is feasible, the corpus includes roughly comparable proportions of fiction, non-fiction, and journalistic writing, thereby enabling meaningful comparisons of lexical developments across time and genre.
For lexicographical and broader linguistic research, uniform, transparent and, above all, consistent annotation over time is essential. A key methodological challenge is how to handle both graphical and morphological variability so that it becomes possible to track lexical items across centuries. For example, the adjective estetický ‘aesthetic’ is attested in the 19th-century material in numerous spelling variants (e.g. aestetického, Esthetickému, aesthetickou). These forms must be recognized as instances of a single lemma in order to study its history in a principled way. To address this, we compiled three manually annotated etalon corpora representing different historical stages of Czech and used them as training data for a lemmatizer and POS/morphological tagger developed within the Universal Dependencies (UD) framework (de Marneffe et al., 2021; Zeman et al., 2021). The result is a diachronic corpus that is lemmatized and morphologically annotated according to a single, coherent scheme.
The benefits of this annotation become especially clear when investigating developmental tendencies that would be difficult or impossible to access using only raw text. By systematically linking historically and graphically divergent word forms under a single lemma and assigning each token a detailed morphological description, we can reliably estimate frequencies, identify collocational patterns, and detect changes in distribution across grammatical categories and constructions. The paper will describe how historical variants were handled in the annotation workflow, and how the resulting data can be exploited to study lexical change.
As concrete case studies illustrating how the Monitor Corpus and Timeline application can be used in practice, we focus on examples of nouns that have undergone semantic change or diachronic shifts in usage, such as nápad ‘attack’ / ‘idea’, puška ‘container’ / ‘rifle’, and národ ‘family’ / ‘nation’. For example, the word národ has undergone a complex semantic evolution since the earliest periods of Czech: attested meanings include ‘fruit’, ‘family/kin’, ‘pagans’, and progressively the now dominant sense ‘community of people sharing a territory or language’. Its sociopolitical salience in the 19th century is reflected in a marked increase in frequency in that period. The annotated corpus allows us to examine not only this overall frequency development, but also changes in the distribution of grammatical forms and constructions, and to relate them to semantic and discursive shifts.
In Old Czech, plural forms are strongly predominant, and one important meaning of the time—‘pagans’—is expressed exclusively in the plural (národové ‘nations’). From the early 19th century onward we observe a significant rise in genitive forms, and the lexical content of národ becomes increasingly abstract and context-dependent. The heads of noun phrases in which národ appears make an important contribution to its discursive interpretation. In the 19th century, typical genitive constructions include duch národa ‘the spirit of the nation’, and rozkvět národa ‘the flourishing of the nation’, reflecting nationalist and Romantic discourses. By contrast, in 21st-century texts, constructions such as společenství národů ‘community/commonwealth of nations’ and organizace národů ‘organization of nations’ become prominent, mirroring a shift towards international and institutional frames.
Using a small set of nouns as case studies, the paper will show how lemmatization enables the automatic extraction of statistically robust lemma-based collocations across periods and genres, and how morphological annotation supports the identification of typical semantic–syntactic patterns. Both lines of research use the Timeline application, which allows various frequency data and collocation profiles to be extracted and visualized on a timeline. The Monitor corpus and the Timeline application are currently in beta and are expected to be made publicly available by the end of 2026.
Speakers: Václav Cvrček (Charles University), Martin Stluka (Charles University), Klára Pivoňková (Charles University)
-
144
-
Lexical Semantics, Neology & Phraseology 01 | Sitzungssaal (01 | Sitzungssaal, OeAW Main Seat, 1st floor)
01 | Sitzungssaal
01 | Sitzungssaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Tanara Zingano Kuhn (CELGA-ILTEC, University of Coimbra)-
147
Emotions in Context: A Semantic Field Viewed through Grammar and Collocations
We examine Czech emotion nouns as a structured semantic field, adopting a field-level perspective that complements entry-level lexicographic description. The analysis is based on 65 nouns denoting emotions and related affective states, selected with reference to psychological classifications and corpus frequency. It combines grammatical profiling, correspondence analysis, and a brief collocational overview in order to show that these lexemes share certain patterns, some of them characteristic of the field as a whole and others pointing to its internal structuring. The results suggest that grammatical behaviour does not simply reflect semantic similarity in the usual psychological sense, but reveals different ways in which emotions are represented in language, for example as experienced states, emotional contents, agent-like forces, or states one is driven into. Collocational evidence points to broader tendencies across the field, including loss of agency, emotional intensity, bodily manifestation, regulation, and combination with other emotions. We argue that this kind of field-based description can complement traditional lexicographic practice and may also be attractive to dictionary users interested in contrasts between related words.
Speakers: Dominika Kováříková (Charles University), Laura A. Janda (UiT The Arctic University of Norway) -
148
GPT as a Sense-Tagger for Selecting Dictionary Examples
Aligning a target word in a sentence with a dictionary sense is a challenging task, given the dynamicity of meaning (in context) and the fixed discrete senses in dictionaries, among other reasons. However, it is still a practical necessity for a lexicographer to select example sentences that best illustrate dictionary senses and, hence, find a way to address this challenge. Further challenges appeared after the established use of large corpora as sources of lexicographic evidence and the need to update dictionary entries with additional illustrative examples. The present study tested the usability of an LLM as a facilitator in sense-aligning and example-selecting tasks. The role of the model was limited to assigning a score that reflects the match between (corpus-based and dictionary-cited) examples and existing dictionary senses under several experimental conditions. Quantitative and qualitative analyses showed the usefulness of the model in selecting candidate senses that best represent a dictionary sense, detecting misalignment cases in dictionaries, excluding less representative examples despite their high GDEX scores and identifying highly overlapping senses (especially in the category of adjectives) that are likely to be puzzling to the users.
Speakers: Ágoston Tóth (University of Debrecen), Esra Abdelzaher (University of Debrecen) -
149
Exploring Large Language Models in Word Sense Disambiguation: A Study of Corpus Examples for Perception Verbs in Learners' Dictionaries
Recent developments in Large Language Models (LLMs) have opened new possibilities for supporting lexicographic work. Studies suggest that LLMs can assist in tasks such as generating definitions, drafting example sentences, and assigning examples to dictionary senses (Lew, 2023; Jakubíček and Rundell, 2023). At the same time, the status of word senses remains one of the most debated issues in lexicography, with scholars highlighting the fluidity and context-dependence of meaning. Traditional lexicographic practice involves abstracting discrete senses from potentially unlimited numbers of contextualised corpus attestations. This process is particularly challenging in the case of highly frequent and polysemous lexical items, where meanings often overlap and form complex semantic networks. Dictionaries may adopt different approaches to sense division, “lumping” or “splitting” senses, which leads to considerable variation in sense inventories (Bond et al., 2024).
Nowadays, apart from assigning corpus-based examples to particular senses, online learners’ dictionaries increasingly include large collections of automatically extracted corpus examples that are not explicitly linked to individual senses. While these collections provide valuable data, their pedagogical usefulness remains limited, as examples are not systematically organised. Recent advances in generative AI suggest that LLMs may help bridge this gap by analysing contextual meaning and aligning examples with sense definitions. Research indicates that such models are capable of handling metaphor, semantic relatedness, and graded meaning variation (Lin et al., 2024; Bond et al., 2024). Nevertheless, their performance in dealing with fine-grained sense distinctions and competing lexicographic models of word sense disambiguation remains underexplored.
The present study examines the potential of LLMs as tools for analysing sense granularity taking a cross-dictionary perspective. This paper investigates the ability of ChatGPT (GPT-5.2) to interpret corpus examples of English perception verbs and align them with the existing sense inventories. Both sense divisions and sections of unassigned corpus examples were drawn from two major online monolingual learners’ dictionaries: Cambridge Advanced Learner’s Dictionary and Longman Dictionary of Contemporary English. As the number of unassigned corpus examples for each examined entry differs between the dictionaries, the first twenty examples per verb were analysed. The dataset includes eleven high-frequency and highly polysemous perception verbs: see, hear, touch, taste, smell, look, watch, observe, notice, listen and feel. These verbs display systematic extensions from the physical domain into cognitive and evaluative domains, hence are suitable for examining sense granularity and semantic overlap. Given their systematic polysemy, the study examines how ChatGPT handles semantic shifts from physical to abstract and metaphorical meanings.
Speaker: Sylwia Wojciechowska (Adam Mickiewicz University, Poznań)
-
147
-
150
Conference Dinner Fuhrgassl-Huber
Fuhrgassl-Huber
Neustift am Walde 68 1190 WienOnly for registered participants.
If you want to join us for the Dinner and haven't registered, please get in touch with the local organisers.
-
-
-
Computational & AI-Based Lexicography 01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)
01 | Johannessaal
01 | Johannessaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Miloš Jakubíček (Lexical Computing)-
151
Are Bears More Offensive than Dogs When Drunk? Using LLMs to Systematize the Labelling of Estonian Synonyms
This paper investigates the potential of large language models (LLMs) as assistants with corpus analysis in lexicography, focusing on the assignment of register labels in the EKI Combined Dictionary (CombiDic). Inconsistencies in register labels become evident across synonym sets in CombiDic, which is the result of different lexicographers working on words at different times. We examine how LLMs from Anthropic, Google, and OpenAI handle the challenge of distinguishing between offensive and colloquial usage, which pose a challenge in lexicographic practice. Using a evaluation dataset of 297 words evaluated by five native Estonian-speaking annotators and three LLMs queried via API, we find that LLMs are quite stable tools for register categorisation and that corpus context significantly improves their performance. Depending on the LLM, 79.8–88.6% of suggestions were deemed acceptable for lexicographic use, although since LLMs may be trained to detect offensive language, they tended to select the OFFENSIVE label more often than human annotators. Nevertheless, Gemini 3.1 Pro performed best overall. The experiment already enabled corrections to CombiDic, demonstrating practical value. While full systematisation has not yet been achieved – the broader goal is labels grounded in usage data – LLMs represent a promising and time-efficient aid in semi-automatic lexicographic workflows.
Speakers: Lydia Risberg (Estonian Language Institute, University of Tartu), Kristina Koppel (Estonian Language Institute), Margit Langemets (Estonian Language Institute), Hanna Maask (Estonian Language Institute), Esta Prangel (Estonian Language Institute), Maria Tuulik (Estonian Language Institute), Silver Vapper (Estonian Language Institute) -
152
The Integration of British and American Slang in Contemporary Polish and the Role of LLM-Driven Dynamic Lexicography
This paper investigates the accelerating influx of English slang into the Polish language, driven by digital globalization and social media usage. While historical Anglicisms often focused on technical or professional domains, modern borrowing increasingly penetrates the affective realm—informal expressions of identity, emotion, and subcultural belonging. Central to this exploratory research is the evaluation of Large Language Model (LLM) assistance as a supplementary tool in dynamic lexicography, where manual editorial processes traditionally struggle to capture short-lived online neologisms. This paper proposes a framework utilizing LLMs for initial candidate extraction and semantic mapping. Using a small-scale pilot survey of Polish speakers, we evaluate LLM outputs, discuss the systemic risks of automated lexicography (such as hallucinated forms and semantic overinterpretation), and outline the vital role of human validation in tracking modern slang.
Speaker: Barbara Lewandowska-Tomaszczyk (University of Applied Sciences in Konin) -
153
Reflections on the Semantic Understanding of Generative AI: From LLM Experiments to Lexicographical Tools
This article presents an application evaluation of the use of large language models (LLMs) for the semantic classification of lexicographical content in three German dialect dictionary projects. Building on earlier experiments with LLM-based semantic classification that showed hit rates of over 80%, the study investigates whether generative AI can support the assignment of meanings to controlled semantic categories while preserving editorial oversight. Evaluation data from user ratings, surveys, and interviews indicate that the tool produces useful classifications in a majority of cases, reduces manual research effort, and contributes to greater consistency in semantic annotation. At the same time, the results reveal recurring limitations, especially in cases involving semantic granularity, context-dependent meanings, and definitions where modifiers are overemphasized. The paper argues that LLMs are most valuable not as replacements for expert lexicographical work, but as components of hybrid workflows that combine automated suggestion with human validation.
Speakers: Ines Röhrer (Bavarian Academy of Sciences and Humanities), Manuel Raaf (Bavarian Academy of Sciences and Humanities) -
154
Introducing Human-Centeredness in AI-Assisted Lexicography
This paper proposes a human-centered artificial intelligence (HCAI) framework for AI-assisted lexicography. While generative AI offers significant opportunities to enhance lexicographic work, it also raises concerns regarding the future role of lexicographers and the preservation of linguistic and cultural diversity. Drawing on HCAI principles and previous applications in other language professions, the paper identifies four interrelated dimensions through which AI integration in lexicography can be understood and critically examined: the augmented lexicographer, the sociotechnical context of AI integration, bias, and the design of AI-powered lexicographic tools. The framework argues that AI should augment rather than replace lexicographers, combining automation with meaningful human control. It further emphasizes the importance of preserving professional agency, mitigating AI-generated biases, and designing tools around the needs of lexicographers. By doing so, the paper provides a foundation for future research and the beneficial integration of AI into lexicographic workflows.
Speakers: Antonio San Martín (University of Quebec in Trois-Rivières), Catherine Trekker (University of Quebec in Trois-Rivières)
-
151
-
Learner's Lexicography & Dictionary Use 00 | Anton Zeilinger Salon (00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor)
00 | Anton Zeilinger Salon
00 | Anton Zeilinger Salon, OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Orin Hargraves (University of Colorado Boulder)-
155
Spoilt for Choice? ChatGPT Versus a Bilingual Dictionary in Selecting L2 Equivalents for Polysemous L1 Words
A common challenge for language learners is finding appropriate L2 equivalents for L1 words. The difficulty is compounded when the L1 word is polysemous, because different senses typically call for different equivalents. When selecting an L2 equivalent for a polysemous L1 word, learners therefore face a twofold task: identifying the intended sense of the L1 word and choosing an L2 equivalent that fits the context.
Traditionally, learners have dealt with this problem by turning to bilingual dictionaries. High-quality bilingual dictionaries may provide glosses, labels, collocations, examples, and other information that helps users select the most appropriate equivalent from among those listed and then use it in context. Yet, as previous studies have shown, the mere availability of information does not guarantee that users will find it or use it effectively (Lew & Tokarek, 2010; Dziemianko, 2012).
With the arrival of ChatGPT and other large language models (LLMs), learners may now turn to AI tools instead of consulting dictionaries. A recent study by Ptasznik and Lew (2025) indicates that at least some learners already do so, although their preferences appear to be task-dependent. Learners’ willingness to use such tools is perhaps unsurprising, given evidence from user studies showing that LLMs can match and sometimes outperform dictionaries in receptive and productive tasks (Rees & Lew, 2023; Ptasznik, Wolfer & Lew, 2024). Yet expert evaluations of entries generated by LLMs have been mixed. An often-noted problem is false polysemy, which occurs when “the system enumerates multiple senses, with different definitions, in cases where there is really only one” (Jakubíček & Rundell, 2023, p. 525). A recent user study found that false polysemy can have a detrimental effect on dictionary users, increasing the time needed to select a definition and lowering users’ confidence in the definition chosen (Michta & Frankenberg-Garcia, 2025).
To the best of our knowledge, no study to date has specifically examined learners’ performance in selecting L2 equivalents for polysemous L1 words when consulting a bilingual dictionary versus an LLM, a task that is common yet under-researched. We investigate learners’ performance using the following measures: response accuracy, response time, and delayed recall of the target equivalents one week later. In addition, we examine learners’ confidence in the equivalents they selected and changes in their tool preferences after the task.
A total of 105 L1 Polish speakers, all enrolled in an English philology programme, took part in the study. Before completing the task, participants indicated which of the two tools, ChatGPT or the online Polish-English PONS dictionary, they would prefer to use for the task and rated how useful they expected each tool to be. They were then randomly assigned to one of the two conditions. Participants completed an online task involving 20 sentences presented in random order. Each item consisted of an English sentence with a gap and a Polish polysemous word in parentheses. Participants first indicated whether they were able to supply an English equivalent without consulting the assigned tool, and if so, provided it. Next, they consulted the assigned tool and filled in the gap using the information obtained during consultation; response time was recorded. After each response, they rated how confident they were that they had selected an appropriate equivalent. One week later, 70 participants completed a post-test in which the same Polish words and senses were presented to them in new contexts. For each item, they provided an English equivalent and rated their confidence. Participants also rated the overall helpfulness of the assigned tool after both the main task and the post-test.
The results suggest that the dictionary group provided more accurate responses, required less time to supply an equivalent, and reported higher confidence in their choices. After the task, participants more often selected the dictionary than ChatGPT as the tool they would choose for a similar task in the future. Full analyses will be presented at the conference.
Speakers: Tomasz Michta (University of Bialystok), Katarzyna Mroczyńska (University of Siedlce) -
156
Learner’s Dictionaries Based on Etymological Principles: The Example of a German Textbook Dictionary for Swedish Learners
Today there are numerous frequency dictionaries marketed as learner’s dictionaries, with various source languages and mostly with English as target language. They typically contain 5000 lemmas, which happens to coincide with a rule of thumb (though without much academic substance) that this number of words provides the lexical text coverage (95-98%) necessary for an adequate text comprehension.
For a learner of a foreign language it must be a daunting task to learn so many words in frequency or alphabetical order, or rearranged in thematic groups, as is often the case for learner’s dictionaries. From a receptive perspective, it is however not necessary to learn all of them as vocabulary items, thanks to intralingual and interlingual facilitating factors.
Intralingual aspects: In a compounding language like German, complex words are more or less transparent and can be interpreted from their component affixes and root morphemes. The concept of “word family” in frequency studies usually means words in different parts of speech transparently derived from each other, for example information, inform, informative. Word family can also denote groups of words with a deeper morphological relationship, such as presented in the German “Wortfamilienwörterbuch” by August (2009), for instance Mund: Mundwinkel, Mundart, mundfaul, mündlich, münden, Mündung, ausmünden, einmünden, etc. To distinguish these deeper word formation families from the more superficial word family, the term used here is “root family” when etymologically related words are grouped around a primary word. A “primary word” is a word with one lexical root and at the same time the most central in its word and root family, such as Mund/mouth.
Interlingual aspects: Different language pairs, such as French–English or English–Finnish, have a varying proportion of cognates, understood here as any two words with an etymological link. A structural analysis of all the primary words in a German frequency dictionary (Tschirner & Möhring 2020) shows that over 90% of them has a distinguishable Swedish cognate. A study conducted with adult beginning learners of German at Stockholm University (Winnerlöv 2013) suggests that 70% of the primary content words can be understood thanks to their Swedish (or sometimes English) cognate. The rest of the cognates follow certain patterns and could be deciphered with some instruction, and when the patterns are less clear, they could at least serve as mnemonic help.
The new etymology-based German–Swedish learner’s dictionary proposed here incorporates the above aspects and covers the majority of the words from the analysed frequency dictionary as well as some other words. During the presentation samples will be shown (where relevant with English rather than Swedish translations) from its main chapters, consisting of 1) some full cognates for pronunciation practices 2) cognate pairs according to predictable patterns 3) other more diverging cognates 4) non-cognate primary words with 1-1 translation equivalents 5) analytical articles for false friends, polysemic words, and other words that need comments 6) function words in paradigms 7) other words learnt in semantic paradigms or series, and finally 8) root family groups. Since the dictionary is meant to be read in its entirety, and reviewed repeatedly, it is called a “textbook dictionary”.
The dictionary is conceived to serve as a model for learner’s monodirectional textbook dictionaries for other language pairs. The relative size of the intralingual versus the interlingual parts will depend on the structural distance between the languages in question, and also on the extent to which other commonly known languages such as English come into play. Common for all possible language versions is the design to reduce the relevant word population to a minimum number of primary words that need to be memorized as traditional vocabulary items, while relying on other cognitive strategies for grasping the meaning of the remaining words.
While the psycholinguistic aspects of cognates are well researched, there are still few scientific studies on their practical use in language pedagogy, which may reflect the traditional reluctance of language instructors to make the etymological dimension explicit, especially for younger and less advanced learners. Yet in a vocabulary learning experiment conducted with US university students of German, Stratton (2022) found that instruction in historical linguistics was helpful for learning German–English cognates that were not readily recognizable at first sight.
The prospective intralingual method with root families, though commonly used by expert language learners, appears to be unexplored in the scientific literature. The proposed textbook dictionary, in addition to being an immediate help for language students who want to expand their receptive lexical knowledge rapidly, could serve as a foundation for research into the effectiveness of explicit etymology – both interlingual and intralingual – as a pedagogical tool in second language learning.
Speaker: Jonas Winnerlöv (European Parliament) -
157
Ostensive Aids in Learner’s Dictionaries: A Diachronic, Comparative Analysis
This paper proposes a diachronic analysis of the use of ostensive aids in 35 print editions and the current online versions of six major English monolingual learners’ dictionaries (Oxford, Longman, Collins, Cambridge, Macmillan, Merriam‑Webster). Alongside macrostructural features of illustrations (density, types, placement, labeling), I analyze their content, performance, and consistency with respect to three questions: (1) whether principles for selecting lemmas to illustrate are stated or inferable; (2) how polysemy and semantic fields are represented; and (3) how culture‑specific concepts are handled. The findings show a persistent lack of explicit, operational selection criteria or principles (with LDOCE2 as a partial exception), a substantial reduction in the use of illustrations in the microstructure in print editions, and an unsystematic implementation of ostensive aids in online editions despite the absence of space constraints. The treatment of polysemy, prototypes, and culture‑bound items remains uneven. The present analysis is a preliminary step toward outlining a theoretically grounded methodology for the implementation of ostensive aids in the Phrase-based Active Dictionary (see Giacomini, 2025; DiMuccio-Failla, 2025).
Speaker: Laura Rebosio (University of Innsbruck) -
158
ChatGPT or a Dictionary? Resolving Polysemy in L2 Reception
Recent empirical studies suggest that ChatGPT can assist users with a range of receptive and productive language tasks (Rees & Lew, 2023; Ptasznik, Wolfer & Lew, 2024). However, users consult dictionaries in many situations not covered in previous studies comparing ChatGPT with dictionaries, and ChatGPT’s usefulness is likely to vary depending on the task. This is plausible because ChatGPT does not appear to perform equally well when asked to produce different types of dictionary information: it has been found to produce definitions that are “practically indistinguishable” from those found in the COBUILD dictionary (Lew, 2023, p. 8), while ChatGPT-generated examples tend to be repetitive, unimaginative and unnatural (Jakubíček & Rundell, 2023; Lew, 2023). The present study focuses on a situation often faced by language learners but still underexplored in empirical comparisons of ChatGPT and dictionaries: resolving polysemy in L2 reception.
Polysemous L2 words pose a considerable challenge for language learners because they can be deceptively familiar: learners may assume they understand a word because they know one of its senses, although the context requires another. If they do consult a dictionary, they still have to determine which sense in the entry matches the context. As earlier studies have shown, many learners struggle to do so (Dziemianko, 2019; Kamiński, 2025).
Learners may now also use an LLM to help them understand which sense of a polysemous word is intended in context. LLMs may be useful here, but they may also present learners with false polysemy by proposing senses that are not clearly distinct from one another (Jakubíček & Rundell, 2023). This matters because false polysemy has been shown to make sense selection more time-consuming and to lower users’ confidence in the sense they selected (Michta & Frankenberg-Garcia, 2025).
In this study, we examine how successfully learners identify the contextually appropriate senses of polysemous L2 items. To this end, we compare the effectiveness of ChatGPT with that of a bilingual dictionary in an equivalent-selection task and with that of a monolingual learner’s dictionary in a definition-selection task. In addition to accuracy, we investigate learners’ confidence in their choices, time-on-task, and order effects.
The participants were 131 L1 Polish secondary-school students whose proficiency in English ranged from B1 to B2. Fifteen polysemous English nouns were selected as target items. Each item appeared in a single sentence taken from Sketch Engine for Language Learning, published books, or dictionaries other than those used by participants in the study. Some sentences were minimally edited to remove non-essential information. The task was administered via a worksheet in which sentences were presented in random order.
The study consisted of a pretest and a main test. In the pretest, all participants saw 15 sentences and attempted to translate or paraphrase each target item without tool support. In the main test, participants were assigned to one of four groups and used their phones to complete one of two tasks. In the equivalent-selection task, Group A used Diki.pl, an online English-Polish dictionary, while Group B used ChatGPT. Participants selected a Polish equivalent of the target word and rated their confidence in selecting an appropriate equivalent on a 1–5 Likert scale. In the definition-selection task, Group C used the online Oxford Advanced Learner’s Dictionary, while Group D used ChatGPT. Participants indicated their selected definition by copying three words from it and rated their confidence in their choice on a 1–5 Likert scale. They were then asked to provide a Polish equivalent of the English target word. Total time-on-task was recorded for the main test.
The data were analysed separately for each of the two tasks using logistic mixed-effects models for accuracy and ordinal mixed-effects models for confidence. Preliminary results suggest that, in the equivalent-selection task, Diki.pl users outperformed ChatGPT users and reported higher confidence. Order effects were also observed in this task. Full analyses will be presented at the conference.
Speakers: Tomasz Michta (University of Bialystok), Tomasz Czyż (University of Bialystok)
-
155
-
Lexical Semantics, Neology & Phraseology 01 | Sitzungssaal (01 | Sitzungssaal, OeAW Main Seat, 1st floor)
01 | Sitzungssaal
01 | Sitzungssaal, OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Tinatin Margalitadze (Ilia State University)-
159
Building a Sense Inventory for the Central Lexicographic Infrastructure of Dutch
Like other lexicographic and language institutes in Europe, the Dutch Language Institute (Instituut voor de Nederlandse Taal, INT) is progressively integrating the compilation and management of its individual lexicographic products into a single shared lexicographic infrastructure with the aim of increasing internal efficiency and enable new forms of use, research and development (Depuydt, Tiberius & Heylen 2026). Comparable infrastructural trajectories can be observed for Estonian (Koppel et al. 2019), Danish (Pedersen et al. 2018), or Slovene (Gantar 2020). For the INT, this integration is particularly challenging because its historical and contemporary dictionaries were originally largely developed as stand-alone projects and differ substantially in their internal structure.
Over the past two decades, the INT has systematically invested in addressing this challenge by developing a shared infrastructural backbone for Dutch lexicography. Since 2007, this has resulted in the creation of GiGaNT (Groot Geïntegreerd Lexicon van de Nederlandse Taal), a centralized lexicon providing unified lemmatisation principles, part-of-speech tagging, and form-related information for both historical and contemporary Dutch dictionaries and lexical databases (Ruitenberg, de Does, Depuydt 2010). GiGaNT enables the alignment of lexicographic resources at the lemma level across time periods and projects, and already supports the reuse of formal and morphological information in multiple lexicographic end products.
The present paper focuses on the next infrastructural step: the development of a sense inventory that extends this integration from the lemma level to the sense level so as to share and combine information that is currently distributed across databases. However, these resources, including the historical Woordenboek der Nederlandsche Taal (WNT), the contemporary general Algemeen Nederlands Woordenboek (ANW), and the phraseological Woordcombinaties differ not only in sense granularity, but also in their organising principles and descriptive focus. Therefore, a sense inventory is introduced as an intermediate semantic layer that allows these resources to be linked through core senses, to which definitions, corpus attestations, collocational patterns, usage information, and sense relations can be attached.
The objectives of this infrastructural development are twofold. Internally, the sense inventory is intended to support more efficient lexicographic workflows, enable systematic consistency checking across resources, and facilitate the reuse of semantic information in multiple dictionary products and services. Externally, it forms part of a publicly accessible lexicographic infrastructure that can be used for research on the Dutch lexicon, for the development of language learning materials, and for language-technology applications. In addition, the Dutch sense inventory is explicitly designed to contribute to and represent Dutch within the European lexicographic infrastructure ELEXIS, which will be substantially upgraded in the recently started Horizon Europe infrastructure project ELEXAI (2026-2029). Within ELEXAI, a collaborative and dynamic sense inventory will play a central role in the further development of the Dictionary Matrix and the European lexicographic knowledge graph.
The design of the Dutch sense inventory builds on earlier work at the INT, most notably the DiaMaNT project (Depuydt & de Does 2018), which established a semantic lexicon for historical Dutch dictionaries by linking definitions across resources at a schematic level. This experience provides both methodological insights and reusable data for extending semantic integration to contemporary lexicography. At the same time, the current work takes into account international models and standards for lexicographic data modelling, including the DMLex data model (DMLex-1.0) and earlier proposals for unified sense-level organisation (e.g. Tavast et al. 2018).
The design of the Dutch sense inventory builds on earlier work at the INT, most notably the DiaMaNT project (Depuydt & de Does 2018), which established a semantic lexicon for historical Dutch dictionaries at a schematic level across resources by levering synonym definitions. This experience provides both methodological insights and reusable data for extending semantic integration to contemporary lexicography. At the same time, the current work takes into account international models and standards for lexicographic data modelling, including the DMLex data model (DMLex-1.0) and earlier proposals for unified sense-level organisation (e.g. Tavast et al. 2018).
This paper presents four interrelated components of the ongoing work. First, we introduce a preliminary data model for the Dutch sense inventory, which builds on the existing GiGaNT model while incorporating sense-level components already present in individual dictionaries. The model is designed to accommodate heterogeneous sense structures and to remain compatible with European interoperability efforts. Second, we discuss the principles for sense linking and sense lumping into core senses. These principles explicitly address large differences in sense granularity between dictionaries, the specific organising principles of Woordcombinaties, and the availability of explicit sense relations in the ANW, with attention to recurrent phenomena such as regular polysemy, for which earlier proposed (e.g. Pedersen et al. 2022) are evaluated and adapted.Third, we describe the compilation of a test dataset that serves both as a proof of concept and as an evaluation benchmark. The dataset includes lemmata with different parts of speech, varying degrees and types of polysemy, and contrasting treatments in the historical WNT, the contemporary ANW, and the pattern-based Woordcombinaties. It builds on previously curated material from DiaMaNT and on the Dutch component of the multilingual evaluation dataset for monolingual word sense alignment (Ahmadi et al. 2022). Finally, we discuss how this dataset will function as a baseline for future experiments in automatic sense linking, lumping and splitting using both open-weights models and commercial large language models.
Speakers: Kris Heylen (Dutch Language Institute (INT)), Katrien Depuydt (Dutch Language Institute (INT)), Carole Tiberius (Dutch Language Institute (INT)), Jesse de Does (Dutch Language Institute (INT)) -
160
New Senses, New Methods: Clustering as a Guide for Detecting Semantic Neologisms
Identifying semantic neology remains a significant challenge for lexicography due to the gradual nature of meaning shifts. This study proposes a methodological framework for the automatic detection of verbal semantic neology in Catalan, leveraging recent advances in Natural Language Processing (NLP). Utilizing a general-language corpus spanning 2000-2021, this research employs a Large Language Model to generate contextualized embeddings for 90 target verbs, capturing nuanced semantic variation. The HDBSCAN clustering algorithm is then applied to group these representations by usage similarity. To validate the results, a lexicographic baseline is established using the Diccionari essencial de la llengua catalana; clusters that do not align with registered senses are flagged as candidate neologisms for qualitative analysis. This qualitative analysis focuses on shifts in argument structure and semantic implicatures. The findings emphasize the value of NLP tools for low resource languages like Catalan that have growing but comparatively few digital resources. Ultimately, this study advocates for a hybrid lexicographic model where machine learning serves as an exploratory tool to augment, rather than replace, human expertise. This approach ensures that modern dictionaries can responsibly integrate technological innovation while maintaining academic and ethical standards.
Speaker: Marta Garcia-Casado (Pompeu Fabra University) -
161
Sense Classification For Word Sketches
The Word Sketch is a core tool for modern corpus-based lexicography, providing a concise summary of a word's grammatical and collocational behavior. A longstanding limitation of standard Word Sketches is the lack of polysemy awareness: when a lexicographer examines a highly polysemous word such as bank, the resulting collocate lists mix financial terms like account and deposit with geographical terms like river and sand. In this paper, we present and evaluate a fully automated method for enriching Word Sketches with word sense information. Using the Adaptive Skip-gram model, we perform unsupervised word sense induction and project the induced senses onto the type-level collocate lists of a Word Sketch, so that collocates with sufficiently confident predictions are associated with the sense they most strongly represent. On a manually annotated test set of 23 polysemous English words, our method achieves a Precision of 0.64 and Recall of 0.77 (F1 ≈ 0.70). The method requires no manual training data and is already deployed in Sketch Engine for eight languages. We describe the user interface, which allows lexicographers to filter Word Sketches by sense, and discuss implications for dictionary-writing workflows.
Speaker: Ondřej Herman (Lexical Computing) -
162
Neologisms in Evaristo.ai vs. ChatGPT: Implications for LLM-Based Lexicography in European Portuguese
The diffusion of large language models (LLMs) has generated growing interest in their application to lexicographic practice, particularly for tasks such as the automatic generation of dictionary entries. Existing studies generally report promising results; however, they have primarily focused on major languages, most notably English, and rely on international, multilingual LLMs, such as ChatGPT (De Schryver, 2023; Lew, 2023; Rees & Lew, 2023), thereby overlooking the performance of language-specific, native LLMs. Accordingly, the lexicographic treatment of neologisms by LLMs has also received little attention. As emerging and often unstable lexical units that are sparsely attested or entirely absent from training data, neologisms limit empirical evidence for reliable semantic generalization and force models to rely on internal mechanisms of meaning composition rather than memorized patterns. This raises specific concerns regarding the extent to which LLMs can model lexical innovation in ways that are consistent with established lexicographic standards (Poix & Shevchenko, 2025). Some of the studies on LLMs and neologisms have approached the issue from a morphological perspective. Studies conducted in English and Polish (Manova, 2023) and Greek (Georgiou, 2025) suggest that LLMs perform unevenly across neologisms derived from different word-formation processes, generally achieving better processing results for derivatives than for compounds. However, since languages differ in their morphology and word-formation processes, it remains unclear to what extent these observations extend to other languages. In this context, the present study presents a comparative analysis of Evaristo.ai, the first chatbot designed for the Portuguese language, and ChatGPT-5.2, a multilingual LLM, focusing on their ability to generate lexicographic entries for neologisms arising from different word-formation processes, namely compounds, derivatives, and loanwords. Specifically, it examines whether and how word-formation influences the automatic lexicographic treatment of neologisms, assessing whether the contrasts previously observed between derivatives and compounds are reproduced in generated dictionary entries, and the extent to which the resulting entries adhere to established lexicographic standards.
This analysis was based on a set of fifteen neological units extracted from the Observatório Lexical (OL), a repository dedicated to the monitoring of lexical units that have not yet been incorporated into the Dicionário da Língua Portuguesa (DLP) of the Academy of Sciences of Lisbon, but whose attestation in contemporary usage warrants linguistic observation. These neologisms were subsequently categorized according to their word-formation process (Correia & Payo de Lemos, 2005; Rio-Torto et al., 2013) into (i) compounds (ciberautoritarismo, criptocorrupção, lgbtfóbico, neonepotismo, tecnomelancólico), (ii) derivatives (almirantismo, bolhificação, lepenismo, urgencializar, trumpização) and (iii) loanwords (brain rot, bromance, doomscrolling, gamekeeper, vlogger). A set of prototype dictionary entries was manually constructed in accordance with the DLP’s criteria and used as a gold standard. Subsequently, all neologisms were processed using a single standardised prompt (Lo, 2023), which specified the required entry structure and the mandatory lexicographic fields. The generated entries were evaluated by two annotators according to a systematic methodology based on the lexicographic conventions of the DLP, considering dimensions such as structural compliance, lemma presentation, grammatical category, definition, example of use, etymology, variety sensitivity (pt-PT), lexicographic style and register, and error typology. Each dimension was rated on a three-point scale (0–2), and entries were then classified as inadequate, partially adequate, or adequate, according to the deficiencies in core lexicographic fields and the degree of editorial intervention required. The results indicate that ChatGPT-5.2 generates lexicographic entries of higher overall adequacy across all neologism types, an expected outcome plausibly related to its training on substantially larger volumes of data. More importantly, however, the distribution of performance across word-formation processes diverges from the patterns reported in previous studies. In the present dataset, lexicographic entries for derivatives do not outperform those for compounds and instead display substantially lower adequacy, a pattern that is especially pronounced for Evaristo.ai: among the derivatives, Evaristo.ai produced three inadequate definitions, one partially adequate definition, and one adequate definition, whereas for compounds it produced two adequate definitions and three partially adequate definitions. This underperformance is primarily attributable to flawed etymological analyses, often involving implausible, hallucinated historical details or misidentified source languages or formative elements, as well as to deficiencies in the example field, which is frequently unclear, implausible, or poorly integrated with the lemma. These findings suggest that the automatic lexicographic treatment of neologisms by LLMs in European Portuguese may be influenced by word-formation processes, though not in the manner previously reported for other languages. Contrary to earlier findings, derivational neologisms appear to be more challenging than compounds, particularly for the language-specific model Evaristo.ai. This indicates that language-specific training does not systematically mitigate previously observed weaknesses. Overall, the study contributes to LLM-based lexicography by showing that word-formation processes affect not only neologism recognition, but also the quality and editorial reliability of automatically generated dictionary entries for European Portuguese.
Speakers: Leonor Martins (Academy of Sciences of Lisbon), Ana Salgado (University of Porto & Academy of Sciences of Lisbon)
-
159
-
11:00
Coffee Break 00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
00 | Aula (Entrance Hall), OeAW Main Seat, ground floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna -
Keynote: Large Language Models, Symbolic Meaning Representations and Ontologies: Perspectives and Questions 01 | Festsaal (Festive Hall), OeAW Main Seat, 1st floor
01 | Festsaal (Festive Hall), OeAW Main Seat, 1st floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaConvener: Jan Hajič (Charles University)-
163
Large Language Models, Symbolic Meaning Representations and Ontologies: Perspectives and Questions
In the talk, I will present the basics of Large Language Models (LLMs) and the status of the European LLM developments, in particular the OpenEuroLLM project that aims to develop open, multilingual foundational models in Europe and its challenges. I will then move on to summarise some recent developments on the other side of the research spectrum: symbolic meaning representations and latest developments of the most prominent ones, such as the Uniform Meaning Representation and the Prague Dependency Treebank. Such representations make heavy use of lexical resources – and I will show how these can be linked together and unified for making current and future meaning representation annotation efforts more consistent, while preserving all the information and alignment between form (plain text) and the representations. Towards the end of the talk, I will connect the current work on LLMs and meaning representations by postulating new research questions for which we now have resources and tools to work on.
Speaker: Jan Hajič (Charles University Prague)
-
163
-
164
Closing & Goodbye 01 | Festsaal (Festive Hall), OeAW Main Seat, first floor
01 | Festsaal (Festive Hall), OeAW Main Seat, first floor
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna -
165
Adam Kilgarriff Memorial Hike Meeting Point: Main Entrance, OeAW Main Seat
Meeting Point: Main Entrance, OeAW Main Seat
Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 ViennaSaturday, 3 October 2026, 2:00–7:00 p.m. (open end)
Our official conference image features one of Vienna’s most beautiful panoramic views—the vineyards of the Vienna Woods (Wienerwald) glowing in their autumn colours. On 3 October 2026, we invite you to experience this landscape in person.
Join us for a relaxed two-hour hike along a shortened section of Vienna’s City Hiking Trail 2 ( Stadtwanderweg 2 ), offering magnificent panoramic views over Vienna and its surrounding vineyards.
The hike concludes with a convivial gathering at Waldgrill Cobenzl, where we will bring the conference to a pleasant close in a relaxed atmosphere.
This excursion is much more than a traditional social event—it brings the visual identity of EURALEX Austria 2026 to life and offers our guests an unforgettable experience of Vienna in the golden colours of autumn.
The hike follows a shortened (approximately two-hour) section of City Hiking Trail 2 and does not include the ascent to Hermannskogel.
Meeting Point
The hike will take place only in good weather.
Please register by e-mail or let us know your participation no later than your arrival at the conference.
Meeting time: 2:00 p.m.
Meeting point: Main entrance of the Austrian Academy of Sciences (ÖAW)From there, we will travel together to the Sievering terminus of bus line 39A, where the hike begins.
Hike Details
- Distance: approximately 5.5 km
- Elevation gain: approximately 180 m
- Duration: approximately 2 hours
The route leads through the open vineyards and via Bellevue Hill (Bellevue-Höhe), offering spectacular views over Vienna. The hike concludes with refreshments at Waldgrill Cobenzl, where participants are welcome to relax and enjoy a drink or an optional meal at their leisure.
Return Journey
Following our stop at Waldgrill Cobenzl, participants may return by bus 38A to Heiligenstadt station and continue on Underground line U4.Recommended Connections
Return to the Austrian Academy of Sciences (ÖAW):
- Take the U4 to Landstraße – Wien Mitte.
Direct return to Vienna Airport:
- Take the U4 to Schwedenplatz.
- Change to bus 2A to Schwarzenbergplatz
- Walk to Salztorbrücke / Morzinplatz and take the Vienna Airport Lines (VAL) bus to Vienna Airport.
Approximate travel time: 1 hour 40 minutes.
-