29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

Neologisms in Evaristo.ai vs. ChatGPT: Implications for LLM-Based Lexicography in European Portuguese

3 Oct 2026, 10:30
30m
01 | Sitzungssaal (01 | Sitzungssaal, OeAW Main Seat, 1st floor)

01 | Sitzungssaal

01 | Sitzungssaal, OeAW Main Seat, 1st floor

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speakers

Leonor Martins (Academy of Sciences of Lisbon) Ana Salgado (University of Porto & Academy of Sciences of Lisbon)

Description

The diffusion of large language models (LLMs) has generated growing interest in their application to lexicographic practice, particularly for tasks such as the automatic generation of dictionary entries. Existing studies generally report promising results; however, they have primarily focused on major languages, most notably English, and rely on international, multilingual LLMs, such as ChatGPT (De Schryver, 2023; Lew, 2023; Rees & Lew, 2023), thereby overlooking the performance of language-specific, native LLMs. Accordingly, the lexicographic treatment of neologisms by LLMs has also received little attention. As emerging and often unstable lexical units that are sparsely attested or entirely absent from training data, neologisms limit empirical evidence for reliable semantic generalization and force models to rely on internal mechanisms of meaning composition rather than memorized patterns. This raises specific concerns regarding the extent to which LLMs can model lexical innovation in ways that are consistent with established lexicographic standards (Poix & Shevchenko, 2025). Some of the studies on LLMs and neologisms have approached the issue from a morphological perspective. Studies conducted in English and Polish (Manova, 2023) and Greek (Georgiou, 2025) suggest that LLMs perform unevenly across neologisms derived from different word-formation processes, generally achieving better processing results for derivatives than for compounds. However, since languages differ in their morphology and word-formation processes, it remains unclear to what extent these observations extend to other languages. In this context, the present study presents a comparative analysis of Evaristo.ai, the first chatbot designed for the Portuguese language, and ChatGPT-5.2, a multilingual LLM, focusing on their ability to generate lexicographic entries for neologisms arising from different word-formation processes, namely compounds, derivatives, and loanwords. Specifically, it examines whether and how word-formation influences the automatic lexicographic treatment of neologisms, assessing whether the contrasts previously observed between derivatives and compounds are reproduced in generated dictionary entries, and the extent to which the resulting entries adhere to established lexicographic standards.

This analysis was based on a set of fifteen neological units extracted from the Observatório Lexical (OL), a repository dedicated to the monitoring of lexical units that have not yet been incorporated into the Dicionário da Língua Portuguesa (DLP) of the Academy of Sciences of Lisbon, but whose attestation in contemporary usage warrants linguistic observation. These neologisms were subsequently categorized according to their word-formation process (Correia & Payo de Lemos, 2005; Rio-Torto et al., 2013) into (i) compounds (ciberautoritarismo, criptocorrupção, lgbtfóbico, neonepotismo, tecnomelancólico), (ii) derivatives (almirantismo, bolhificação, lepenismo, urgencializar, trumpização) and (iii) loanwords (brain rot, bromance, doomscrolling, gamekeeper, vlogger). A set of prototype dictionary entries was manually constructed in accordance with the DLP’s criteria and used as a gold standard. Subsequently, all neologisms were processed using a single standardised prompt (Lo, 2023), which specified the required entry structure and the mandatory lexicographic fields. The generated entries were evaluated by two annotators according to a systematic methodology based on the lexicographic conventions of the DLP, considering dimensions such as structural compliance, lemma presentation, grammatical category, definition, example of use, etymology, variety sensitivity (pt-PT), lexicographic style and register, and error typology. Each dimension was rated on a three-point scale (0–2), and entries were then classified as inadequate, partially adequate, or adequate, according to the deficiencies in core lexicographic fields and the degree of editorial intervention required. The results indicate that ChatGPT-5.2 generates lexicographic entries of higher overall adequacy across all neologism types, an expected outcome plausibly related to its training on substantially larger volumes of data. More importantly, however, the distribution of performance across word-formation processes diverges from the patterns reported in previous studies. In the present dataset, lexicographic entries for derivatives do not outperform those for compounds and instead display substantially lower adequacy, a pattern that is especially pronounced for Evaristo.ai: among the derivatives, Evaristo.ai produced three inadequate definitions, one partially adequate definition, and one adequate definition, whereas for compounds it produced two adequate definitions and three partially adequate definitions. This underperformance is primarily attributable to flawed etymological analyses, often involving implausible, hallucinated historical details or misidentified source languages or formative elements, as well as to deficiencies in the example field, which is frequently unclear, implausible, or poorly integrated with the lemma. These findings suggest that the automatic lexicographic treatment of neologisms by LLMs in European Portuguese may be influenced by word-formation processes, though not in the manner previously reported for other languages. Contrary to earlier findings, derivational neologisms appear to be more challenging than compounds, particularly for the language-specific model Evaristo.ai. This indicates that language-specific training does not systematically mitigate previously observed weaknesses. Overall, the study contributes to LLM-based lexicography by showing that word-formation processes affect not only neologism recognition, but also the quality and editorial reliability of automatically generated dictionary entries for European Portuguese.

Presentation materials

There are no materials yet.