29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

Analysis of Synthetic Clinical Neologisms: A Corpus-Based Analysis of Large Language Model-Generated Content in the Context of Artificial Intelligence-Assisted Clinical Documentation

30 Sept 2026, 14:00
1h 30m
Aula (Entrance Hall) & 1st floor (00/01 | Aula (Entrance Hall) & 1st floor, OeAW Main Seat)

Aula (Entrance Hall) & 1st floor

00/01 | Aula (Entrance Hall) & 1st floor, OeAW Main Seat

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speaker

Amila Kugic (Medical University of Graz)

Description

With generative artificial intelligence (AI) tools becoming more widely available and integrated into clinical environments, physicians are more frequently exploring ways to integrate AI into their clinical documentation and decision-making processes (Brender et al., 2025). Research already shows that large language models (LLMs) have the capability for summarization and analysis of clinical documents but carry also risks towards introduced hallucinated or biased content that can potentially skew further diagnostic steps and treatment options (Busch et al., 2025). Previous research suggested that the addition of neologisms to texts ended up decreasing machine translation quality by an average of 43% (Zheng et al., 2024).

This current investigation examined synthetically LLM-generated German clinical corpora and the use of neologisms and lexical patterns that differ from established controlled vocabularies, such as the Unified Medical Language System (UMLS) and SNOMED CT. Four synthetic German clinical training datasets were generated using GPT-4o and GPT-5 models. Each dataset comprising 900 snippets was constructed to compare zero-shot and few-shot prompting strategies, where both prompts instructed the generation of short, approximately 100-character length clinical-style snippets. LLM outputs were designed to reflect realistic clinical documentation features, such as abbreviated syntax, incomplete sentences, and laboratory-style reporting. Words or multiword expressions absent from reference biomedical corpora were flagged as candidate neologisms. These were then analyzed based on their morphological and structural properties.

To assess semantic relatedness between candidate neologisms and established biomedical terminology, embedding-based similarity analysis was performed using a pretrained multilingual transformer model (sentence-transformers/paraphrase-multilingual-mpnet-base-v2). Each candidate term and each controlled vocabulary entry was converted into a vector representation. The nearest-neighbor approach via cosine similarity allowed each candidate form to be linked to its most semantically similar standardized form. The method compared candidate terms that closely match existing medical concepts linked in terminologies, and those with low similarity scores, which were more likely to represent hallucinated expressions. The similarity scores were furthermore combined with frequency information to highlight commonly occurring but semantically distant forms for further analysis. This embedding-based comparison supported also frequency-based filtering of candidate terms for manual validation and review.

Preliminary analysis of synthetically generated German clinical narrative snippets revealed that the majority of candidate neologisms were not entirely novel lexical terms with reference to SNOMED CT and UMLS, but rather orthographic variants, abbreviation modifications, and fragmented forms derived from existing biomedical terminology. Co-occurrence analysis demonstrated clustering around conventional documentation templates when few-shot prompting was implemented in the synthetic generation pipeline. This suggested that neologisms primarily emerged through recombination and modification of established documentation patterns. Embedding-based similarity analysis further indicated that most candidate forms remained semantically close to controlled-vocabulary concepts, while only a small subset showed characteristics consistent with hallucinations. The comparison between prompting strategies revealed differences in lexical variability. Few-shot prompting led towards repetitive documentation templates, e.g., repeating the same structure and information found within the example prompt though varied slightly by changing abbreviations or contents. This meant that the few-shot method introduced rigidity with prompt examples and reduced lexical diversity, whereas zero-shot prompting produced a higher degree of structural variability and a greater proportion of candidate non-standard forms.

From a lexicographic perspective, this approach offers a systematic way to identify emerging terminology and non-standard usages in specialized corpora. By mapping novel forms to established vocabularies, lexicographers can better document language innovation, improve term standardization, and support the creation of more comprehensive, up-to-date lexical resources for clinical, technical and scientific domains.

Presentation materials

There are no materials yet.