29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

What LLMs Learn and What They Miss: Frequency Effects in the Automated Description of Korean Verbal Inflection

29 Sept 2026, 16:30
30m
01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)

01 | Johannessaal

01 | Johannessaal, OeAW Main Seat, 1st floor

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speakers

Li Cui (Yonsei University) Seol Namkung (Yonsei University) Jun Lee (Yonsei University) Hae-Yun Jung (Kyungpook National University) Kilim Nam (Yonsei University)

Description

The present study aims to empirically examine the possibilities and limitations of LLMs as a lexicographic tool for Korean, by automatically generating inflected forms of Korean verb and adjective headwords using Large Language Models and contrasting the results with corpus frequency data. Focusing on the description of inflected forms of Korean verbs and adjectives, the study analyses in detail the extent to which LLMs reproduce corpus-based frequency rankings and the accuracy of inflected form generation, divided into the categories of regular, irregular, and defective predicates. The results show that LLMs consistently produce various errors in low-frequency inflected forms and reveal clear limitations in precisely reproducing frequency information. This empirical examination points to the possibilities and risks inherent in LLMs as lexicographic assistance tools from multiple angles. While LLMs hold meaningful value as an assistive means for drafting dictionary descriptions, lexicographic reliability can only be guaranteed when this is complemented by systematic corpus-based verification and the distinctive expertise of lexicographers. These findings reaffirm the fundamental role that lexicographers must play in the automation of dictionary compilation.

Presentation materials

There are no materials yet.