Speaker
Description
Identifying semantic neology remains a significant challenge for lexicography due to the gradual nature of meaning shifts. This study proposes a methodological framework for the automatic detection of verbal semantic neology in Catalan, leveraging recent advances in Natural Language Processing (NLP). Utilizing a general-language corpus spanning 2000-2021, this research employs a Large Language Model to generate contextualized embeddings for 90 target verbs, capturing nuanced semantic variation. The HDBSCAN clustering algorithm is then applied to group these representations by usage similarity. To validate the results, a lexicographic baseline is established using the Diccionari essencial de la llengua catalana; clusters that do not align with registered senses are flagged as candidate neologisms for qualitative analysis. This qualitative analysis focuses on shifts in argument structure and semantic implicatures. The findings emphasize the value of NLP tools for low resource languages like Catalan that have growing but comparatively few digital resources. Ultimately, this study advocates for a hybrid lexicographic model where machine learning serves as an exploratory tool to augment, rather than replace, human expertise. This approach ensures that modern dictionaries can responsibly integrate technological innovation while maintaining academic and ethical standards.