Speakers
Description
This paper presents the methodological framework for building the Similes Repository for Contemporary Serbian (Simili-SR), a corpus-based linguistic resource designed for the collection, annotation, and analysis of similes in contemporary Serbian. The study presents a multi-stage extraction and validation pipeline based on the genre-balanced corpus SrpKor2021+ of the Reference Corpus of Contemporary Serbian, combining CQL queries, local grammars implemented in UNITEX, and LLM-assisted semantic annotation. The initial extraction phase produced 574,195 candidate sentences across six major simile patterns, confirming both the productivity and structural diversity of simile constructions in Serbian. Subsequent rule-based filtering and LLM-assisted validation substantially reduced noise while preserving a large pool of linguistically relevant figurative comparisons. Manual evaluation performed on a stratified sample additionally revealed that simile interpretation frequently involves distributed semantic structure, and complex interactions between grammatical realization and figurative meaning, and highlighted the importance of linguistic post-processing, as well as several challenges related to semantic annotation in morphologically rich languages such as Serbian. By combining rule-based methods, LLMs, and detailed linguistic analysis, the paper contributes to both the development of computational resources for Serbian and broader research on figurative language annotation and representation.