Speakers
Description
Large language models (LLMs) continue to perform suboptimally in tasks involving metaphor detection and explanation (Devlin et al., 2019; Liu et al., 2022; Klemen & Robnik Šikonja, 2023, Kim et al., 2023; He et al., 2023; Despot, Ostroški Anić, & Veale, 2023; Tong et al., 2024; Puraivan et al., 2024; Yang et al., 2025; Chen & Wang, 2025). Core challenges lie in the inherent semantic complexity of metaphorical expressions, and in the dearth of structured metaphor corpora. In response to this limitation, we introduce CroSloMet, a novel structured metaphor dataset for two South Slavic languages, Croatian and Slovene, designed to support both metaphor identification and explanation tasks (Ge et al., 2023, Despot et al., 2025). CroSloMet comprises over 1,120 metaphorical sentences and 1,120 matched literal sentences for each language, annotated with semantic metadata. Each data point is aligned with corresponding conceptual metaphors, multi-word expressions, canonical forms, and literal usages, providing a multi-layered annotation scheme that goes beyond surface metaphor tagging to encode conceptual and linguistic structure. This resource is grounded in the MetaNet.HR framework (Despot et al., 2019) enabling researchers to explore generality and specificity in metaphoric expressions. The dataset supports metaphor identification tasks—the classification of sentences as metaphorical vs. literal through annotated pairs that preserve lexical overlap between figurative and non-figurative contexts. Moreover, it facilitates metaphor explanation (conceptual metaphor type recognition, where the conceptual metaphor name is used as a proxy for explanation), allowing evaluation of model output not merely on binary classification but on the degree to which generated conceptual metaphor type capture intended conceptual mappings. CroSloMet’s dual annotation quality makes it suitable for corpus-based analysis, cognitive linguistics studies, and the development of metaphor understanding modules in NLP systems. To demonstrate the dataset’s utility, the original work reports preliminary experiments using a fine-tuned CroSloEngual BERT model for metaphor classification, achieving an accuracy of 88.5%, and an evaluation of LLaMA 3-8B for metaphor detection. While classification results were promising, strict exact-match evaluation for explanation generation (conceptual metaphor type recognition) yielded low scores, revealing a gap between model outputs and human interpretive expectations despite the semantic validity of many generated texts. This discrepancy pointed to the need for improved evaluation metrics that capture semantic similarity and interpretive nuance.
Building on these insights, the current paper proposes to extend CroSloMet with a multi-level validation framework based on the comparison of manual annotation to natural language inference (NLI), semantic similarity scoring, and LLM-based judgment to assess metaphor explanations more holistically. By integrating human expert judgments with automatic evaluation signals, the proposed research seeks to close the gap between computational performance and cognitive plausibility in metaphor interpretation. This study leverages this dataset and evaluation framework to address several core research questions: How can structured metaphor annotations be exploited to improve cross-lingual metaphor understanding? What evaluation strategies best capture the quality of metaphor explanations in LLM-generated text? And how can conceptual metaphors be operationalized to support both symbolic and distributional approaches to figurative language modelling? To answer these questions, the paper outlines an experimental pipeline that integrates CroSloMet with advanced representation learning techniques, such as transformer architectures and semantic embedding spaces calibrated for figurative meaning. A key aspect of this research is the incorporation of semantic generalization layers within explanation models that can abstract over lexical variation to capture deep conceptual mappings. We take advantage of the hierarchical organization of conceptual metaphors in the MetaNet.HR database (Despot et al., 2019) to refine the evaluation of model-generated metaphor explanations. Instead of relying solely on exact matches between predicted and gold-standard conceptual metaphors, we implement a graded evaluation scheme that accounts for metaphor hierarchy and allows us to calculate conceptual distance or similarity between predicted and reference metaphors even if its formulation is more abstract or differently phrased. By combining frame-based representations with contextualized embeddings, the study aims to evaluate metaphor understanding at the level of conceptual pattern recognition – a capability that bridges computational modelling and cognitive theories of metaphor (Lakoff and Johnson, 1980; 1999 – for an overview of different and more recent approaches see Dancygier & Sweetser, 2014 and Despot, 2024). The dataset’s parallel nature further enables cross-lingual transfer experiments, where models trained on one language can be tested on another. To ensure that CroSloMet can serve as a shared resource for language technology research, supporting tasks such as frame mapping across languages and constructive evaluation of metaphor explanation models aligned with human interpretive criteria, CroSloMet is aligned with standardized semantic tagging schemes, and can be exported in interoperable formats providing benchmarks for corpus-based metrics.
This paper presents CroSloMet as both a robust dataset for metaphor research in under-resourced languages and an impetus for future method development in metaphor interpretation. By articulating a clear research agenda centered on structured annotations and nuanced evaluation, the work addresses persistent gaps in metaphor understanding with computational models. The proposed multi-level validation framework and cross-lingual experimental design position CroSloMet as a contribution to lexicography, computational linguistics, and cognitive semantics, making it relevant for figurative language modelling.