Speakers
Description
We introduce and describe the DICI-A, a learner dictionary of Italian collocations which is the output of a two-year project funded by the Italian Ministry of University and Research. In the Italian lexicographic landscape, three dictionaries of collocations have been published in the last fifteen years (Lo Cascio, 2013; Tiberii, 2012; Urzì, 2009). However, none of the three has two fundamental features that can be found in DICI-A: it is specifically targeted at L2 learners of Italian, and it has been created according to corpus-based criteria. Other features of the DICI-A are:
- it is monolingual;
- it includes over 10,000 collocation entries belonging to six syntactic configurations;
- each collocation is assigned to a specific proficiency label;
- GenAI and human assessment have been integrated for the creation of definitions and examples;
- it is digital and freely available online, at https://dictionary.dici-a.it/.
We relied on a broad definition of collocation: a co-occurrence of two words with a syntactic relation characterised by its conventional meaning, resulting from the number of times it is used in naturally occurring language (frequency), the range of texts where it occurs (dispersion) and the extent to which its components attract each other and are strongly associated (association measures), either adjacently or within a distance.
In addition to a general description, the presentation will focus on three specific features of the dictionary.
1) Identification and selection of the dictionary entries
We have included in the dictionary collocations that fall into six syntactic types: i. Verb + Direct object (vdobj; mantenere una promessa, ‘to keep a promise’); ii. Adjective + Noun/Noun + Adjective, the adjective is a modifier before or after a noun (amod; brutta avventura, ‘bad adventure’; tempo libero, ‘free time’); iii. Verb + Adjective, the adjective functions like an adverb by modifying the verb (advmod1; stare zitto, ‘to stay quiet’); iv. Verb + Adverb, the adverb modifies the verb (advmod2; fare presto, ‘to hurry up’); v. Adverb + Adjective, the adjective is modified by the adverb (advmod3; altamente positive, ‘highly positive’); and vi. Noun + Noun, compounds made of two adjacent nouns (comp; parco divertimenti, ‘amusement park’).
The collocations belonging to these six syntactic configurations were extracted from the PEC24 (Spina et al., 2025), a large reference corpus of written and spoken Italian, by combining pos-tagging and syntactic parsing (Seretan, 2011). These initial 2 million candidate collocations were filtered via a multi-method approach (Spina et al., 2026) involving both automatic stages (based on five different measures: dispersion, frequency, mutual information, log-dice and log-likelihood) and human evaluation, after a comparison between the filtered candidate collocations and two existing non corpus-based Italian collocation dictionaries.
2) Attribution of proficiency labels to each entry
The 10,596 collocations resulting from this selection process were assigned CEFRCV-based proficiency labels (Common European Framework of Reference Companion Volume; Council of Europe, 2020) based on a set of quantitative and qualitative criteria, including corpus frequency and dispersion, semantic transparency, register, and CEFRCV vocabulary range descriptors. This procedure was entirely based on human assessment: each collocation was annotated independently by two annotators; then a third annotator resolved the cases of disagreement. Inter-annotator agreement varied across collocation types; as an example, it reached 80% for the verb-direct object collocations.
3) Creation of definitions and examples with the support of GenAI
We relied on recent literature (Lew, 2023; Lew et al., 2024) showing that GenAI can be effective in speeding up the process of dictionary creation, performing well both in writing definitions and in producing examples. Thus, we asked ChatGPT-4o to assist us in identifying collocation’s meaning(s) and in providing us with definitions and examples for each meaning. In our prompt, we strongly emphasised the need for the output to contain learner-friendly vocabulary suitable for learners of Italian. The 10,596 AI-generated pairs of definitions and examples were validated by human lexicographers and replaced or modified accordingly. After this evaluation process 72% pairs could be accepted as they were in the AI-generated version; 12% required changes in the definition, 12% had to be modified in the example, and only 4% had to be completely rewritten.