Speakers
Description
This paper describes efforts to extract frequency information from corpora and to rank the collocates in Woordcombinaties, a Dutch phraseological lexicographic resource, in a uniform and systematic way, taking various factors such as the size and composition of the corpus, linguistic processing and the mathematical properties of various association measures into account. A small sample of 2032 combinations, divided over 19 verbs and 19 nouns, was used, and their frequencies were extracted from the Corpus Contemporary Dutch, using surface co-occurrence, textual co-occurrence, and syntactic co-occurrence. Results were verified manually. It was found that: 1) corpus size increases coverage; 2) syntactic co-occurrence leads to loss of coverage; 3) combination of a narrow window span and within-sentence context seems to have the highest coverage; 4) separate treatment is needed to deal with multi-word collocates. However, given that surface/textual co-occurrence still produces a substantial amount of false positives, we intend to use syntactic co-occurrence to retrieve the frequency information for the combinations in Woordcombinaties, because it has a higher precision and offers the greatest scope for improvement.