Speaker
Description
When non-specialist users consult general dictionaries to understand legally significant terms, they risk encountering definitions that are fundamentally incomplete. This study introduces a reproducible computational methodology to quantify the level of semantic reduction and explore AI tools to resolve it. Five socio-legal terms were examined from the Cambridge Dictionary, Oxford English Dictionary, Merriam-Webster (general) and Black’s Law Dictionary, Oxford Dictionary of Law, Merriam-Webster’s Dictionary of Law (legal): harassment, stalking, discrimination, abuse, and neglect. General and legal definitions (35 and 15 respectively) were vectorised using Sentence-BERT and compared via cosine similarity. None of the five terms reached the 0.70 similarity threshold. Mean similarity ranges from 0.491 (abuse) to 0.571 (stalking); 92.4% of all pairwise comparisons fall below this threshold. Statistical significance was confirmed by one-sample t-tests (p < 0.01). The Qwen2.5-3B-Instruct model extracted key legal components for each term, which can serve as editorial checklists for lexicographers reviewing dictionary entries. The findings demonstrate that semantic reduction is systematic, measurable, and varies across terms. AI tools can assist lexicographers in detecting and addressing such gaps, though final decisions remain with the human expert.