Speakers
Description
Recent research in lexicography has shown that the integration of generative AI into dictionary-making can work well for bilingual tasks (Chen, 2025; de Schryver, 2025; Han & de Schryver, 2025). The present research presents a corpus-based study of Chinese–English parallel business news and explores its implications for bilingual business lexicography in the age of AI. The study is based on a purpose-built Chinese–English parallel corpus drawn from the “Bilingual News” section of China Daily, comprising over 36,000 Chinese tokens and 42,000 English tokens. Although the corpus is relatively small and restricted to a single source, it includes all available bilingual business news published by this major official outlet in 2024. It can therefore be seen as a complete dataset within a clearly defined domain.
The English texts are typically original reports aimed at international audiences, while the Chinese texts are translations for domestic readers. Using the Sketch Engine, the analysis combines single-word keyword extraction and multi-word term extraction, with comparisons against large general-news reference corpora in both languages. A minimum frequency threshold of 5 is applied, and items with very low keyness scores are excluded, as they do not show domain-specific salience. This approach makes it possible to identify lexically salient items that are characteristic of contemporary business and business-related policy discourse.
The findings show clear cross-linguistic asymmetries. Chinese business texts tend to be data-driven and domestically oriented, favouring compact noun compounds, numerical indicators, and institution-specific terminology that compress complex economic relations into dense nominal structures. English texts, by contrast, rely more heavily on abstract policy framing, analytic syntactic constructions, and standardised acronyms and compound expressions, with repeated use of established policy phrases. These asymmetries should be interpreted with caution. As the English texts are produced within Chinese official media contexts, they may reflect patterns of China English and may differ from the expressions typically used in media from English-speaking countries, rather than purely language-internal variation.
Across both languages, multi-word units show markedly higher keyness values than single-word items. This suggests that meaning in this domain is primarily conveyed through fixed expressions, collocations, and policy phrases. While some high-frequency Chinese–English multi-word pairs have stabilised as standard equivalents, many Chinese business terms do not have direct lexical counterparts in English. These items are often culture- or institution-specific and therefore call for a paraphrastic or definition-based treatment rather than a word-for-word translation.
Building on these findings, the study proposes several lexicographic implications. First, it argues for prioritising multi-word units as dictionary headwords, as these units capture recurrent usage and encode domain-specific concepts more directly. Second, it supports a morpheme-plus-collocation entry structure. Under this structure, core lexical elements are systematically linked to their most frequent and productive multi-word patterns. Third, it advances a dual-track dictionary model. This model is designed to accommodate both the production-oriented needs of Chinese users and the reception-oriented needs of foreign users, allowing lexicographic content to be organised according to different user situations and functions.
To explore the role of genAI in contemporary lexicographic practice, the study further includes an AI chatbot, DeepSeek-V3, to generate preliminary dictionary entries. These are generated entirely by AI, with prompts iteratively refined to improve the quality and structure of the output. The prompt is shown in Addendum A. DeepSeek-V3 was provided with compilation notes, as seen in Addendum B. An example of generated entries is shown in Addendum C.
The study compares AI-generated entries with human-produced lexical resources. When accessed via dictionary apps, reference works such as the New Century Chinese–English / English–Chinese Dictionary (2026) and the Youdao Dictionary (2026) mainly provide equivalents, but do not offer full definitions nor usage examples. Overall, then, the present study demonstrates how corpus evidence and genAI chatbots can jointly inform the next-generation bilingual specialised dictionary design. By integrating corpus-based insights and AI assistance, the paper contributes to ongoing discussions on dictionary-making processes in the age of AI.