29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

Camfranglais in the Age of Artificial Intelligence: Lexicographic Challenges, Opportunities and a Human-AI Approach to Documenting Linguistic Hybridity in Multilingual Cameroon

1 Oct 2026, 16:00
30m
02 | Museumszimmer (02 | Museumszimmer, OeAW Main Seat, 2nd floor)

02 | Museumszimmer

02 | Museumszimmer, OeAW Main Seat, 2nd floor

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speaker

Emmanuel Sylvain Fomat (University of Hildesheim)

Description

In postcolonial societies, language often functions as both a site of power and resistance, a reality profoundly exemplified by Camfranglais in multilingual Cameroon. As Foucault (1978, pp. 95–96) argues, power inevitably generates resistance, and in Cameroon this resistance manifests linguistically through a hybrid sociolect that disrupts inherited colonial hierarchies. Emerging in the 1970s within a “communicative vacuum” (Kießling, 2015, p. 9) created by exoglossic language policies, Camfranglais combines French, English, Cameroonian Pidgin English, and numerous indigenous languages. Despite its widespread use and cultural significance, it remains substantially under-documented. Existing lexicographic resources resemble glossaries or linguistic guides constrained by static print formats, limited corpora, and the absence of systematic phonetic transcription. They also reveal a metalanguage mismatch: Kamdem Fonkoua (2015) provides English definitions with untranslated Camfranglais examples, creating barriers for Francophone users. Essome Bouti (2024) and Ndongo (2015) offer French metalanguage, which serves Francophone speakers but excludes Anglophone Cameroonians. Neither orientation serves speakers from the less-educated variety identified by Ebongue and Fonkoua (2010), who may have acquired French or English informally rather than through formal schooling. Consequently, a significant gap persists between the sociolinguistic reality of Camfranglais and its lexicographic representation.

The rise of artificial intelligence opens new possibilities for low-resource language documentation through large-scale data processing, pattern recognition, and phonetic prediction (Zhong et al., 2024; Wang, 2024; Midigo 2025). However, these opportunities are constrained by structural biases in current NLP systems, which are predominantly trained on high-resource and typologically stable languages. Ebrahimi et al. (2022, p. 6279) show that multilingual models achieve only 38.48% average zero-shot accuracy on genuinely low-resource languages. To address this tension, this paper proposes a Human–AI collaborative framework for documenting Camfranglais, grounded in the Function Theory of Lexicography (Bergenholtz & Tarp, 2002, 2003; Tarp, 2008a, 2008b). This framework informs the design of a bilingual, bidirectional digital dictionary with dual metalanguage access (French–Camfranglais and English–Camfranglais), intended to serve both Francophone and Anglophone users. Critical Discourse Analysis (CDA) is also applied to capture how Camfranglais expresses identity, ideology, and power relations in context (Fairclough, 1989; Blommaert & Bulcaen, 2000; Hidalgo Tenorio, 2011; Discourse Analyzer, 2024).

The study is further informed by an exploratory survey of forty Cameroonian speakers from diverse sociolinguistic backgrounds. Results confirm widespread informal use of Camfranglais and strong support for its documentation as a marker of national identity, while also revealing concerns that excessive standardisation could undermine its creativity and fluidity. In response, AI is positioned as a “copilot” subordinated to human expertise, ensuring linguistic sovereignty and aligning with principles of anti-extractivism (Riofrancos, 2020).

Methodologically, the study relies on a deliberately constructed multi-source corpus capturing both historical written forms and contemporary oral-digital manifestations of Camfranglais. The written component includes the Grioo Forum (2004–2006), documenting early diasporic digital interaction, and the Bonaberi Forum (2008–2026), reflecting spontaneous written exchanges on everyday topics. The oral-digital component is drawn from the YouTube channel Warman du Terre à Terre (2013–2026). From this source, 1,583 video URLs were extracted to map lexical diversity, and 100 videos were fully transcribed to capture spontaneous speech, prosody, and emerging vocabulary. Together, these sources provide a heterogeneous corpus documenting the evolution of Camfranglais from street speech to digital discourse.

The analytical workflow operationalises a three-stage Human–AI feedback loop. First, computational tools identify candidate entries, frequency patterns, and collocations. Python was used for statistical extraction, while Sketch Engine (Kilgarriff et al., 2014) enabled deeper contextual analysis through concordances and Key Word In Context (KWIC) functions, particularly useful for identifying hybrid forms often missed by automated methods. Second, the Large Language Model Claude generated contextual definitions and etymological notes, while ChatGPT produced candidate phonetic transcriptions. Only standard French items were excluded; Cameroonian French forms with semantic or phonological deviation, as well as English items phonologically adapted by Camfranglophones, were retained as authentic lemmas. This procedure yielded approximately 2,850 lexical items, including multi-word expressions.

Finally, all outputs were validated by fourteen native Camfranglais speakers through independent five-point Likert-scale evaluations followed by consensus sessions. Results highlight both the strengths and limitations of AI-assisted lexicography. While AI accelerates extraction and contextual analysis, it struggles with tonal minimal pairs, pragmatic inversion, neologisms, hybrid morphology, and distinctions between Cameroonian and standard French. Nevertheless, validation achieved substantial inter-rater agreement (Fleiss’ kappa = 0.74) (Landis & Koch, 1977), with high phonetic concordance: 94.2% perceptual agreement; 96.8% acoustic agreement using Praat (Boersma & Weenink, 2026).

This paper argues that a principled Human–AI approach is not merely a technical convenience but an epistemic necessity for documenting Camfranglais. By combining computational efficiency with native-speaker expertise, the study produces a pilot electronic dictionary that is phonologically informed, sociolinguistically grounded, and adaptable to ongoing lexical change. More broadly, it offers a transferable model for documenting hybrid urban varieties and demonstrates how artificial intelligence can support marginalized speech communities without reproducing extractive dynamics.

Presentation materials

There are no materials yet.