29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

Mapping Sociolinguistic Theory into Practice: Headword Candidate Selection in the Dictionary of Pluricentric Portuguese

2 Oct 2026, 10:40
20m
01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)

01 | Johannessaal

01 | Johannessaal, OeAW Main Seat, 1st floor

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speakers

Tanara Zingano Kuhn (CELGA-ILTEC, University of Coimbra) Vojtěch Kovář (Lexical Computing) Miloš Jakubíček (Lexical Computing)

Description

Presently in its preliminary phase at the Research Centre for General and Applied Linguistics at the University of Coimbra (CELGA-ILTEC), the dictionary of pluricentric Portuguese project seeks to develop an online dictionary that documents the usage of Portuguese across diverse global contexts. Considering the socio-historical complexities inherent to the Portuguese language area (Faraco, 2016; Correia, in print), this project’s lexicographic work is grounded in a deeper understanding of language use and its social context, with sociolinguistic theory serving as a core guiding framework. The purpose of this paper is to present how theoretical insights from sociolinguistics translate into practical lexicographic decisions regarding headword candidate selection.

Portuguese is the official language of nine countries (Angola, Brazil, Cabo Verde, Guinea-Bissau, Equatorial Guinea, Mozambique, Portugal, Sao Tome and Principe, and Timor-Leste) and one territory (Macao), although its actual functional status varies significantly in these multilingual regions. Traditionally, Portuguese is considered to have two dominant varieties, Brazilian Portuguese and European Portuguese, with the latter being adopted as the norm in the other countries. The lexicography of the language reflects this bicentrism and dictionary production is largely centred in Brazil and Portugal (Kuhn & Correia, 2025). However, in countries where Portuguese was introduced as a result of colonisation (all countries except Portugal), research has revealed discrepancy between real used norms and the official discourse legislating the language, including teaching materials, reference instruments such as dictionaries, and official documents. In Brazil, for instance, the dominant exonormative view of the language has led to a significant gap between how language is actually used by formally educated people in contexts requiring more language monitoring, i.e., the cultivated standard, and its prescriptive use, or the ideal language standard (Faraco, 2008). Relatedly, studies on the real norms used in the other countries include the description of Mozambican Portuguese by Gonçalves P. (2010) and Firmino (2011), Angolan Portuguese by Adriano (2015) and Inverno (2009), and Sao Tomean Portuguese by Gonçalves R. (2016) and Bouchard (2017). Considering all this in this dictionary project requires, in addition to a robust foundational sociolinguistic element, the compilation of a corpus that reflects the real use of Portuguese, in a variety of contexts and in different territories. Since a corpus with these characteristics is not available, a decision was made to carry out a pilot-dictionary project, using a small corpus of tweets (Rodrigues Gomide, 2022). The purpose of the pilot is to try methodologies, evaluate results, and develop solutions to the problems found, in a smaller scale, alongside the compilation of a larger corpus and further construction of the theoretical framework. For that, we have adopted the Dictionary Express method (Baisa et al., 2019). The planned workflow consists of several steps of iterative collaboration between human lexicographers and automatically extracted data, with headword candidate selection as the first step. This involves headword annotation using a specially created interface, with different lexicographers assigning flags (e.g., not a lemma, non-standard, etc.) to the same set of candidates and final headword list definition resulting from general agreement among lexicographers. Differently from previous projects for other languages (Baisa et al., 2019; Blahuš et al., 2023; Kovařík et al., 2024), in our project this method is used to make a dictionary that describes, at the moment, eight varieties of one language. One additional challenge stems from the sociolinguistic grounding of the project, by which traditional views on standard language are questioned. As a result, some changes have been made, such as the decision to not filter the list of automatically extracted candidate headwords for part of speech. This is because POS-taggers have been developed for Brazilian and European Portuguese, thus words from other varieties or words that are not considered to be standard might not be identified, leading to wrong POS-tagging which, in turn, would affect the frequency count. Another aspect of the headword candidate annotation process that is especially crucial to our project refers to the flag “non-standard”. This type of decision-making requires a previous common-ground perspective on standardisation and standard language ideologies (McLelland, 2021) and how they affect the representation of language in the dictionary. This consideration becomes even more critical in the case of Portuguese, which is spoken across multilingual regions marked by complex socio-political-cultural contexts, particularly considering the impact of colonial history and the enduring legacy of colonial language policies. At the time of writing, this step of the workflow is still ongoing; however, further results are expected for the time of the conference. With the presentation of how the theoretical framing of this project maps into practical lexicographic work regarding headword candidate selection, we hope to share our experience with our peers, as well as contribute to questioning and changing long-established language ideologies and language attitudes prevailing in the Portuguese language area.

Presentation materials

There are no materials yet.