29 September 2026 to 3 October 2026
OeAW Main Seat
Europe/Vienna timezone

From Content to Code and Back: Evaluating Agentic Coding in LLMs for Building R Shiny Dictionary Interfaces

1 Oct 2026, 14:00
30m
01 | Johannessaal (01 | Johannessaal, OeAW Main Seat, 1st floor)

01 | Johannessaal

01 | Johannessaal, OeAW Main Seat, 1st floor

Austrian Academy of Sciences Dr. Ignaz Seipel-Platz 2 1010 Vienna

Speakers

Iván Arias-Arias (Universidade de Santiago de Compostela) María José Domínguez Vázquez (Universidade de Santiago de Compostela) María Teresa Sanmarco Bande (Universidade de Santiago de Compostela) Carlos Valcárcel Riveiro (Universidade de Vigo)

Description

While recent scholarship highlights the generative capabilities of LLMs for drafting lexicographic entries, their “black box” nature poses challenges regarding training corpora, linguistic representativity, and semantic hallu-cination. We propose that this opacity is less problematic when LLMs are deployed not as linguistic experts, but as software engineers. By leveraging LLMs as agentic coding tools, lexicographers can bridge the gap between linguistic data and digital publication, democratizing access to custom lexicographic tools. This study investigates how LLMs conceptualize “user-friendliness” when tasked with designing lexicographic interfaces. We evaluate three leading models in agentic coding – Claude Opus 4.6, ChatGPT-5.5, and Gemini 3.1 Pro – using a zero-shot prompting strategy within a case study on noun valency. Each model generates an R Shiny application from curated CSV data, accompanied by explicit justifications for its design choices. The generated interfaces are assessed against criteria fundamental to e-lexicography: access structures, visual design, information architec-ture, and alignment with user-centred principles for language learners. Results reveal substantial differences across models in how they conceptualize and operationalize usability, with each model emphasizing fundamen-tally different design aspects. Crucially, not all models prove equally reliable as agentic coders for generating lexicographic interfaces that meet professional standards, highlighting the need for critical evaluation before adoption.

Presentation materials

There are no materials yet.