Intertextuality as a Retrieval Task: Benchmarking Text Reuse in Classical and Medieval Latin

Sep 8, 2026, 12:00 PM
20m
Room 3

Room 3

Speaker

Martin Roček (Institute for Medieval Research, Austrian Academy of Sciences and Faculty of Arts, Charles University)

Description

What if we reframe intertextuality detection as a retrieval task? In this talk, I will take quotations from an author and rank them against corpus of possible sources. This allows me to measure recall@10, precision@1, nDCG@10 and compare four different methods that are commonly used: BM25, dense embeddings (generic and Latin-specific), reciprocal-rank fusion, and reranking by cross-encoder or large LLM. Due to the lack of a gold standard for medieval Latin, I used classical Latin (Loci Similes; Jerome and Lactantius against ~90,000 passages), where verified links are available as a published benchmark, and created my own for medieval Latin (Bernard of Clairvaux against the Vulgate), where no such benchmark exists and the reference set has to be assembled from a machine-parsed Patrologia Latina apparatus and a BiblIndex export. Results show that no single method wins across corpora. Fusion of BM25 and an embedding model has potential to give the best coverage on both corpora and reranking lifts an already strong lists. Part of this benchmarking experiment was a discovery run on Bernard. The pipeline proposed 934 candidates, of which 308 were known references independently re-found; of the remaining 626, a machine verifier flagged 166 for reading, and hand review confirmed 28 previously unrecorded scriptural links — a third of them invisible to word-matching, and several of them explicitly introduced, word-for-word quotations recorded in neither Migne nor BiblIndex. Such systems work best as calibrated assistants whose "false positives" are often worth following.

Presentation materials

There are no materials yet.