Ars Inquirendi 2026 · your paper
Victoria Fendel (University of Oxford)
Strand
Mapping the Old World Graphosphere, and Beyond
Day
Day 1 · Friday 20 November 2026
Part
Afternoon — exact times to follow
Role
Talk
Participation
In person
Keynote of the strand
Marieke Meelen
Chair
Roger Martínez-Dávila
How do you apply universal, categorical, and computationally inspired annotation guidelines to language varieties that are historical, fragmented, and abound with variation? This is the question we, the PARSEME Ancient Greek team, have been faced with for years. In a nutshell, we use an existing Universal Dependencies treebank and make manual adjustments where necessary for the task at hand. We then enhance this treebank by means of manual annotation based on the PARSEME 2.0 universal annotation guidelines for multi-word expressions (MWEs). MWEs are expressions made up of multiple words, such as in front of or to make a suggestion. For PARSEME, words are orthographically defined; linguistically, the question is more complicated (see e.g. Taylor 2014; Dixon & Aikhenvald 2021; Haspelmath 2023). The PARSEME 2.0 guidelines are universal, in that they provide comparative concepts which can be applied to a range of language varieties (cf. Haspelmath 2010). Decision-trees based on the hypothesis that semantic idiomaticity goes hand in hand with morpho-syntactic inflexibility ensure replicability. Thus, diversity can be measured inter-lingually, but what about intra-lingual diversity? This is where even high-resource corpus varieties like classical Greek (5th c. BCE) and Latin (1st c. BCE to 1st c. CE) have challenged us (e.g. Fendel 2025). From sampling, through finding the so-called neutral form of each MWE token, to assessing the (in)flexibility of this neutral form during annotation by means of corpus queries, we have had to adapt (see Fendel, Squeri & Platanou 2026). As an illustrative example, I will draw on the Shared Task 2025 GRC data (classical literary Attic Greek courtroom oratory) (cf. Savary et al. 2026) and the UniDive WG1 Sub-task 1.6 Latin data (classical literary Latin historiography). Both datasets show a significant skew in the MWE tokens towards verbal MWEs rather than nominal, adverbial, adjectival, or functional MWEs. Both samples also have faced us with significant issues when distinguishing between categories of MWEs and between MWEs and fully compositional, flexible structures. The skew towards verbal MWEs can be explained diachronically in the languages’ history. The difficulties in distinguishing between structures can be explained not only synchronically in each language but also based on the annotation process. Finally, I will raise open questions that relate to future work we are planning.
Editions & philologyOld World written cultures
All talks except the live keynotes are pre-recorded and released a week ahead of the conference, by 13 November 2026.
The live sessions are discussions of the pre-recorded talks, held in the room at St Edmund Hall, Oxford, and online.
Speakers do not register or pay.
Invitation or visa letters, and accommodation questions: write to arsinquirendi@gmail.com.