Ars Inquirendi

Ars Inquirendi 2026 · your paper

Work References across the OpenITI

Sarah Bowen Savant (Aga Khan University)

Your slot

Strand

Mapping the Old World Graphosphere, and Beyond

Day

Day 1 · Friday 20 November 2026

Part

Afternoon — exact times to follow

Role

Talk

Participation

In person

Keynote of the strand

Marieke Meelen

Chair

Roger Martínez-Dávila

Abstract

This paper discusses an in-progress project to create a list of works cited in Arabic books in the period 700-1800. The goal is to create a list for every book in the OpenITI corpus (8,810 Arabic books, exceeding 1 billion word-tokens) detailing citations, including of titled books and more ephemeral texts (such as notes). Such compiled lists will function as new metadata to accompany the books that survive today in the OpenITI corpus, and will also contain further information per work cited, including authorship attribution, where possible; frequency of citation in the OpenITI book; locations within the OpenITI book of the work citations; and judgments about the certitude of identifications. These lists can be used to address crucial questions about the Arabic tradition as a whole, such as:

  1. Authorial practices. For example, which works, and which types of works, do book authors cite most frequently? For which authors do we have the most citations – and do they cite differently in different works that they wrote?
  2. Survival. The book citations can be compared to titles in the OpenITI corpus, as well as to historical and modern book lists and catalogues. To what extent is the OpenITI corpus itself representative of the works its authors cited?
  3. Lost texts. There are many texts embedded in later texts, for which we have no surviving independent witness. The titles and authorial attributions can assist scholars now working on retrieval of passages in their proximity.

The project will be iterative — producing, reproducing and improving the lists over time. All items on a list should be considered candidates, requiring subsequent scholarly judgment. Datasets will be released with the OpenITI corpus and can be ingested into the KITAB/OpenITI web application (kitab-project.org/explore) and other applications.

The method is a pipeline that teaches a model to recognise work references, treating detection — does this span refer to a specific written work? — separately from identification of which work it is, and evaluating the two separately. The pipeline will be discussed in the presentation, and likewise, the contribution that conversations with Claude made to its development.

Subjects

Islamicate & AfricaArchives, libraries & lossOld World written cultures

What we need from you

How the conference works

All talks except the live keynotes are pre-recorded and released a week ahead of the conference, by 13 November 2026.

The live sessions are discussions of the pre-recorded talks, held in the room at St Edmund Hall, Oxford, and online.

Speakers do not register or pay.

Invitation or visa letters, and accommodation questions: write to arsinquirendi@gmail.com.