Ars Inquirendi 2026 · 20–22 November · St Edmund Hall, Oxford, and online
The programme is provisional and may change slightly. All times are Oxford time (GMT). Talks are pre-recorded and released a week ahead, so that the live sessions can be given over to discussion; live keynotes and workshops are labelled. Abstracts open beneath each title.
The pre-print-era graphosphere: the totality of what was written, daubed, etched and carved in a culture before movable-type printing spread through it. Come and hear about the unprecedented power that LLMs are giving us to explore and depict that totality, and about the attendant challenges. Not least, machines have begun to transcribe manuscript images en masse; how far that holds across scripts and languages with few exemplars is one of the day's open questions.
09:15–09:30
Live In-Person
09:30–11:30
Workshop
Live Workshop In-Person
First law-facing contribution to the programme; plausible angles include evidence and proof standards, or legal-historical records as corpus.
13:00–16:15
Panel discussion
Charting the world before print as a totality, from many traditions and many witnesses: writing above all, but also material remains, environments, and whatever holds across time in language, mind and body.
Chair: Roger Martínez-Dávila (University of Colorado & Plus Ultra Collective)
What records are available to LLMs to reveal the world before print? Working across written evidence, languages, and any other forms of witness, the panel asks how surviving material becomes usable evidence, both within traditions and across them, considered as a totality or Gestalt. A proper mapping would show, within a culture, what was written, where and when, and how much of it survives; across cultures, it would allow the comparisons that only a totality permits, from the travels of scripts and genres to the shape of loss itself. Machine-Scale History, on Saturday, comes at the same problem from the other side: not what the evidence is, but what can be inferred from it.
Pre-Recorded Talk Online
In the Regional Great Russian Dictionary of 1852, красота ('beauty') denotes a bride's ribbon, placed by the priest into the Gospel book during the wedding rite. In the Suprasl Codex of the 10th–11th centuries, its root renders Greek κόσμος 'ornament, order, universe'. The familiar aesthetic sense—'attractiveness'—is a late and derived development. Historical lexicography preserves this stratigraphy; large language models, trained overwhelmingly on post-print text, may not. This pilot study tests whether LLMs can read pre-print-era dictionary definitions without projecting modern semantics backwards.
The material comprises entries for красота in eleven Church Slavonic and historical Russian dictionaries documenting the language of the pre-print era—from the earliest Slavonic manuscripts of the 10th–11th centuries onwards—in editions from the eighteenth to the twenty-first century. The pre-print past becomes machine-readable only through this later lexicographic mediation, which is precisely where models may substitute their own training distribution for the historical record. A completed manual componential and Greek–Slavonic analysis of these sources serves as the gold standard: humanistic scholarship as the evaluation benchmark for AI. Definitions are presented to frontier models under three conditions—unattributed, source-attributed, and with an explicit question about the relation to modern meaning—separating inference from the text, recognition of the source, and knowledge of later semantics. Errors are classified as epistemic upgrades, anachronisms, fabricated components, fabricated Greek glosses, and invented continuity narratives.
Two single-model series (42 responses) show fabrication migrating to the continuity condition: the model refused to reproduce unseen dictionary content in direct probes, read the authentic regional label correctly, yet invented two contradictory expansions for its corrupted variant, produced an anachronistic attribution, and inserted a nonexistent word into a Septuagint quote. The paper proposes an anti-fabrication protocol for LLM-assisted work with historical lexicography—a contribution to defending pre-print scholarship against dangerously plausible AI output.
Pre-Recorded Talk In-Person
This paper presents an experiment in using a commercially available frontier LLM to extract, normalize, and critically verify Slavic toponymic evidence for Livonia and its adjoining eastern Baltic borderlands before 1700. The corpus belongs to a transitional manuscript and early-print ecology, in which Latvian and Estonian vernacular textual traditions remained comparatively sparse while German, Latin, Polish, Ruthenian, and Russian documentary practices intersected.
The Slavic evidence comes principally from three early modern traditions—Middle Russian, Polish, and Ruthenian—with a smaller Old East Slavic chronicle layer. Place-names survive in chronicles, legal acts, boundary descriptions, maps and atlases, and later editions of manuscript material. The workflow combines candidate extraction, scholarly transliteration, separation of attested historical forms from modern translations, geographical filtering, duplicate detection, and comparison with an existing historical-toponymic database. The current working table contains 538 attestations and preserves uncertain identifications, competing normalizations, and editorial decisions.
A preliminary normalized-Levenshtein analysis asks whether orthographic affinity between Slavic forms and German or modern local comparison forms varies geographically and by source tradition. Aggregated by identified place, Slavic forms are significantly closer to German comparison forms in Estonia (paired p=.003) and the Daugava corridor (p=.028), while northern Latvia shows no comparable asymmetry. The balance between German and local affinity also differs significantly among Middle Russian, Polish, and Ruthenian groups (Kruskal–Wallis p=.004), with Polish forms showing the strongest German affinity. The smaller Old East Slavic layer is treated separately.
The paper examines LLM failure modes including hallucinated identifications, over-inclusion, confusion between ethnonyms and toponyms, and false equivalence between historical and modern forms. The LLM serves as an interface for querying a fragmented multilingual archive under explicit linguistic control.
This paper discusses an in-progress project to create a list of works cited in Arabic books in the period 700-1800. The goal is to create a list for every book in the OpenITI corpus (8,810 Arabic books, exceeding 1 billion word-tokens) detailing citations, including of titled books and more ephemeral texts (such as notes). Such compiled lists will function as new metadata to accompany the books that survive today in the OpenITI corpus, and will also contain further information per work cited, including authorship attribution, where possible; frequency of citation in the OpenITI book; locations within the OpenITI book of the work citations; and judgments about the certitude of identifications. These lists can be used to address crucial questions about the Arabic tradition as a whole, such as:
The project will be iterative — producing, reproducing and improving the lists over time. All items on a list should be considered candidates, requiring subsequent scholarly judgment. Datasets will be released with the OpenITI corpus and can be ingested into the KITAB/OpenITI web application (kitab-project.org/explore) and other applications.
The method is a pipeline that teaches a model to recognise work references, treating detection — does this span refer to a specific written work? — separately from identification of which work it is, and evaluating the two separately. The pipeline will be discussed in the presentation, and likewise, the contribution that conversations with Claude made to its development.
Pre-Recorded Talk Online
Xuelong Li, with Zheng Liu, Tieshan Zhang & Yu Weng
Pre-Recorded Talk Online
The Dunhuang manuscripts are among the richest sources for the written cultures of medieval Central Asia and the Silk Road, yet more than ninety percent survive only as fragments scattered across collections worldwide. For a century, rejoining them has relied on the chance encounters of scholarly memory. We turn reassembly into a reproducible computational process: it first perceives what survives on each fragment, then synthesises the pieces into whole pages, with a large language model proposing layouts and scholars making the final decision. We first build a perception layer to make fragments machine-readable. Boundary-geometry matching identifies sibling fragments; patch-level handwriting recognition groups leaves by scribal hand; a codebook of 512 learnable visual primitives captures fine-grained style; and glyph-augmentation networks restore damaged characters into readable form. Where material is lost, a diffusion-based simulator, Fate Twin, generates plausible degradation paths, so each reassembly decision rests on simulated evidence rather than intuition. From these fragment-level signals, SRP performs global reassembly. It fuses fragment images, edge maps and OCR text to predict adjacency and relative direction, then a large language model reasons under placement constraints to reconstruct whole-page layouts. Reassembly is thus lifted from local pairwise matching to constrained whole-document inference. Every proposed placement is then scrutinised by domain experts, who make the final determination. Our methods extend beyond Chinese manuscripts to the endangered Khotanese script, for which we have assembled a dataset of 256 character classes and 201,452 images through self-supervised contrastive learning and iterative clustering. This human-machine collaborative approach offers a reproducible and generalisable path toward the digital reassembly of scattered manuscripts worldwide.
Pre-Recorded Talk Online
The Saharan manuscript tradition inverts misguided intuitions about "low-resource" regions of intellectual production. In fact, the Sahara is a region of extraordinary manuscript wealth and scholarly production, with thousands of volumes across family libraries and a nomadic civilization rooted in mobile knowledge transmission. Notwithstanding, the Saharan archive remains minimally catalogued, scarcely digitized, virtually unattested in LLM training corpora. This poverty of learning is the machine's, not the archive's. Drawing on in situ research on Saharan intellectual history, this paper reports two lines of inquiry conducted with frontier models, set against an exemplary archive of digitized microfilms at the University of Illinois Urbana-Champaign.
I begin with a stress test. Trained overwhelmingly on printed and born-digital text, models probed on reception history, scholarly networks, and bibliography in this tradition fail in a characteristic way: fluency inversely tracks reliability. They produce inauthentic if mimetic Islamicate prose, plausible-enough transmission chains, and confected dialectic precisely where training data is thinnest. In a field of plausible-sounding lineages and attributions, deadly nonsense can escape notice. The philologist's inherited disciplines of source criticism and transmission-evaluation turn out to be the working method for using these systems at all, after all.
Moreover, the Stewart microforms of Mauritanian mss at Urbana-Champaign, now digitized and with (occasionally) attributed hands, make a paleographic pilot testable. Can vision-capable models, taught explicit diagnostic criteria, triage manuscripts by regional script type (Andalusī, Maghribī, ṣaḥrāwī/shinqīṭī)? Harder still, can they sort unattributed leaves by scribal hand better than chance, at measurable levels of confidence, against adversarial cases like teacher-student stylistic continuity? What becomes findable matters: women's copying and annotation, legal-network geographies, visual graphs of modal logic applied to theology. I present this as experiment design and early probing, not results, offered by a scholar still learning how to put models to work beyond the chat window, and arguing that neglected traditions are the true test of whether a pre-print AI ecosystem serves the whole pre-print world or only its well-digitized provinces.
Pre-Recorded Talk In-Person
How do you apply universal, categorical, and computationally inspired annotation guidelines to language varieties that are historical, fragmented, and abound with variation? This is the question we, the PARSEME Ancient Greek team, have been faced with for years. In a nutshell, we use an existing Universal Dependencies treebank and make manual adjustments where necessary for the task at hand. We then enhance this treebank by means of manual annotation based on the PARSEME 2.0 universal annotation guidelines for multi-word expressions (MWEs). MWEs are expressions made up of multiple words, such as in front of or to make a suggestion. For PARSEME, words are orthographically defined; linguistically, the question is more complicated (see e.g. Taylor 2014; Dixon & Aikhenvald 2021; Haspelmath 2023). The PARSEME 2.0 guidelines are universal, in that they provide comparative concepts which can be applied to a range of language varieties (cf. Haspelmath 2010). Decision-trees based on the hypothesis that semantic idiomaticity goes hand in hand with morpho-syntactic inflexibility ensure replicability. Thus, diversity can be measured inter-lingually, but what about intra-lingual diversity? This is where even high-resource corpus varieties like classical Greek (5th c. BCE) and Latin (1st c. BCE to 1st c. CE) have challenged us (e.g. Fendel 2025). From sampling, through finding the so-called neutral form of each MWE token, to assessing the (in)flexibility of this neutral form during annotation by means of corpus queries, we have had to adapt (see Fendel, Squeri & Platanou 2026). As an illustrative example, I will draw on the Shared Task 2025 GRC data (classical literary Attic Greek courtroom oratory) (cf. Savary et al. 2026) and the UniDive WG1 Sub-task 1.6 Latin data (classical literary Latin historiography). Both datasets show a significant skew in the MWE tokens towards verbal MWEs rather than nominal, adverbial, adjectival, or functional MWEs. Both samples also have faced us with significant issues when distinguishing between categories of MWEs and between MWEs and fully compositional, flexible structures. The skew towards verbal MWEs can be explained diachronically in the languages’ history. The difficulties in distinguishing between structures can be explained not only synchronically in each language but also based on the annotation process. Finally, I will raise open questions that relate to future work we are planning.
16:30–18:30
Panel discussion
Recent progress in LLM-assisted online editions and machine transcription raises particular challenges and opportunities for scholarly presentation. How can tools that assemble and compare available print editions meet the needs of manuscript traditions? And how can one adequately represent the wealth of raw data that LLMs are enabling?
A report on a UPenn fellowship (August 2026) assessing whether AI/HTR techniques can help read the Medici account books — repetitive, historically rich documents that are also very hard to decipher.
The visual/spatial archive line: maps as pre-print documents that are image, text and worldview at once.
Badr Hamed Al-Harbi , with Abdulaziz Suleiman Al-Harbi
Pre-Recorded Talk Online
Large Language Models are becoming a common point of entry for students and researchers working with premodern Islamic texts. They can locate a reported saying, explain a legal or theological position, and suggest a source within seconds. But an answer may be broadly correct while its evidence is not. A quotation can be slightly altered, a statement assigned to the wrong scholar, or a convincing reference given to a text in which it does not actually appear.
This paper examines this problem through a small comparative experiment focused on source fidelity. A set of questions will be selected from premodern Islamic materials in hadith, law, theology, and Qur'anic interpretation and submitted to several widely used language models under the same conditions. The references will first be checked against the relevant primary sources, providing a basis against which the models' answers can be assessed.
The analysis will distinguish between different kinds of reliability. Does the model give the right information? Does it attribute a statement to the right person or school? Does the cited work exist, and does it actually contain the material attributed to it? When the model presents words as a quotation, how closely do they correspond to the source? Particular attention will be given to references that sound credible but cannot be verified.
These questions matter especially for Islamic textual traditions, where the provenance of a statement is often inseparable from its scholarly value. The study therefore suggests that methods familiar from textual criticism and source verification can also help us evaluate AI-generated scholarship. Rather than treating accuracy as a single measure, it proposes source fidelity as a distinct category for assessing the use of LLMs in the study of premodern texts.
Pre-Recorded Talk Online
The Pseudo-Isidorian Collectio Decretalium (c. 830–850) is the pre-print era’s most ambitious enterprise of forgery: a vast canonical collection interweaving authentic and fabricated material, from conciliar acts to papal letters. Its opening part consists of letters attributed to the thirty pre-Nicene bishops of Rome, from Clement I to Miltiades – fabricated in their entirety, stitched from thousands of authentic biblical, patristic and legal excerpts to sound credible to contemporaries and posterity. This section – the field of my current research – most fully lays bare the workshop of a forger who deceived readers for seven centuries. The corpus offers a unique laboratory for one question: how can machines that generate plausible text today assist in studying plausible text “generated” in the ninth century?
The paper demonstrates a pilot LLM-assisted source inventory on a sample of the corpus: automatic detection of biblical quotations and allusions, distinguishing Vulgate wording, Vetus Latina readings and borrowings mediated by the Fathers, alongside patristic and canonistic excerpts. Results are validated against Hinschius’s apparatus and Karl-Georg Schon’s Pseudoisidor transcriptions, measuring precision and recall and typologizing errors. I will show where the model outperforms the traditional toolkit (scale, unflagged quotations, paraphrase) and where it fails (hallucinated attributions, mistaking a quotation’s intermediary for its source).
The demonstration is paired with a hermeneutic reflection: medieval forgery as a mirror of today’s anxieties about synthetic text. Pseudo-Isidore proves that dangerously plausible text is no invention of LLMs – and philology has long possessed the tools for unmasking it, worth translating into standards for working with language models. I close by situating the method within a wider pre-print ecosystem – retrieval over editions, verification on MDZ scans, the Clavis Canonum – as a working model for the lone scholar without programming skills or grant infrastructure.
LLMs are opening up the reconstruction of lives, networks and societies at dramatically new scales, bringing closer the visions of macro-historians such as Peter Turchin. However, their access to the pre-print past is often filtered through later editions, translations and historical interpretations, and material and environmental evidence matter too. Machine-Scale History asks not only what can be inferred at scale and how such inferences are tested, but also how the field might be transformed by mining the world's untapped primary evidence about the world before print.
09:00–11:00
Workshop
Live Workshop In-Person
"Vibe-coding", the practice of directing large language models to write, run, and improve code through natural-language instruction rather than active human programming, has moved quickly from a curiosity to a genuinely viable research method. This workshop teaches medievalists, whatever their coding background, how to vibe-code effectively and safely, and to recognise when the technique may give trustworthy results and when it will not.
A central question of responsible use: when can we trust the outputs of vibe-coded tools? The answer is predominantly one of verifiability — tasks with an easily checkable output, such as valid TEI-XML encodings automatically produced from a manual transcription, are far safer territory than tasks whose correctness cannot easily be confirmed. The practical core covers tool and environment selection, precise prompt specification, planning phases, architectural pre-decisions, incremental testing and review, and safe version control — consolidated through a guided project building a simple TEI-encoding web application.
13:00 – up to 18:45
Panel discussion
LLMs are opening up the reconstruction of lives, networks and societies at dramatically new scales, bringing closer the visions of macro-historians such as Peter Turchin. However, their access to the pre-print past is overwhelmingly filtered through later editions, translations and historical interpretations, as is material and environmental evidence. Machine-Scale History asks not only what can be inferred at scale and how such inferences are tested, but also how the field might be transformed by mining the world's untapped primary evidence about the world before print.
David Zbíral, with Zoltan Brys, Robert L. J. Shaw & Gideon Kotzé
Pre-Recorded Talk In-Person
The contestation of religious authorities – direct verbal challenges to ecclesiastical legitimacy, clerical conduct, and sacramental power – is dispersed thinly across medieval inquisition records and overshadowed by other topics. Scholarship recognizes its presence in dissident milieus but has never quantified its frequency or examined the social and temporal factors shaping its expression. This dispersal across thousands of documents in dozens of registers made systematic analysis technically unreachable for unassisted scholarship.
We deployed Anthropic's Claude Sonnet 4 to classify and extract relevant passages from 4,357 testimonies spanning 20 inquisition registers (South-Western France, North-Central Italy, Switzerland, England; 1243–1522). LLM classification was validated against independent human coding on random samples of 200 testimonies per variable, achieving 85–99% agreement. We then examined whether four factors correlate with contestation frequency: religious culture, urban versus rural setting, temporal period, and gender.
Results reveal distinct patterns: reformistic dissidents (Waldensians, Beguins, Lollards) articulated authority contestation much more frequently than separatistic groups (Cathars, Apostles, Guglielmites). Urban-dominated registers paradoxically contain fewer such contestations. Temporal analysis demonstrates genuine growth in authority contestation toward the Reformation, affirming expanding lay engagement with ecclesiastical legitimacy. Gender showed no significant effect.
This study illustrates how LLM-assisted extraction, when grounded in rigorous prompt design and human validation, transforms what was previously dispersed beyond scholarly reach into a dataset amenable to quantitative cultural-historical analysis.
Eltjo Buringh, with Auke Rijpma
Pre-Recorded Talk In-Person
A database of some 75,000 picture descriptions from nearly 11,000 European manuscripts (800–1600), used to estimate the prevalence of medieval capital goods. Deliberately non-LLM, manual-extraction work — the baseline the machine-assisted programmes must beat.
Pre-Recorded Talk Online
Drawing on ongoing research with the St. Augustine Jewish Historical Society, this paper explores how large language models can assist in reconstructing Jewish and converso kinship networks connecting Iberia, Spanish Florida, Mexico City, and the Caribbean during the sixteenth and early seventeenth centuries.
Military musters, inquisitorial proceedings, wills, and genealogical inquiries provide the basis for investigating relationships sustained through marriage, patronage, and officeholding. LLM-assisted transcription, translation, and comparison help identify possible connections across dispersed records, which require verification against original manuscripts. The inquiry distinguishes documented kinship from plausible association, and Jewish ancestry from religious practice or accusation.
Collaboration between historians and volunteer researchers grounds these methods in shared questions of belonging and historical memory. The paper considers how AI can support historical inquiry while preserving uncertainty and the complexity of lives shaped by migration, conversion, and colonization.
The convenor's own talk: QUINCUNX as registered falsification for historical inference — pre-registration, inference/evidence firewalls, kill criteria, synthetic written worlds as calibration (public article: arsinq.com/articles/quincunx/).
Remy Levin, with Daniela Vidart
Pre-Recorded Talk In-Person
We design a method for measuring the risk preferences of agents in the deep past. The method combines a structural model of crop choice as a portfolio allocation with machine-learning prediction of expected crop returns, using historic agronomic and climate data. We estimate county-level risk preferences for the United States and farmer-level preferences in Kansas from 1889 to 1929. More risk averse farmers leveraged less, were less likely to purchase novel WWI Liberty Bonds, and were more likely to participate in local risk-sharing institutions. We show that higher risk aversion predicts slower tractor adoption and farm mechanization during the 1920s.
Xuelong Li , with Zimeng Qu
Pre-Recorded Talk Online
The palace-memorial (zouzhe) system reached its fullest articulation under the Yongzheng emperor (1723–1735): memorials bearing vermilion rescripts (zhupi) circulated between the throne and provincial officials, forming a vast administrative record. Historians have relied on close reading of selected exemplars. We ask what changes when the corpus becomes computable. In a system whose bottleneck was imperial attention itself, these documents record how a state managed overflow: prioritisation, verification, and the reproduction of administrative knowledge. We present a computational-history framework that fuses SikuBERT, a BERT model pre-trained on classical and historical Chinese, with the three-level coding of grounded theory. SikuBERT supplies dense semantic representations of memorial texts; open, axial and selective coding organises these into historically meaningful categories, from genre and subject to the tone and urgency of imperial feedback. The two components run iteratively, not sequentially: codings are checked against model outputs and re-derived from the representations, keeping distant reading accountable to close reading. The framework makes systematic analysis of massive historical collections tractable, and keeps interpretation firmly in the historians' hands, replacing exemplar-driven narrative with corpus-scale evidence. Applied to the Yongzheng corpus, the framework traces how information was generated, fed back and reproduced under the memorial system, offering new empirical evidence for information behaviour in Qing state governance. The framework is designed to extend across reigns and document types; we outline how large language models will enter the pipeline next, through assisted coding, corpus-wide hypothesis generation, and dialogue with the archive. The memorial system thus offers an instructive precedent for the computational study of pre-print cultures: a state-run information regime whose “language model” was the emperor himself.
Pre-Recorded Talk In-Person
Medieval genealogy is exceptionally vulnerable to the structure of its archive. Surviving documents privilege inheritance, title, property and institutional continuity, while cadets, cousins, wives, clerics, household associates, witnesses, patrons and other lateral actors frequently disappear from conventional narratives. Later pedigrees can therefore impose linearity upon relationships that may once have been experienced as far more distributed communities of kinship, obligation, patronage and memory.
This paper asks whether generative AI can make a longstanding but previously impractical form of historical inquiry more achievable: reconstructing relationships dispersed across people, places, archives and documentary traditions without collapsing possibility into proof. Developed through The Invisible House, a study of Ridel/Rudel networks across eleventh- and twelfth-century western France, southern Italy and England, the method treats the aristocratic family not simply as a succession of inheritors but as a distributed social system.
An LLM-assisted research environment interrogates charters, cartularies, editions, prosopographical databases, spatial relationships and scholarship iteratively and at scale, allowing weak connections between otherwise separated actors and archives to become visible. Rather than treating nominal similarity as evidence of identity, the method distinguishes source assertions, relational inferences and identity hypotheses, preserving the epistemic status of each claim. Generated hypotheses must survive human source inspection and attempts at falsification before entering the historical argument.
The paper argues that generative AI's most productive role in medieval prosopography is therefore not automated genealogy, but an interrogative layer between archive and scholar. By recovering overlooked actors and relationships, it may also expose the limits of later pedigree traditions themselves: beneath the linear genealogies through which medieval families were subsequently remembered lie older, dispersed networks whose surviving traces suggest that kinship, identity and family memory followed more complicated paths.
19:45
Social
Who controls the reading? Sunday asks how scholars keep authority over models they did not build: whether to train models centred on historical sources or to feed the frontier with primary material; how AI should acknowledge, and answer to, the human scholarship it draws upon; and how to work with agents without surrendering judgement.
09:00–11:00
Workshop
Live Workshop In-Person
An update of his 2025 workshop, reflecting a year of rapid movement on fine-tuning LLMs for pre-modern studies — from understanding the instrument to constructing one.
11:15–12:45
Panel discussion
How should we make LLMs better suited to pre-print research? Should we build models centred on historical sources, increase the primary material available for frontier-model training, or combine the two? This panel considers where scholarly effort will make the greatest difference, from digitisation and transcription to model development, and what these choices mean for accuracy, access and scholarly control. That general models can now transcribe manuscripts raises the incentive to scan them, and the stakes of the choice: the readable archive can grow quickly, but on models that scholars neither own nor control.
Offered: an overview of the latest web version of the Polyscriptor handwritten-text-recognition tool; and/or a demonstration of local agentic harnesses and open-weights models hosted on university servers, avoiding commercial cloud-based models.
2025 alumnus ('Agentic AI and Homoiconic Coding'). Natural kin to any workshop track on building research instruments rather than applying them.
13:15–14:30
Roundtable
Katryna Peart's roundtable opens the final afternoon. Using concrete pre-print cases, it examines how archival survival, digitisation and model behaviour shape scholarly inference, what prompting can and cannot establish about those losses, and which safeguards scholars should apply before carrying AI-generated claims into research or teaching. Its three anchor questions — which layer of the compression chain sheds fidelity, how models inherit archival silences, and what a pre-aggregation diagnostic would look like — then feed the conference-wide discussion that follows.
Scholarly sovereignty is the roundtable's second theme: who sets the questions, understands the results and receives credit; how responsibilities, training and career opportunities should change; and how the sources and people behind a model-assisted result can remain visible.
Live Roundtable Online
The mechanisms by which pre-print scriptoria controlled knowledge — selecting which texts survived, authorizing certain traditions over others, enforcing Latin as the language of institutional truth — are not historical curiosities. They are the structural antecedents of how AI systems trained on digitized pre-print materials currently behave.
Drawing on the Cumulative Compression Model — a four-layer framework mapping how institutional records lose fidelity before any AI system ever touches them — and on empirical findings from a structured adversarial testing study of four major AI systems across contested historical case studies, this roundtable invites medievalists to examine how the mechanisms of pre-print archival control are reproduced, amplified, and rendered invisible in AI-generated scholarly inference.
Designed for conversation rather than presentation: participants bring their own encounters with AI error, synthetic confidence, and archival compression from their domains.
15:00–17:00
Roundtable
A structured open discussion among all the conference's presenters.
Chair: Sarah Bowen Savant (Aga Khan University)
As AI agents let scholars undertake more technical work themselves, how should collaboration with computer scientists, research software engineers and other researchers change? This concluding panel considers how teams can share tools, organise independent checks and sustain the expertise on which their work depends.
17:00
Points of agreement, unresolved questions and next steps from all three days, brought together in the final fifteen minutes.