Ars Inquirendi

Ars Inquirendi 2026 · your paper

Reassembling Dunhuang: Reading Old World Written Cultures at LLM Scale

Xuelong Li, with Zheng Liu, Tieshan Zhang & Yu Weng (Minzu University of China)

Your slot

Strand

Mapping the Old World Graphosphere, and Beyond

Day

Day 1 · Friday 20 November 2026

Part

Afternoon — exact times to follow

Role

Talk

Participation

Online

Keynote of the strand

Marieke Meelen

Chair

Roger Martínez-Dávila

Abstract

The Dunhuang manuscripts are among the richest sources for the written cultures of medieval Central Asia and the Silk Road, yet more than ninety percent survive only as fragments scattered across collections worldwide. For a century, rejoining them has relied on the chance encounters of scholarly memory. We turn reassembly into a reproducible computational process: it first perceives what survives on each fragment, then synthesises the pieces into whole pages, with a large language model proposing layouts and scholars making the final decision. We first build a perception layer to make fragments machine-readable. Boundary-geometry matching identifies sibling fragments; patch-level handwriting recognition groups leaves by scribal hand; a codebook of 512 learnable visual primitives captures fine-grained style; and glyph-augmentation networks restore damaged characters into readable form. Where material is lost, a diffusion-based simulator, Fate Twin, generates plausible degradation paths, so each reassembly decision rests on simulated evidence rather than intuition. From these fragment-level signals, SRP performs global reassembly. It fuses fragment images, edge maps and OCR text to predict adjacency and relative direction, then a large language model reasons under placement constraints to reconstruct whole-page layouts. Reassembly is thus lifted from local pairwise matching to constrained whole-document inference. Every proposed placement is then scrutinised by domain experts, who make the final determination. Our methods extend beyond Chinese manuscripts to the endangered Khotanese script, for which we have assembled a dataset of 256 character classes and 201,452 images through self-supervised contrastive learning and iterative clustering. This human-machine collaborative approach offers a reproducible and generalisable path toward the digital reassembly of scattered manuscripts worldwide.

Subjects

South, East & Central AsiaOld World written culturesHTR & instruments

What we need from you

How the conference works

All talks except the live keynotes are pre-recorded and released a week ahead of the conference, by 13 November 2026.

The live sessions are discussions of the pre-recorded talks, held in the room at St Edmund Hall, Oxford, and online.

Speakers do not register or pay.

Invitation or visa letters, and accommodation questions: write to arsinquirendi@gmail.com.