Ars Inquirendi

← All conjectures · Philology, editions & stemmatics

Sogdian Under the Skin

Status: Anticipated in print · untested

Status is derived only from the shepherd-authored triage/prediction data above -- community submissions and claims are a separate overlay and can never change it (see the participation panel below).

This is a conjecture imagined by a language model — drawn from its trained weights and held to falsifiability, novelty, and plausibility, not to any one method: it may join two or more fields, or none. It is not an article and not evidence: it sits below the evidence/publication boundary. A quantitative prediction and a named kill-dataset are attached (when registered) so the claim stays falsifiable rather than merely evocative.
“Anticipated in print” means the direction of this claim already appears in scholarship, while this exact test has never been run — it records prior art, not proof; nothing here counts as supported or falsified until a registered resolution says so.

Claim (verbatim)

A community's first missionaries leave fingerprints in its deepest vocabulary that later, larger contacts cannot dislodge. Old Uyghur Buddhism of the 9th-14th centuries took most of its texts from Chinese, yet its core technical lexicon - nom for dharma (Greek nomos through Sogdian nwm), chshapt for shiksapada, nizvani for klesha, shmnu for Mara, hormuzta for Indra - wears Sogdian phonology, because the Turks' first Buddhism (8th-10th centuries) came through Sogdian carriers whose terminology fossilized before the Chinese translation wave. Restated as prediction: a loan-etymology census of high-frequency Uyghur Buddhist terms will show Iranian-mediated shapes dominating even inside texts translated from Chinese originals - the bridge language over-represented far beyond its share of the surviving manuscripts.

Prediction clause (verbatim)

Prediction: among the 40 highest-frequency Buddhist technical terms attested in the Uigurica corpus, at least 50% show phonological shapes derivable only through a Sogdian (or other Iranian/Tocharian) intermediary rather than directly from Sanskrit or Chinese, and this share exceeds the Sogdian fraction of Turfan Buddhist manuscripts by a factor of at least 3; coverage guard: at least 30 of the 40 terms must carry established etymologies, else inconclusive.

Kill-dataset (verbatim)

Kill (not yet built): an etymology table over the glossaries of F. W. K. Mueller's Uigurica I-III (Berlin, 1908-1922, archive.org), classifying each high-frequency doctrinal term's donor path as direct-Sanskrit, direct-Chinese, or Iranian-mediated, and counting the proportions.

Nobody has run this test. The kill-data is named above. If you can run it — or you know the paper that already settles it — claim the kill or submit the prior scholarship. Kills and prior scholarship are credited here, by name, as they come in.

On Inferpedia

This conjecture is linked to the following pages on Inferpedia, an encyclopedia of the missing — working atlas pages, some still early scaffolding.

Provenance

Run: Fresh agent generation · model: claude-fable-5

Fresh blind generation by claude-fable-5, 2026-07-20, for the connected-Old-World (movement-signatures) wave. The sixteen items lean on digit-string error phylogeny, first-translation-lag profiles (backlog cliffs and two-pulse corridors), bridge-language fingerprints (Sogdian, Aramaic-chancery, Judeo-Arabic), attribution drift, and pulse/directional asymmetries, with five anti-connection wagers (ordinals 2, 9, 13, 14, 15) where independence or non-transmission is the predicted outcome.

Novelty / leakage triage

anticipated in the literature — this exact test has never been run

That the core doctrinal lexicon of Old Uyghur Buddhism wears Sogdian phonology even inside texts translated from Chinese originals is firmly established Turcology: the examples in the claim (nom, chshapt, nizvani, shmnu, hormuzta) are the standard set and are correctly described, and the Sogdian-first versus Chinese-first debate over the origins of Uyghur Buddhism is explicit in the literature. What has not been done, to my knowledge, is the pinned census: a frequency-ranked forty-term etymological table over the Uigurica glossaries with an over-representation ratio computed against the Sogdian share of Turfan Buddhist manuscripts. The bridge-language over-representation number is therefore new even though the qualitative fingerprint is settled, and the by-3x manuscript-share comparison is a genuine measurement rather than a restatement. Adjacent.

Sources cited by the triage
  • F. W. K. Muller, Uigurica I-III, Abhandlungen der Preussischen Akademie der Wissenschaften, Berlin, 1908-1922. — The pinned corpus; its glossaries already note the Sogdian forms.
  • Annemarie von Gabain, Altturkische Grammatik, Leipzig, 1941 (3rd ed., Wiesbaden, 1974).
  • Jens Wilkens, “Buddhism in the West Uyghur Kingdom and Beyond,” in Carmen Meinert (ed.), Transfer of Buddhism Across Central Asian Networks (7th to 13th Centuries), Brill, 2016. — Surveys the Sogdian-mediated terminology and the origins debate.

Predictions

No prediction registered yet.

Weigh in

No community feedback yet.

New here? Create an account first

Create an account or sign in and your feedback is tied to you — you can track it, get replies, and claim this conjecture so others know you’re working on it. Prefer not to? Just leave your take below as a guest — only the name you type is shown.

Add your take

Posted immediately (spam is removed). Community feedback is never an adjudicated verdict and never changes this conjecture's triage label or status above.