Ars Inquirendi

← All conjectures · Central & Inner Asian texts

A script born old

Status: Already answered

Status is derived only from the shepherd-authored triage/prediction data above -- community submissions and claims are a separate overlay and can never change it (see the participation panel below).

This is a conjecture imagined by a language model — drawn from its trained weights and held to falsifiability, novelty, and plausibility, not to any one method: it may join two or more fields, or none. It is not an article and not evidence: it sits below the evidence/publication boundary. A quantitative prediction and a named kill-dataset are attached (when registered) so the claim stays falsifiable rather than merely evocative.

Claim (verbatim)

The Tangut script was promulgated at a stroke by imperial decree in 1036 under Weiming Yuanhao — tradition credits the scholar Yeli Renrong with its design — as roughly six thousand logographs built in deliberate visual differentiation from Chinese. Natural logographies are eroded by use: in Chinese, frequent characters tend to be simpler, because centuries of writing economy wear the common signs smooth. A script issued complete, used for under three centuries by a chancery-and-clergy class, and never passed through mass scribal attrition should lack this wear pattern: its complexity should be flat across the frequency spectrum. Prediction: joining per-character frequencies from digitized Khara-Khoto texts to stroke counts, the rank correlation between frequency and stroke count in Tangut will be no stronger than -0.1, while the matched computation on a pre-modern Chinese corpus yields -0.25 or stronger; secondarily, Tangut characters will average at least 25% more strokes than the Chinese characters glossing them in Li Fanwen's dictionary (primary clause: the correlation contrast at the stated thresholds; the verdict follows it). Kill (not yet built): a Tangut frequency-by-stroke table, buildable now from the digitized Khara-Khoto facsimiles in Ecang Heishuicheng wenxian (the Kozlov collection of the IOM, St Petersburg; Shanghai Guji facsimile series) joined to the character data of the BabelStone Tangut database (the dataset behind Unicode Tangut) and Li Fanwen's Xia-Han zidian (1997).

Prediction clause (verbatim)

Prediction: joining per-character frequencies from digitized Khara-Khoto texts to stroke counts, the rank correlation between frequency and stroke count in Tangut will be no stronger than -0.1, while the matched computation on a pre-modern Chinese corpus yields -0.25 or stronger; secondarily, Tangut characters will average at least 25% more strokes than the Chinese characters glossing them in Li Fanwen's dictionary (primary clause: the correlation contrast at the stated thresholds; the verdict follows it).

Kill-dataset (verbatim)

Kill (not yet built): a Tangut frequency-by-stroke table, buildable now from the digitized Khara-Khoto facsimiles in Ecang Heishuicheng wenxian (the Kozlov collection of the IOM, St Petersburg; Shanghai Guji facsimile series) joined to the character data of the BabelStone Tangut database (the dataset behind Unicode Tangut) and Li Fanwen's Xia-Han zidian (1997).

Browse the registry: Dunhuang & Turfan hoards →

How this item is catalogued

Catalogue tags are curatorial metadata that make the conjecture findable. They are not findings, not a novelty verdict, and naming a source family is not a promise that the data exist or are accessible. Each dimension below says whether it was assigned, reviewed with none applying, uncertain, or not yet reviewed, and who recorded it.

Place & era tags — Assigned
Steppe & Central Asia; Medieval Tagged by Claude (Opus 4.8)
Dating — Phenomenon
Dated c. 1030 CE – 1380 CE Tangut script, promulgated 1036 CE, in chancery/clergy use through Khara-Khoto's abandonment (1372); BabelStone/Li Fanwen digitized data are the instrument, not the phenomenon. Tagged by claude-sonnet-5 + claude-fable-5 (gate)
Subject — Assigned
Texts, scribes & transmission Tangut script's complexity untouched by centuries of scribal-use wear Tagged by claude-sonnet-5
Finer subject topic — Reviewed — none applies
the only listed script-death page is the cuneiform one Tagged by claude-opus-5 draft (requested max effort) adjudicated by claude-fable-5-1 driver, lane 1da5; runtime not independently attested
Claim level — Assigned
Record claim the prediction is a frequency-to-complexity relation inside a script's own sign inventory, a property of the writing system and its use Tagged by claude-opus-5 draft (requested max effort) adjudicated by claude-fable-5-1 driver, lane 1da5; runtime not independently attested
Method — Assigned
Information theory: redundancy, error-correction, and compression; The comparative method and natural experiments primary: writing economy is treated as a compression pressure that shortens frequent signs, and the prediction is that a decreed script never subjected to it lacks that coding signature; second: a matched computation on a pre-modern Chinese corpus is the control against which the contrast is read Tagged by claude-opus-5 draft (requested max effort) adjudicated by claude-fable-5-1 driver, lane 1da5; runtime not independently attested
Evidence source — Assigned
Dunhuang & Turfan hoards; Not yet built the digitized Khara-Khoto facsimiles are the text base the family covers, and the kill text states the frequency-by-stroke table is not yet built Tagged by claude-opus-5 draft (requested max effort) adjudicated by claude-fable-5-1 driver, lane 1da5; runtime not independently attested
Research stage — Idea — no test designed

On Inferpedia

This conjecture is linked to the following pages on Inferpedia, an encyclopedia of the missing — working atlas pages, some still early scaffolding.

Provenance

Run: Fresh agent generation · model: claude-fable-5

Fresh blind generation, claude-fable-5, 2026-07-16, breadth wave: under-represented cultures & places (Southeast Asia + Central/Inner Asia), produced from model knowledge; grounded in real works/inscriptions/corpora; no fabricated citations.

Novelty / leakage triage

already answered in the literature

Andrew West (author of the BabelStone Tangut database the conjecture names as its kill-source) has published exactly this analysis: there is no relationship between frequency and stroke count for Tangut characters — normal Tangut text is uniformly composed of characters of about 12 plus-or-minus 6 strokes — whereas high-frequency Chinese characters are disproportionately simple. The primary contrast is his stated result.

Sources cited by the triage
  • A. West, 'How Complex is Tangut?', BabelStone Blog, August 2009 (babelstone.co.uk/Blog/2009/08/how-complex-is-tangut.html)
  • Li Fanwen, Xia-Han zidian (Zhongguo shehui kexue chubanshe, 1997)

Predictions

No prediction registered yet.

Weigh in

No community feedback yet.

New here? Create an account first

Create an account or sign in and your feedback is tied to you — you can track it, get replies, and claim this conjecture so others know you’re working on it. Prefer not to? Just leave your take below as a guest — only the name you type is shown.

Add your take

Posted immediately (spam is removed). Community feedback is never an adjudicated verdict and never changes this conjecture's triage label or status above.