Ars Inquirendi

← All conjectures · Jewish book cultures

The code that erased its own sources

Status: Falsified

The verdict’s fine print — quoted from the resolution record: “KILLED BY 6% ON A PINNED INSTRUMENT - the caveats carry what the number means.” Read the full caveats ↓

Status is derived only from the shepherd-authored triage/prediction data above -- community submissions and claims are a separate overlay and can never change it (see the participation panel below).

This is a conjecture imagined by a language model — drawn from its trained weights and held to falsifiability, novelty, and plausibility, not to any one method: it may join two or more fields, or none. It is not an article and not evidence: it sits below the evidence/publication boundary. A quantitative prediction and a named kill-dataset are attached (when registered) so the claim stays falsifiable rather than merely evocative.

Claim (verbatim)

Maimonides built the Mishneh Torah as a self-standing law: fourteen books of ruled halakhah in his own Hebrew, deliberately and programmatically without citing sources - he says as much in the introduction, promising a code one could read in place of the whole prior literature. That choice made the work a citation black hole. Every ruling has an ancestry - Mishnah, the two Talmuds, the halakhic midrashim, and a heavy but silent debt to geonic writings - but the code names none of it, which is why an entire genre of source-hunting (the Maggid Mishneh, the Kesef Mishneh, down to the modern Frankel apparatus) grew up around it, and why some rulings resisted source-identification for centuries: where Maimonides leaned on a geonic responsum or a reading now lost, the code preserves the ruling while destroying the pointer back to it. Mechanism: a source survives partly by being cited, and a code that cites nothing converts its sources into anonymous law, severing the quotation-trail that recovery needs. Prediction restated: the density of explicit external source-attribution formulae in the Mishneh Torah is a small fraction of that in a citing code of comparable scope such as the Tur - the measurable signature of a work that swallowed its own bibliography, and with it the traceability of the lost geonic layer beneath.

Prediction clause (verbatim)

Prediction: counting explicit external source-attribution formulae (naming a tractate, a named sage, or "the Geonim" as authority) per 10,000 words across Sefaria's machine-readable Mishneh Torah versus the Arba'ah Turim (Tur), the Mishneh Torah's rate will be under one-tenth (a ratio < 0.1) of the Tur's - primary clause: the < 0.1 Mishneh-Torah-to-Tur source-formula density ratio; the verdict follows it. Internal self-references ("as we explained in Hilkhot...") are excluded so that only outward source-attribution is counted; the test voids for coverage if either code is under 80% digitized.

Kill-dataset (verbatim)

Kill: Sefaria's open Mishneh Torah and Arba'ah Turim (Tur) - pattern-match external source-attribution formulae, normalize per 10,000 words, and compare densities, with the Frankel-edition source apparatus and the Kesef Mishneh as the control on what the Mishneh Torah left unnamed. Distinct from an internal cross-reference audit: only outward source-attribution is scored.

Provenance

Run: Fresh agent generation · model: claude-fable-5

Fresh blind generation by claude-fable-5, 2026-07-17, Jewish text-culture wave instrument-anchored on the open Sefaria corpus and its cross-reference link data, with standard critical editions as controls: every kill names a real corpus and a countable operation (coverage maps, citation-formula counts, link-orphan shares, digest-fraction, recension divergence, citation-decay), thresholds far from 1 with explicit coverage guards distinguishing what Sefaria holds from what existed. Ground is CITATION-GEOMETRIC and disjoint from the 2026-07-10 w19 Jewish wave, which was material-culture-of-loss (colophons, parchment, genizah, masora, binding fragments): no material-culture re-posing here. Two candidates were dropped after a grep of all fresh packets: a Tosafot-density-by-tractate survival item (pre-empted by w02-philosophy #18, which already correlates per-folio Tosafot density with manuscript survival), and a piyyut liturgy-vs-anthology survival item (pre-empted by w01-literature #28 'Liturgy out-survives fame', keyed to Davidson's Thesaurus). The Mishneh Torah item here tests EXTERNAL source-attribution erasure vs the Tur (loss of source-geometry), deliberately distinct from minds-w02 #25, which tests the code's INTERNAL recall/promise cross-references. Two items are marked Kill (not yet built) where the deciding corpus (Kohut's Arukh apparatus; Lewin's parallel Iggeret recensions) is not yet digitized.

Novelty / leakage triage

anticipated in the literature — this exact test has never been run

The design fact is printed in the primary source itself - Maimonides' introduction announces a code without source citations, the Rabad's stricture attacks the choice, and Twersky's Introduction to the Code analyzes both the program and the scattered exceptions ('the Geonim ruled', named teachers) at length - and the Tur's character as a densely citing code is equally standard. But neither code's source-attribution formulae have ever been counted, the Mishneh Torah's rate is small but not zero, and the < 0.1 per-10,000-word density ratio is a specific figure no reference work states or strictly guarantees. Direction print-established, arithmetic un-run.

Sources cited by the triage
  • I. Twersky, Introduction to the Code of Maimonides (Mishneh Torah) (New Haven: Yale UP, 1980), on the omission of sources and its critics
  • Maimonides, Mishneh Torah, author's introduction (with the hassagah of Rabad ad loc.)

Predictions

Killed registered 2026-07-19 variant: v2-pinned-formula-set calibration prediction (parent triage: leaked/adjacent)

Resolution: Killed

Caveats: KILLED BY 6% ON A PINNED INSTRUMENT - the caveats carry what the number means. (a) THE KILL: R=0.1059 against the ex-ante line of 0.1, a 5.9% overshoot on a wager of a full 10x density separation. This time the formula set was pinned at registration with eyes open (lesson 3, adopted after v1's unpinned set forced an inconclusive), the only registered escape hatch - the 200-match floor guard - did not trigger, and the REFUTED clause fires; the verdict follows it mechanically. Declining to kill on a post-hoc audit of a set the house itself pinned would make every future pin re-litigable the moment a number comes in unfavourable, and would gut pre-registration; the discipline is worth more than this conjecture. (b) NOISE DIAGNOSTIC (disclosed, NO verdict weight): the compute's own audit flags families b/c (construct bet/dalet+tractate-name) as dominated by generic-word polysemy rather than genuine citation - בשבת is overwhelmingly 'on the Sabbath', not 'in tractate Shabbat', and all 10 MT family-c audit samples were generic-sense on inspection - at near-identical ABSOLUTE scale in both works (673 vs 715 matches; 565 vs 530 on the shabbat token alone). Equal absolute noise is proportionally unequal: it is 91% of MT's numerator (673 of 739) but only 11% of the Tur's (715 of 6,379), so it inflates measured R upward AGAINST the conjecture. The unregistered diagnostic excluding families b/c lands at MT ~0.86/10k vs Tur ~81/10k, R~0.011 - deeply under 0.1. A second quantified spec-inherent bias points the same way: lamed-prefixed acronym forms fall outside the pinned 5-letter prefix set (~9% relative undercount on the sampled Rif item; the rule is identical for both works, but its practical incidence is on the Tur side, since MT has zero acronyms to undercount). Readers should take this kill as a fact about the pinned instrument and the 10x wager, not as evidence that the Mishneh Torah genuinely cites at more than a tenth of the Tur's rate - the cleaned diagnostic says it does not. (c) QUALITATIVE PICTURE: the direction is resoundingly right. A 9.4x citation-density gap on the pinned set (90.72 vs 9.61 per 10k); 5,185 Rishonic/Geonic acronym citations in the Tur against literally ZERO in the Mishneh Torah - verified real, not an engine defect: Rambam predates most of the listed Rishonim and wrote in his famously self-contained uncited style, while the Tur cites them by design (including the author citing his own father as the Rosh). The 'code that erased its own sources' signature is plainly present in the data; what died is the specific <0.1 pin. (d) CALIBRATION: another direction-right / multiplicative-overshoot datum for the calibration series, and the closest miss yet - the registrant wagered a 10x multiple and the world gave 9.4x. File it with the standing pattern of correct-direction conjectures killed by overconfident multiplicative thresholds. (e) NO V3 CHASE: any polysemy-cleaned v3 set would need INDEPENDENT justification stated BEFORE any recompute; re-registering now, after seeing that the cleaned diagnostic would support, is threshold-shopping - exactly what lesson 3 exists to prevent. The shepherd recommends leaving this conjecture killed and letting the calibration essay absorb the miss.

RE-REGISTRATION under lesson 3 (pin the full operationalization ex ante). Same conjecture, same threshold; the citation-formula set is now pinned AT REGISTRATION, symmetric across corpora, including the Tur's dominant Rishonic-acronym citation idiom that v1's compute-chosen set missed. Primary clause unchanged: the Mishneh Torah's external source-attribution density per 10,000 words is under one-tenth of the Tur's (R < 0.1).

Resolution criteria — the registered fine print

Resolution criteria: PINNED SYMMETRIC FORMULA SET (counted identically in BOTH works): (a) Rishonic/Geonic acronym-citations: הרא"ש, הרמב"ם, רמב"ן, רשב"א, רשב"ם, רי"ף, ר"י, ר"ת, בה"ג, הראב"ד, בעל העיטור, סמ"ג, רס"ג; (b) construct-state tractate citation: ד + gemara-bearing tractate name (e.g. דברכות, דשבת); (c) ב + tractate name and כדאיתא ב־ + tractate; (d) מס'/מסכת + tractate name; (e) named-sage formulae: אמר רבי/רבי + name-whitelist, רב + name-whitelist (the v1 whitelists); (f) גאון/הגאונים. EXCLUDED: within-work self-pointers (כמו שביארנו; MT's בהלכות X self-references). DATA: Sefaria's 84 MT sub-books + Tur's 4 Turim, Hebrew, HTML-stripped, whitespace tokens (v1's fetch route). R = MT matches-per-10k-words / Tur matches-per-10k-words under the pinned set. CLAUSE PRECEDENCE: (1) INCONCLUSIVE_BY_DESIGN if the pinned set yields fewer than 200 total matches in either work (floor guard); (2) SUPPORTED if R < 0.1; (3) REFUTED if R >= 0.1. Compute firewalled from threshold and direction (it receives the set, not the expectation).

Known-priors disclosure — what the registrant already knew

Known priors disclosure: v1 (2026-07-18) went INCONCLUSIVE by operationalization: raw R=1.06 invalid (topical vocabulary), conservative R=0.1199 with a documented one-sided Tur-undercount bias (acronym citations invisible). With acronyms now pinned in, the Tur numerator should rise substantially; the registrant expects R to fall BELOW 0.1 (supported) but regards it as genuinely open — the MT side also gains some acronym hits (e.g. citations of the Rif).

Method and dataset — how it was measured

Register-before-compute with the FULL operationalization pinned at registration per lesson 3 (adopted after v1 went inconclusive on an unpinned, compute-chosen set): formula families, token whitelists, prefix-letter set, gershayim normalization, self-pointer exclusions, 200-match floor guard and clause precedence all fixed ex ante in conjecture_prediction_batch6_20260719.json (written 2026-07-19T09:51Z). Firewalled compute (claude-sonnet, blind to threshold and direction - it received the set, not the expectation) fetched all 88 text bodies from the Sefaria API and computed at 2026-07-19T10:28:12Z; computed_at postdates registered_at. Matching by precedence-ordered non-overlapping family scan over a shared claimed[] token array; exclusions verified by a with/without counterfactual re-run. Registered clause precedence applied mechanically: (1) floor guard NOT triggered (MT 739 and Tur 6,379 both >= 200 matches); (2) SUPPORTED if R < 0.1 not met; (3) R = 0.1059 >= 0.1, the REFUTED clause fires.

Dataset: Sefaria Hebrew corpora fetched live 2026-07-19: Mishneh Torah as 84 leaf books under 14 Sefer categories (769,174 words) vs the Arba'ah Turim as 4 Turim nodes (703,141 words), identical preprocessing for both (HTML-strip; niqud-strip, essential because MT's default version is vocalized and the Tur's is not; 3-way gershayim normalization). The v2 formula set, pinned IN FULL at registration and applied symmetrically: (a) 13 Rishonic/Geonic acronym items; (b) construct-state dalet+tractate; (c) construct-state bet+tractate incl. the kedeita-be subset; (d) masechet-marker+tractate; (e) named-sage formulae over pinned whitelists (rabbi 40 names / rav 26 names); (f) gaon/hageonim tokens; within-work self-pointers excluded (exclusion logic verified active, zero net effect). Per-family shape: Tur 6,379 total matches dominated by family a = 5,185 genuine Rishonic acronym citations (הרא"ש 1,754, הרמב"ם 1,598, ר"י 335, הראב"ד 320); MT 739 total with family a = 0 and 673 of 739 (91%) in the polysemy-flagged families b/c; family d = 0 in both works, verified rather than assumed.

computed 2026-07-19

Inconclusive registered 2026-07-19 calibration prediction (parent triage: leaked/adjacent)

Resolution: Inconclusive

Caveats: TWO OPERATIONALIZATIONS, ONE INVALID. (1) RAW (all tractate-name tokens + named-sage + geonim): MT 53.16/10k vs Tur 50.23/10k, R=1.0583 - invalidated as a citation measure by the compute's own QC: all 37 Bavli tractate names double as the ordinary Hebrew noun for their own subject matter (Tamid='always', Shabbat=the Sabbath day, Gittin=divorce documents) and both codes are organized topically around exactly those subjects; only 1 of MT's 3,988 raw tractate hits (0.03%) and 50 of Tur's 2,804 (1.8%) carry the construct-state ד- citation marker, and hand-inspected samples confirm the hits are topical vocabulary (every sampled MT 'Tamid' is the adverb 'always' in a cosmology passage). No verdict may rest on the raw reading. (2) CONSERVATIVE (ד--marked tractate citations + named-sage + geonim), the faithful operationalization of 'explicit external source-attribution': MT 1.3261/10k vs Tur 11.0646/10k, R=0.1199. WHY NOT KILLED ON 0.1199 >= 0.1: the number sits 20% above the threshold - within the measurement slop of mechanical string-matching - AND carries a documented ONE-SIDED bias: the Tur's dominant citation idiom, Rishonic surname acronyms and epithets (הרא"ש, הרמב"ם, רמב"ן, רשב"א, רי"ף, בעל העיטור, ר"י), lies entirely outside the pinned formula families and is structurally invisible to them, while MT has almost no such hidden layer to miss (18 named-sage hits and 1 construct citation in 769,174 words); the Tur denominator is therefore systematically UNDERCOUNTED (as are its ב-/כדאיתא and מס' tractate-citation forms, sample-confirmed), so the true conservative R is LOWER than 0.1199, plausibly below 0.1. A kill declared on a number the compute itself certifies as biased upward against the conjecture would be an artifact of the instrument, not a finding about the texts - a weak kill. WHY NOT SUPPORTED: no measured reading gives R < 0.1; support would require an unmeasured bias-corrected extrapolation. REGISTRATION DEFECT AND LESSON: the registration delegated the concrete formula set to the compute ('define a pinned external-attribution formula set') instead of pinning it ex ante; the resulting order-of-magnitude operationalization-dependence (1.06 vs 0.12) is therefore partly a registration failure. Per house precedent (KP-SD, DBBE, Kushan), near-threshold results resting on unpinned operationalizations with known one-sided biases go INCONCLUSIVE with a re-registration path rather than to a verdict. Lesson: pin the formula set, token lists and aggregation rule ex ante, not just the threshold. The registered clause precedence itself ranks 'formula set cannot be applied consistently' (INCONCLUSIVE_BY_DESIGN) ahead of the threshold clauses; mechanically the formula was applied identically to both corpora, but as a measure of the target quantity it captures MT's citation idiom nearly completely while missing the Tur's dominant idiom - an asymmetry of measurement validity in the spirit of clause (1). RE-REGISTRATION PATH: re-register with a pre-pinned, symmetric citation-formula set that INCLUDES the Rishonic acronym/epithet citations (הרא"ש, הרמב"ם, רמב"ן, רשב"א, רי"ף, ר"י, בעל העיטור and peers), construct ד- tractate citations, ב-/כדאיתא + tractate and מס' abbreviation forms, the named-sage formulae, and גאון/גאונים; apply identically to both corpora and recompute R against the same <0.1 threshold. WHAT SURVIVES REGARDLESS: under the faithful conservative reading the Mishneh Torah cites drastically less than the Tur - 1.33 vs 11.06 per 10k words, an ~8.3x gap that the documented Tur undercount can only widen - fully consistent with Twersky's account of Maimonides' programmatic no-citation design and with the conjecture's qualitative mechanism; only the registered <0.1 magnitude pin remains undecided. Calibration note carried from registration: '<0.1 is a very bold pin' was flagged ex ante; the pin nearly held (0.1199) even on an instrument biased against it.

Registered against Sefaria's open Mishneh Torah and Arba'ah Turim (Tur) (verified-real Sefaria indices). PREDICTION VERBATIM: Prediction: counting explicit external source-attribution formulae (naming a tractate, a named sage, or "the Geonim" as authority) per 10,000 words across Sefaria's machine-readable Mishneh Torah versus the Arba'ah Turim (Tur), the Mishneh Torah's rate will be under one-tenth (a ratio < 0.1) of the Tur's - primary clause: the < 0.1 Mishneh-Torah-to-Tur source-formula density ratio; the verdict follows it. Internal self-references ("as we explained in Hilkhot...") are excluded so that only outward source-attribution is counted; the test voids for coverage if either code is under 80% digitized.

Resolution criteria — the registered fine print

Resolution criteria: DENOMINATOR (verbatim): explicit EXTERNAL source-attribution formulae (naming a tractate, a named sage, or 'the Geonim' as authority) per 10,000 words, across Sefaria's Mishneh Torah vs the Tur. INTERNAL self-references ('as we explained in Hilkhot...') EXCLUDED — only outward source-attribution counted. DATA: fetch both works' Hebrew text from Sefaria; define a pinned external-attribution formula set (tractate names; 'רבי X'/named-sage citation formulae; 'הגאונים'); count per 10,000 words in each; report the formula set used and both densities. PRIMARY R = MT density / Tur density. CLAUSE PRECEDENCE: (1) INCONCLUSIVE_BY_DESIGN if either text is absent or the formula set cannot be applied consistently; (2) SUPPORTED if R < 0.1; (3) REFUTED if R >= 0.1. Compute firewalled from threshold and direction.

Known-priors disclosure — what the registrant already knew

Known priors disclosure: Triage (2026-07-17) graded ADJACENT: Maimonides' programmatic no-citation style in the Mishneh Torah is textbook (Twersky, Introduction to the Code) but the per-10k-word density RATIO vs the Tur is never counted, and MT is not literally zero-citation. Direction certain (MT cites far less); the <0.1 MAGNITUDE is the real test. CALIBRATION NOTE: <0.1 is a very bold pin — recent kills show over-pinned magnitudes; MT may cite less-than-Tur but not <one-tenth.

Method and dataset — how it was measured

Register-before-compute: the registration (conjecture_prediction_batch5_20260718.json, written 2026-07-19T09:14:28Z) pinned threshold (SUPPORTED if R = MT_density/Tur_density < 0.1), direction, self-reference exclusions and void conditions ex ante, but delegated the concrete formula set to the compute. Firewalled compute (claude-sonnet, blind to threshold and direction) fetched all 88 text bodies 2026-07-19T09:25-09:27Z and ran 2026-07-19T09:38:49Z - computed_at postdates registered_at. The compute pinned one formula-family set (37 Bavli tractate names with prefix-stripping; whitelisted 'Rabbi/Rav + name' sage formulae; geonim tokens; MT-only self-pointer exclusion) and reported TWO defensible aggregations an order of magnitude apart: raw tractate-token counting (R=1.0583) and construct-marker-restricted counting (R=0.1199), each with QC samples and per-component breakdowns. Verdict authored by claude-fable-5 against the registered clause precedence.

Dataset: Sefaria machine-readable Hebrew texts: Mishneh Torah reconstructed from the site TOC as 14 Sefer categories / 84 Hilchot leaf indices (769,174 words; no unified 'Mishneh Torah' index exists on Sefaria) vs the Arba'ah Turim as the single 'Tur' index, 4 Turim fetched whole (703,145 words; OC 697 / YD 403 / EH 178 / CM 426 simanim). Both operationalizations of 'explicit external source-attribution formulae per 10k words' are in evidence: RAW (every tractate-name token + named-sage formulae + geonim) and CONSERVATIVE (construct-state ד- tractate citations only + named-sage formulae + geonim).

computed 2026-07-19

Weigh in

No community feedback yet.

New here? Create an account first

Create an account or sign in and your feedback is tied to you — you can track it, get replies, and claim this conjecture so others know you’re working on it. Prefer not to? Just leave your take below as a guest — only the name you type is shown.

Add your take

Posted immediately (spam is removed). Community feedback is never an adjudicated verdict and never changes this conjecture's triage label or status above.