Ars Inquirendi

QUINCUNX: A Discovery Engine for the World's Sparse Domains, & Where They Join

Blueprint and First Registered Test on Historical Pre-Print-Era Evidence

Edoardo Tresoldi, Basilica di Siponto (2016) — the lost early-Christian basilica partly reconstructed in wire mesh above its own ruins. Photo: mauritius images GmbH / Alamy.
Edoardo Tresoldi, Basilica di Siponto (2016) — the lost early-Christian basilica partly reconstructed in wire mesh above its own ruins. Photo: mauritius images GmbH / Alamy.

Abstract

For fifty years, machines have generated conjectures and killed them against evidence one domain at a time: Graffiti in graph theory, BACON and Eureqa in physical law, Robot Scientist in yeast genetics, FunSearch and its successors in problems a program can score. Each required a hand-built hypothesis language and evaluator. QUINCUNX instead attempts a general loop that has not previously been implemented as a whole: Mint conjectures; Select them against evidence; Grade the mechanisms behind the survivors; Promote the strongest into reusable, explicitly scoped inference instruments; and apply those instruments to incomplete records to Infer what the surviving record does not contain.

What makes such an assembly newly possible is the conjunction of large language models, modern computation and extensive machine-accessible data. Language models provide something approaching a general conjecture language: they can generate hypotheses in volume, operationalize prose claims against documented datasets, write machinery to test them, and move ideas across domains that previously required separate specialist systems. QUINCUNX consequently moves the principal bottleneck from conjecture generation towards severe selection. A possible further product follows from that generality: at sufficient scale, the engine may reveal where apparently separate domains of knowledge actually join, as mechanisms or inference instruments generated in one body of evidence prove testable in another.

Its intended scope is broad, but particularly important are fields in which decisive experimental or computational oracles are unavailable. These are domains of found data: evidence accumulated for other purposes, often unrepeatable, incomplete, survivorship-shaped, unevenly measured and mutually dependent. Historical evidence provides an extreme case. QUINCUNX’s first registered masked-ground-truth test therefore created an artificial historical lacuna in the Epigraphic Database Heidelberg (EDH), a completed scholarly corpus of roughly 82,000 Roman inscriptions. Thirty-three of its sixty-six provincial files were left unfetched while the engine froze twenty-five conjectures and sixteen quantitative predictions derived from the half available to it; the concealed files were then fetched, reconciled against the source and mechanically scored.

The first registered run demonstrated inference into a deliberately concealed part of a known historical dataset: information extracted from the visible part had measurable predictive purchase on pre-specified properties of the hidden part. Several estimates landed close to their targets, including an overall non-Latin share predicted at 2.42% and observed at 2.42%. More importantly, a subsequently registered comparison with structure-blind baselines showed substantial positive predictive skill for the two linguistic-geography differentiations, while the corresponding marble differentiations performed worse than the naive baseline. The same procedure therefore rewarded structural inference where it travelled and penalized it where it did not. Of twenty-five registered conjectures, nineteen survived their first tests, five were killed and one became undecidable.

The run also demonstrated why inference over found data requires more than successful prediction. Defects were encountered in statistical construction and data instruments; these were preserved in the record, diagnosed, repaired under registered rules where the correction was mechanically determined, and separated from subsequent questions of independent certification. More generally, dataset integrity and construction-path independence must be demonstrated rather than assumed; model memory is a possible contamination route; model-coded variables require provenance and independent validation; mechanisms must survive adversarial rivals rather than merely accompany correlations with plausible stories; and promotion must account for statistical power, multiplicity, calibration and the effective number of independent evidential paths. Data cleaning, reconciliation and instrument repair are therefore part of the engine itself rather than preliminary housekeeping.

Above all, inferred quantities must never harden into evidence for later rounds. QUINCUNX therefore builds breadth, not height: its inferences may direct the search for new evidence, but they may not themselves become evidence. The first programme shows that the procedure can freeze bets before outcomes are available, preserve holdouts, infer properties of concealed data, distinguish useful structural differentiation from harmful differentiation, kill attractive conjectures, expose defects in its own machinery, enforce pre-outcome Grade and park gates, and withhold mechanism credit while rival explanations remain alive.

No inference instrument has yet been promoted, and no genuinely unknown historical quantity has yet been inferred. That is the present boundary of the claim. Nor should the eventual power of such an engine be judged from a single batch: QUINCUNX is a population process whose real performance, if the architecture succeeds, will emerge only through repeated selection across large numbers of conjectures and heterogeneous, genuinely independent bodies of evidence.

The machine is not finished. The first result is that the loop can now be run.

Keywords: automated discovery, inference, conjecture generation, falsification, large language models, found data, preregistration, masked ground truth, calibration, survivorship, epigraphy, Roman inscriptions, manuscript studies, digital humanities, the pre-print world

DOI: paper (this version) 10.5281/zenodo.21908761 · all versions 10.5281/zenodo.21878807 · data & audit record 10.5281/zenodo.21878845.

Part I — Blueprint

1. Conjecture, selection and the missing loop

The idea that knowledge can advance through the generation of variation followed by severe selection is older than machine learning. Donald Campbell described creative thought in terms of "blind variation and selective retention" 1; Karl Popper cast the growth of knowledge in an explicitly evolutionary form, through conjecture and refutation 2. Computer scientists subsequently built working pieces of that process. AM generated mathematical concepts 3; Graffiti produced graph-theoretic conjectures 4; BACON searched for regularities in physical data 5; Robot Scientist generated biological hypotheses and tested them experimentally 6; Eureqa searched large spaces of symbolic relationships 7; more recently, FunSearch and related systems have coupled generative models to machine-checkable evaluators and searched spaces too large for hand construction 8 9.

These systems establish an important precedent: conjecture generation can be mechanized. They also expose a historical limitation. Each operates within a domain whose hypothesis language, evaluator or experimental oracle has been built for that domain. A graph conjecture can be tested by graph machinery, an equation by numerical fit, a program by running it, a biological hypothesis by an assay. The machine knows what counts as a candidate and what counts as success because those definitions have been engineered in advance.

QUINCUNX asks whether that architecture can become substantially more general. Its proposed loop is:

Mint → Select → Grade → Promote → Infer

The first two stages resemble earlier conjecture engines: generate candidate claims and subject them to tests. The remaining three become necessary when the engine is moved into domains where evidence is partial, dependent and incapable of providing an immediate oracle. A survivor must therefore face a second question: did the relation survive because the proposed mechanism has genuine support, or because some other process created the same pattern? Hence Grade. Even a graded survivor is not automatically knowledge portable beyond its test environment; it requires a specified scope, calibration, provenance and surviving falsifiers before it can become a reusable inference instrument. Hence Promote. Only then may such an instrument be applied to an incomplete record in order to estimate something that the surviving record itself does not contain. Hence Infer.

The distinction is fundamental. QUINCUNX is not intended as a system that asks a language model what probably happened and appends a confidence score to the answer. It attempts instead to manufacture instruments whose right to make such estimates has itself survived selection.

2. Why QUINCUNX?

A quincunx is the five-point arrangement familiar from the face of a die: four points surrounding one at the centre. The name first suggested itself through the Roman quincunx, a bronze denomination whose value and five-pellet mark gave the word its ancient form 10 — an appropriately minted object for an engine whose first stage is Mint. The same fivefold figure appears in the Saturn V engine arrangement, with four steerable outer F-1 engines surrounding a central fixed engine 11; Francis Galton used the word for the disposition of pins in his probability apparatus, where innumerable individual deflections resolve into aggregate order 12.

Thomas Browne made the figure itself a method. In The Garden of Cyrus (1658), published with Urne-Buriall, his meditation on what survives the dead 13, Browne pursued the quincunx through botany, art, nature and antiquity. He found the pattern everywhere. The name therefore carries a warning as well as an emblem. Language models are extraordinarily good at finding patterns everywhere they look; QUINCUNX exists to make finding one expensive. Its five stages are not five opportunities for interpretation. Four exist largely to discipline what the first produces, with the central problem always being selection: what survives, why it survived, and how much permission survival earns.

3. The five stages

3.1 Mint

A language model generates large batches of novel and falsifiable conjectures about a body of evidence. Mint is intentionally promiscuous. The generator is not required to be reliably right; indeed, once generation becomes cheap, there is little reason to optimize prematurely for a high survival rate. What matters is producing enough varied, testable claims for severe selection to operate.

A minted conjecture must nevertheless be more than an interesting sentence. Before testing, the claim must acquire an operational form: its population must be stated, its variables defined, the evidence capable of deciding it identified, and the observations that would kill it written down. The important sequence is therefore:

claim → operationalization → kill-condition → registration → evidence → verdict

rather than:

claim → evidence → convenient operationalization → explanation.

The purpose of Mint is not to manufacture conclusions. It is to manufacture things capable of dying.

In practice, Mint is not selection-free, and its internal curation must therefore be disclosed. The size of the batch, the constraints under which it was produced and the fate of candidate ideas considered before the batch was fixed all matter to the denominator against which later survival is read. A system that generated a hundred possibilities, silently discarded seventy-five and registered the remaining twenty-five would be performing an unreported selection step before Select had formally begun. QUINCUNX therefore records not only the conjectures that emerge from Mint but, where practicable, the candidate angles explored and abandoned while the batch is being formed. Creativity is encouraged at this stage, but hidden curation is not.

This places an unusual premium on the generative capacity of the minting model. Select, Grade and the later stages are designed to be severe; Mint has the opposite task of creating enough varied and potentially valuable material for that severity to act upon. For the first EDH batch we therefore used Fable 5, the frontier model selected by the project for its generative and reasoning capacity, deliberately favouring generative range over economy. The underlying hypothesis is that model creativity matters: a more capable minter should be better able not merely to notice obvious associations in a dataset, but to propose surprising variables, mechanisms, comparisons and transfers that a narrower search would never expose to selection. We have not yet calibrated this property, and QUINCUNX does not treat a model’s apparent creativity as evidence of truth. It treats creativity as a source of variation whose value can only be established downstream by what survives.

Because that creative search is domain-general, Mint can also propose relations not merely within datasets but across bodies of knowledge; the larger implications of that capacity are considered in §4.

3.2 Select

Select executes the registered tests against evidence. Wherever possible the deciding evidence is held out until the conjecture, its operationalization and its kill-condition are frozen. Batch size is recorded because one survivor among ten attempted conjectures is not evidentially equivalent to one survivor among ten thousand.

A survival does not prove a conjecture true; it means that a specified attempt to kill it failed. A death, correspondingly, is not discarded. It enters the graveyard with the evidence that killed it, the scope over which the refutation holds, the test that delivered the verdict and the population of sibling conjectures among which it was selected. The graveyard is therefore one of QUINCUNX's intended outputs. This becomes particularly important once Mint begins operating at machine scale, because a system capable of producing thousands or millions of conjectures while recording only the survivors would simply automate publication bias. 14 QUINCUNX publishes the dead beside the living.

3.3 Grade

Survival is not mechanism. A relationship may hold in one body of evidence and be useful for interpolation within it while possessing no warrant whatever for extrapolation elsewhere 15. Language models intensify this problem because they can generate compelling causal narratives around almost any correlation. Plausibility is therefore awarded no mechanism credit.

At Grade, model families disjoint from the original minter generate alternative mechanisms capable of producing the same surviving pattern. Those rivals remain live until discriminating evidence removes them. Where possible, the proposed mechanism must also make surplus predictions: consequences not already used to select the original relationship. A survivor may consequently retain its descriptive result while being refused permission to travel.

The distinction can be stated simply:

Select licenses survival. Grade may license mechanism. Promote may license inference.

They are not synonyms.

3.4 Promote

A sufficiently tested survivor may become a reusable inference instrument. Promotion is deliberately difficult. An instrument must carry more than a finding: it needs a tested scope, an explicit estimand, a mechanism grade, statistical error controls, a provenance record, a construction-path audit, a selection history and still-armed falsifiers.

Its scope should ultimately be machine-readable. If an instrument has been validated only for a particular population, period, data-generating process or range of missingness, a request outside that scope should be refused rather than answered with an attractive caveat beneath the number. No QUINCUNX conjecture has yet been promoted.

3.5 Infer

Only a promoted instrument may be applied to an incomplete record in order to estimate something absent from that record. The resulting quantity remains explicitly inferred and retains the identity and version of the instrument that generated it. An estimate does not become an observation because it is useful, nor does it become evidence because it has been repeated. It may never be fed back into the evidence layer merely because the system produced it.

This gives QUINCUNX its governing architectural rule:

The engine builds breadth, not height.

4. Why large language models change the problem

The underlying epistemology is not new. What is new is the cost structure. Earlier conjecture engines required domain-specific hypothesis spaces and evaluators. Language models alter that constraint in at least four ways.

First, they provide something approaching a general conjecture language. A claim about manuscripts, inscriptions, tumour registries, coin hoards, ecological surveys, administrative records, supply chains or astronomical catalogues can be expressed through the same natural-language interface. This does not make the underlying sciences interchangeable, but it does mean that the machinery required to formulate candidate claims no longer has to be engineered from scratch for every field.

Second, language models make mass generation possible. Conjectures cease to be scarce, shifting the bottleneck towards selection. A filter cannot discover a hypothesis it was never offered; if generation is cheap, it can be broad, speculative and combinatorial, provided that survival is correspondingly difficult. The eventual behaviour of QUINCUNX should therefore be visible not only in individual results but statistically across populations of conjectures: which classes die quickly, which repeatedly survive, which mechanisms travel, which instruments calibrate, and which domains provide too weak a selection environment to support inference at all.

Third, language models can operationalize prose. Given a conjecture and documented data, they can identify potentially relevant fields, propose proxies, formulate queries, write code and convert a qualitative proposition into a measurable one. This is powerful, but it is also one of the engine's largest failure surfaces. A prose claim often admits several defensible operationalizations 16; if the operationalizer sees the outcome before choosing among them, flexibility becomes fitting. The admissible specification set must therefore be constrained, frozen where possible and exposed rather than silently collapsed into whichever implementation happens to deliver a result.

Fourth, language models can create new measurements from previously unstructured evidence. Catalogue descriptions, manuscript transcriptions, inscriptional apparatus and historical prose can be turned into machine-readable variables at scales formerly impractical. But a model-coded classification is not equivalent to a shelfmark, numerical measurement or explicit date supplied by a source: it is an inference about a source and therefore requires provenance and independent validation.

Mint consequently operates at more than one scale. At its simplest it proposes falsifiable relations within a body of evidence: that one population should differ from another, that a variable should predict an outcome, or that a mechanism should leave a specified trace. But a domain-general conjecture language also permits a second kind of search. The engine can propose that a structure, variable, estimator or mechanism developed in one field may apply in another: a selection function from astronomy to a historical catalogue, an unseen-population estimator from ecology to manuscript survival, or a network process from one class of records to another. The creative act here is not merely generating another hypothesis inside an established field; it is proposing that two bodies of knowledge presently treated as separate may in fact share a testable structure.

Such conjectures are therefore candidate relations between knowledge domains. Disciplinary boundaries need not be supplied to the engine as boundaries of nature. Mint proposes the crossing; Select asks whether the proposed relation survives evidence; Grade asks whether the apparent analogy reflects a mechanism capable of travelling rather than a superficial resemblance. Failed transfers matter too, because they locate boundaries at which apparently comparable domains cease to behave alike.

Surviving cross-domain conjectures could do more than map relationships between bodies of knowledge presently treated as separate. Suppose a conjecture combines a structure developed in physics with a mechanism or variable developed in biology, where the two had not previously been related, and the resulting claim survives severe testing against independent real-world evidence. What has then been established is not merely an interdisciplinary analogy. The result is evidence that structures first recognized in different disciplines may correspond to a real relation in the world itself.

At sufficient scale, QUINCUNX could therefore begin to expose latent structure across the existing organization of knowledge: not simply where disciplines overlap, but where phenomena conventionally assigned to different domains participate in common regularities. Failed transfers would be equally informative, marking places where an apparent analogy does not survive contact with reality. In that sense QUINCUNX may eventually become not merely an engine for inference within domains, nor merely a way of reconnecting disciplines, but a means of discovering and testing relations within the structure of the world that our existing division of knowledge has obscured. Demonstrating such relations is a later goal; if successful, it may prove as important as any individual historical inference the engine produces.

5. Found data

The domains in which QUINCUNX may be most useful are precisely those in which evaluation is hardest. These are the world’s sparse domains: fields whose surviving or measurable record is thin relative to the world it attests, and whose evidence cannot be regenerated on demand. A mathematician can ask a proof checker; a program can be rerun; a laboratory scientist can often obtain another measurement or perform another assay. Historical, economic, ecological, epidemiological and many astronomical records belong to a different economy. The evidence already exists, was accumulated for purposes other than the present experiment, may be expensive or impossible to reproduce, and may contain omissions that are systematic rather than random.

The historical record is an extreme case because the surviving data are themselves products of the process we wish to understand. Books survive through copying, collecting, fire, damp, war, rebinding, disposal, institutional preference and chance. Inscriptions survive destruction, reuse, excavation, publication and modern cataloguing. Archives inherit the priorities and accidents of the institutions that created and later preserved them. What remains is not a random sample of what once existed; survival is part of the phenomenon.

Observational astronomy provides a mature analogy. Astronomers do not normally treat a survey catalogue as though the objects entering it were selected at random. They model the survey's selection function 17 18 and measure completeness through procedures such as injection–recovery, in which synthetic objects of known properties are placed into real observations and the pipeline is asked to recover them. History cannot perform the equivalent experiment on the past itself: nobody can rebuild the thirteenth century, destroy its archives again and see what survives on the second run. That forces additional disciplines into the engine.

5.1 Count independent evidence, not datasets

"Supported by three databases" does not mean "supported by three independent lines of evidence." Catalogues copy catalogues, aggregators ingest shared source material, different scholarly databases may inherit classifications from the same handlist, and apparently separate resources may employ overlapping personnel or common technical pipelines. Agreement can therefore be generated by common construction.

The dependence is concrete in both scholarship and code. The count of surviving Piers Plowman manuscripts moved from fifty-eight in Hanna’s 1993 handlist 19 to fifty-nine in Bowers’s 2007 count, as reported by Warner, 20 not because fourteen years had produced two independent enumerations, but because the inherited list had been read at a different boundary, admitting one additional four-line extract; every tally examined for this project ultimately descended from the same handlist. QUINCUNX then supplied its own machine-age version of the same mistake. A later pinned census measurement reproduced an earlier figure of 42,526 manuscript records exactly — apparently reassuring agreement — until audit showed that both measurements shared the same defective regular expression, which silently excluded every document describing exactly one manuscript. Correcting the common error moved the counting floor to 43,058. In one case a handlist, in the other a regex: agreement generated by common construction was not corroboration.

The relevant quantity for QUINCUNX is not simply the number of datasets but the effective number of evidential paths that could have failed separately. Construction-path independence must be demonstrated rather than assumed. This is not a philosophical nicety: selection can become actively worse when dependent datasets are treated as independent. A conjecture detecting a shared catalogue artefact may repeatedly "replicate" across descendants of the same construction process and become enriched among the survivors. Apparent corroboration can therefore be anti-evidence about the quality of the selection environment.

5.2 The generator has memory

Evolutionary variation is blind; an LLM's variation is not. A model may have encountered the dataset or scholarly claim being tested during training. A conjecture that looks novel to the project may be rediscovery, and a public corpus withheld from a particular run may nevertheless have influenced the model's prior expectations.

No prompt instruction can erase model weights. Registration solves a narrower but important problem: it prevents the project from adjusting its claims after the deciding evidence has been inspected. It does not prove that the model had never encountered that evidence before. QUINCUNX therefore distinguishes registration-blind replication from stronger forms of discovery. Public corpora can provide valid holdouts against post-registration fitting while remaining contaminated, in principle, by pretraining exposure.

The strongest future evidence will come from material unavailable to relevant models at source level: newly catalogued collections, newly imaged manuscripts, newly excavated objects, post-training measurements and records not yet digitized when a conjecture was frozen. The pre-print world is particularly promising in this respect because its digital corpus is still being built.

Model memory creates a second problem beyond leakage. The generator’s own priors reflect the digitization state of knowledge from which its training corpus was built. For the pre-print world, primary material remains radically thinner in machine-readable form than later printed and born-digital evidence, so much of what the model knows arrives mediated through editions, translations, catalogues and modern scholarship. The minter can therefore inherit the same survivorship and digitization distortions that QUINCUNX is intended to reason through. This creates a useful three-way tension at the sparse edge: under-digitized domains may offer the largest inferential prize, because so much remains absent; the weakest generator, because comparatively little primary evidence has entered its substrate; and the cleanest future holdouts, because genuinely undigitized material cannot already have entered model weights at source level. That condition is not fixed. As primary corpora enter the machine-readable world, later models should acquire a less mediated representation of those domains. Whether their conjecture and inference performance then improves measurably is itself a falsifiable future prediction of the programme.

5.3 Clean data is not a starting condition

Found data do not arrive experimentally pristine. They arrive with duplicated records, conflicting units, undocumented inheritance, stale identifiers, encoding differences, partial fields, serialization choices, local scholarly conventions, inaccessible licence terms and proxies whose meaning changes from one corpus to another.

For QUINCUNX, data cleaning is not preparation for the engine's work; it is part of the engine's work.

A corpus must therefore be treated as an instrument with properties that can themselves fail. Its unit of count, scope, provenance, serialization conventions and classification systems need to be inspected and recorded. Crude measurements carried across corpora require structural diagnostics capable of revealing that they are measuring file formats, cataloguing practice or software behaviour rather than the historical phenomenon they purport to measure.

Beneath these checks lies the requirement of source fidelity: the evidential chain from source to digital record must preserve the properties on which the claim depends. A transcription, catalogue field, photograph or other surrogate may be adequate for one question and incapable of deciding another; where a claim turns on a physical feature of an artefact, validation must therefore reach back to the artefact itself or to an image capable of showing that feature. The first run enacted this principle when a proxy intended to distinguish physical loss from editorial restoration was validated against photographs of the inscribed objects and failed its gate, parking the analysis that depended on it.

Nor can this audit occur only once, before testing begins. Contact with new evidence will expose defects that were invisible earlier. A parser may strip meaningful markup; a lookup table may encode an external vocabulary incorrectly; a threshold rule in prose may differ from its executable implementation; a model-coded semantic proxy may turn out not to measure the construct assigned to it. The scientific requirement is therefore not that a corpus remain forever untouched after first contact, but that the sequence remain explicit:

registered instrument → contact with evidence → diagnosed defect → registered repair → controlled re-execution → fresh validation.

A repair fully determined by an external specification or by semantics already registered is not equivalent to tuning a procedure until it produces a favourable answer. The distinction must be recorded. Blindness is therefore local and spendable, rather than a mystical state of permanent purity. The first contact can expose a defect; a mechanically determined repair can then be applied and its effect measured on the same material, provided that the corrected result is labelled as such and never confused with a second independent blind validation. General certification still requires fresh evidence.

Data cleaning, in this sense, is capable of producing findings of its own. Catalogue hygiene is an engine product.

5.4 The firewall is absolute

A found-data inference engine can corrupt itself. Suppose an instrument estimates a missing historical quantity. The estimate becomes useful, then familiar, then appears in a table; a later conjecture is tested against that table as though the inferred value were observed. The engine has begun confirming its own previous outputs. 21

Without that firewall the engine becomes a Tower of Babel: locally plausible claim built upon locally plausible claim until the top is detached from observation.

A tower of Babel built of luminous wire mesh, rising from solid stone ruins
A tower of Babel built of luminous wire mesh, rising out of its own stone ruins. Created with Google Gemini, after Pieter Bruegel the Elder’s The Little Tower of Babel (c. 1568, Museum Boijmans Van Beuningen, Rotterdam) and the wire-mesh sculptures of Edoardo Tresoldi, photographed by Roberto Conte.

QUINCUNX therefore distinguishes three broad provenance classes. Found values are directly present in a source. Model-coded values are produced by a model reading a source; these require independent validation before carrying evidential weight and retain their provenance thereafter. Inferred values are produced by an inference instrument and never enter the evidence layer.

The rule still permits inference to direct measurement. An instrument may suggest that a particular class of source should contain a signal; researchers may then scan, catalogue, excavate or measure that source. The result can become new evidence because the world, rather than the inference, supplied it. An inference may point to scaffolding; it may never be scaffolding.

The firewall must ultimately operate across model generations as well as within a single QUINCUNX run. Expanding the digitized primary corpus is desirable: grounded records entering future training substrates should give successor models a richer and less mediated representation of sparse domains. But inferred quantities must not enter those substrates disguised as observations. Otherwise one generation’s conjectures become the next generation’s priors, creating a slow cross-generational form of the same self-confirmation the internal firewall is designed to prevent. The engine may improve the corpus, and the corpus may improve the next engine; inference must remain visibly inference throughout that ratchet.

An older cartographic metaphor, developed with Anthony John Lappin, captures what this permits. 22 QUINCUNX should grow like a map rather than a tower. Found evidence forms the surveyed ground; between and beyond its fragments lies terra incognita. But unknown territory need not remain featureless until it can be directly observed. Independently tested inference instruments may approach the same absence from different directions, constraining its extent, contents and relations until an inferred outline begins to emerge through their agreement. One instrument may suggest a boundary, another a likely magnitude, another exclude a mechanism, another imply that two known regions must connect through what is missing. If those lines are genuinely independent, their convergence may progressively narrow the space of possible worlds even where the missing evidence can never be recovered. A later QUINCUNX stage may formalize this as inferential triangulation: the combination of independently validated inference instruments, grounded in distinct evidential paths, whose convergence constrains the same unknown quantity or structure. The distinction is deliberate: inferential triangulation is not multiple models agreeing; it is multiple independently warranted instruments converging on the same terra incognita. The term is prospective here; no such combined instrument has yet been calibrated or promoted.

The distinction from a tower remains absolute. An inferred contour does not become surveyed ground merely because several instruments agree, and one inferred result is not promoted into evidence for the next. Its inferential status, provenance and degree of warrant remain visible. Some areas of the map may later be directly charted as new records are found, digitized or measured; others may remain permanently inaccessible yet increasingly constrained from their edges. Breadth, not height therefore allows the map to become richly articulated before all of its territory is — or ever can be — directly known.

Hence the rule again: breadth, not height.

6. Selection must be severe

Once conjectures can be generated cheaply, weak selection becomes dangerous. Suppose an engine mints 1,000 conjectures, only fifty of which are true. Imagine that a false conjecture has a 25% chance of surviving one weak test while a true conjecture has a 90% chance. Under three genuinely independent tests, a false conjecture survives all three with probability 0.25³ ≈ 1.6%, leaving roughly fifteen false claims beside about thirty-six true ones.

Now suppose the apparent tests are dependent. Passing the first makes passing each inherited catalogue 60% likely. The false-survival probability becomes 0.25 × 0.60 × 0.60 = 0.09, producing roughly eighty-five false survivors beside the same thirty-six true ones. The survivor pool becomes mostly false. More data made the selection worse.

Four disciplines follow. Severity and power matter: a test that was unlikely to detect a false claim supplies little evidence when the claim survives. 23 Multiplicity matters: hundreds or thousands of minted conjectures cannot be interpreted as though each appeared alone. 24 Effective independence matters: five datasets inherited from two construction paths should not count as five corroborating witnesses. Finally, mechanism must compete: a plausible story surrounding a correlation is not a licence to extrapolate it.

The difficulty of QUINCUNX is therefore not primarily Mint. It is building a selection environment capable of making survival informative.

7. Grade: making coherence expensive

The first registered Heidelberg run supplies a concrete example of why Grade is separate from Select. A conjecture predicted a division between Latin-West and Greek-East provinces in their shares of non-Latin inscriptions, organized around the traditional Jireček-line distinction. The registered directional pattern survived in the concealed half, but that did not promote the conjecture.

A rival mechanism proposed that the apparent province-level linguistic boundary was created by unequal mixtures of sites and inscriptional genres: a few heavily excavated sites might dominate provincial totals and manufacture the contrast. That rival could be tested. The evidence was reweighted so that sites rather than inscriptions carried equal weight, and the registered divide did not dissolve. One rival was therefore discharged.

Another was not. Perhaps the source corpora themselves had been assembled regionally in ways correlated with linguistic classification. The available evidence was insufficient to exclude that construction-path explanation. The pattern therefore survived, while the mechanism remained underdetermined; the conjecture stayed interpolation-only. It may describe, but it may not yet travel.

This is the purpose of Grade. Language models make coherent explanation extraordinarily cheap. Coherence is cheap; the engine's job is to make coherence expensive.

8. What is new — and what is not

Every major component of QUINCUNX has precedent. 25 Program-search systems generate and select at scale 26 27; automated laboratories close experimental loops 28 29 30; hypothesis-validation agents test supplied claims 31; model-assisted systems extract variables from textual corpora 32 33; historical retrodiction experiments test predictions about the past 34; statistical science supplies tools for multiplicity, calibration, equivalence testing, robustness and missing-data inference.

The novelty claim should therefore be narrow. QUINCUNX does not claim to have invented machine conjecture generation, automated falsification, model-assisted coding, preregistration or prediction. Its claim concerns their assembly for found data: mass-mint falsifiable conjectures; freeze their operationalizations and kill-conditions; select them against evidence whose construction paths are audited; generate adversarial rival mechanisms through disjoint model families; refuse extrapolation while rivals remain alive; promote only qualified survivors into explicitly scoped inference instruments; apply those instruments to incomplete data; maintain an absolute evidence–inference firewall; and preserve the dead beside the living.

Earlier engines were built predominantly where evaluation was cheap and decisive. That is precisely where several pieces of this architecture matter least. A proof checker does not need an independence audit of three manuscript catalogues; a program-search system does not need to worry that its inferred output will later be cited by another database as primary evidence about the ancient world; an assay offers causal leverage that a historical corpus cannot.

QUINCUNX attempts the full machinery on the other side of that divide: where claims remain falsifiable, but evidence is partial, found and expensive to replace.

9. Calibration before darkness

A beautifully disciplined engine can still be wrong. Inference into genuinely lost history creates a fundamental validation problem: if the answer no longer exists, the apparent success of a reconstruction cannot calibrate the method that produced it. Before QUINCUNX is allowed to infer into a genuinely unknown region, it must therefore work where truth is known.

Three forms of test are particularly important. Masked-ground-truth reconstruction removes part of a known dataset from the active run, freezes predictions from what remains, then reveals the concealed part. Synthetic written worlds generate known production, copying, loss and cataloguing processes, allowing the engine to infer quantities whose true values are retained by the simulator. Injection–recovery places synthetic records of known properties into real, messy corpora and measures what the complete data and inference pipeline actually recovers.

These tests answer different questions. Masked ground truth asks whether information extracted from one portion of a real corpus predicts a second portion. Synthetic worlds ask whether the full inferential architecture can recover known truth under controlled production and loss. Injection–recovery asks whether the machinery survives contact with real data plumbing.

Calibration cannot prove that an instrument is correct about the historical past. It establishes the minimum condition for taking such an instrument seriously: that it can recover truth when the answer exists. Without that, inference into darkness is assertion in a lab coat.

Part II — First Test

10. Creating an artificial lacuna

The first registered demonstration used the Epigraphic Database Heidelberg, a completed scholarly corpus of roughly 82,000 Roman inscriptions. The experimental idea was deliberately simple: if QUINCUNX is eventually intended to say something disciplined about missing historical information, it should first be able to look at an incomplete version of a world whose complete state remains available to the experimenter and make risky predictions about what has been concealed.

The sixty-six EDH province files were divided into two sets. Thirty-three formed the visible half; the other thirty-three were declared as the holdout and left unfetched by the project.

The partition itself was deterministic rather than randomized. QUINCUNX took the first half of EDH’s own published province list as the visible set and the second half as the holdout, with the rule fixed before inspection of the concealed records. This removed project discretion over which provinces would be hidden, but deliberately accepted a harder risk: historical structure might itself track the alphabetical partition. The registration therefore described the non-random split as a “named wager” and identified cross-half exchangeability as something the experiment was staking rather than assuming.

Against the visible half, the engine froze twenty-five conjectures with their registered tests and kill-conditions, together with sixteen quantitative predictions about the aggregate shape of the concealed half. The targets included provincial concentration, numbers of large and small provinces, chronological composition, non-Latin inscription shares, funerary and votive proportions, dating characteristics, textual-damage markup, photographic coverage and material composition.

The provenance of those twenty-five conjectures was itself retained as part of the experiment. On 5 August 2026 the minting model worked read-only against aggregate statistics from the thirty-three visible EDH provinces; it did not receive record dumps, and the held-out province files were not present on the project machine. Nine substantive rounds of aggregate interrogation supplied the descriptive environment from which it minted a single batch of twenty-five conjectures. All twenty-five were disclosed and all twenty-five proceeded to registration: there was no subsequent triage selection of a preferred subset.

The denominator is nevertheless deeper than twenty-five. The minting record preserves eight candidate angles considered during formation of the batch. Two were reworked, on the evidence available in the visible half, into conjectures that entered the registered batch; six were not minted. Among the latter were proposals rejected as an apparent catalogue artefact, as near-redundant with stronger conjectures, or because the available variable would have required an unvalidated measurement procedure. Selection during creative exploration was therefore not treated as though it had never happened.

A triage pass on 6 August passed all twenty-five minted conjectures with zero kills but imposed five hardening directives before registration: harmonizing denominators, replacing family-level eligibility descriptions with explicit per-conjecture province lists, reconciling two overlapping conjectures with one enumerated carve-out, making two kill-conditions robust to attrition, and fixing one aggregation rule. The purpose of triage was thus not to choose the most attractive conjectures after Mint, but to make the disclosed population sufficiently explicit to be allowed to fail. The resulting registered objects carried enumerated scopes, frozen queries, kill-conditions, minimum-support rules and verification notes. The mint files, considered-and-dropped ledger, triage record and final registrations are preserved in the accompanying data deposit.

The sixteen quantitative recovery targets were produced differently and should not be conflated with this model-minted population. They were derived mechanically by a single committed script from the visible half, using the registered procedure and fixed seed, and frozen in the same registration commit as Block A. In total the registration therefore disclosed forty-one objects: twenty-five model-minted conjectures and sixteen script-derived recovery predictions.

The holdout was then downloaded, and every concealed province reconciled exactly against the database's own declared record total before scoring. That reconciliation establishes integrity relative to EDH's published catalogue counts; it does not turn EDH into a probability sample of the Roman world. The experiment predicts properties of the catalogue, not direct proportions of ancient reality.

Because EDH is public, prior model-training exposure cannot be excluded. The holdout is therefore blind in an operational and registration sense: the project had not fetched or queried those files before freezing the predictions, and the active run was not supplied with them. It is not a claim that the corpus was absent from all relevant model weights. The experiment is accordingly classified as replication-grade.

11. The recovery result

The twenty-five registered conjectures produced 19 survivors, 5 deaths and 1 undecidable case. The sixteen quantitative predictions provide the more direct masked-ground-truth test. The blind registered score is 9/16 overall; on the fifteen targets mechanically eligible for the correction described in §12, the registered intervals covered 8/15 against 11/15 under the corrected instrument.

Target (holdout half)UnitPredicted point90% interval*Landed†Within?Corrected-instrument 90% interval‡Within, corrected?
Largest province’s share of records%18.20[11.97, 26.91]11.87no7.80–29.61Yes
Top-five provinces’ combined share%64.21[50.64, 78.09]48.77no45.09–84.73Yes
Provinces with ≥1,000 recordscount of 328.9[6, 12]15no3.9–13.9No
Provinces with <300 recordscount of 3216.0[11, 20]12yes10.0–23.0Yes
Century-by-century dated mix (11 bins)TV distance≤ 0.1196 (registered ceiling)0.0703yesn/a — abstained§
Non-Latin share, whole holdout%2.42[1.29, 4.59]2.42yes0.22–5.02Yes
Epitaph share%40.46[27.07, 54.68]47.93yes22.47–59.14Yes
Votive share%23.53[16.84, 30.71]19.80yes14.28–32.67Yes
Dated share%80.90[64.43, 91.22]66.40yes62.59–100.60Yes
Range-dated precision, among dated%95.49[93.37, 96.64]92.92no93.27–97.77No
Bracket (damage-markup) density%65.33[62.41, 68.47]61.54no61.11–69.38Yes
Records without photograph%22.29[12.96, 30.72]14.95yes10.85–34.51Yes
Non-Latin share, Latin-West stratum%0.63[0.39, 0.90]0.84yes0.27–0.98Yes
Non-Latin share, Greek-East stratum%24.63[17.71, 38.60]19.03yes9.41–39.96Yes
Marble share, Mediterranean-coast stratum%37.61[31.57, 40.89]22.55no31.11–43.69No
Marble share, continental stratum%0.95[0.65, 1.32]12.85no0.44–1.45No

* “Blind” with two stated limits: the predictions were derived by a committed script and sealed before any holdout record was fetched, but the corpus is public, so prior model-training exposure cannot be excluded (the replication-grade caveat); and the intervals are printed exactly as registered, from the construction §1 describes as later found too narrow — the defect is part of the record, not repaired away. Values rounded to two decimals; the frozen four-decimal registrations, with each target’s exact definition and evaluation set, are in the deposited artifacts.

† “Landed” means the holdout half as fetched and reconciled against the database’s own declared per-province counts — figures about the catalogue, not about the ancient world.

‡ Corrected-instrument diagnostic per the registered protocol (docs/DESIGN_QUINCUNX_EDH_CORRECTED_INTERVAL_DIAGNOSTIC_20260811.md): the registered construction plus the mechanically entailed second variance component — the future half’s own sampling error, whose omission the audit identified — recomputed once from the same frozen visible-half inputs under a public seed, with the construction code barred from reading the landed values (intervals written to disk before scoring joined them). On the fifteen scorable targets the registered intervals covered 8 and the corrected instrument covered 11: three intervals flipped from missed to covered (largest-province share, top-five share, bracket density), none flipped the other way. The four remaining misses — the province count, range-dating precision, and both marble strata — are not explained by the identified omitted-variance defect; they remain attributable to prediction error, structural mismatch, residual interval inadequacy, or some combination. The dated-share corrected upper bound exceeds 100% and is reported unclipped (an artifact of the additive construction near a boundary). § The century-mix row abstains: a total-variation ceiling against a fixed anchor has no mechanically entailed two-sample analogue, and the registration forbids invented constructions. The blind score of record remains 9/16. This column measures the corrected instrument on this holdout; it does not certify it — the holdout selected which defect was repaired, so certification remains the registered fresh-population run.

The point estimates show genuine predictive purchase on the concealed half. The overall non-Latin share landed exactly on the frozen point estimate: 2.42% predicted, 2.42% observed. Several other targets landed comfortably within their registered ranges; others failed badly. Those facts belong together. QUINCUNX did not reconstruct the hidden half wholesale, but neither did it merely produce generic plausibilities. It made pre-specified numerical claims about data concealed from the active run, and some of those claims tracked the concealed values closely.

The strongest statement supported by this first experiment is therefore not that QUINCUNX "recovered the missing half". It is that QUINCUNX demonstrated inference into a deliberately concealed part of a known historical dataset: information extracted from the visible part had measurable predictive purchase on pre-specified properties of the hidden part. The registered baseline comparison makes that claim stronger.

12. The interval failure, diagnosis and corrected-instrument diagnostic

The registered uncertainty intervals failed: only 9 of 16 nominal 90% intervals covered their concealed targets. For a procedure advertised as producing 90% prediction intervals, that is not a successful calibration result. The project's first diagnosis focused on the split itself. The alphabetical partition had created two structurally dissimilar halves of the Roman provincial record. That observation was correct and remains methodologically important, but it was not the principal defect in the interval construction.

Audit found a simpler error. 35 The registered ranges had been generated by repeatedly resampling the visible provinces and measuring how much the visible-side statistic varied. That procedure estimates uncertainty in the observed-side estimate. But the target was not a fixed underlying population parameter; it was the corresponding statistic in a second finite set of provinces, which has its own uncertainty. The registered semantics therefore required two variance components: uncertainty in the visible-half estimate and uncertainty in the future half being predicted. The original construction included only the first.

Under the simplified registered derivation in which the two halves are independent, exchangeable and equally sized, the components have equal variance. The variance of their difference is therefore doubled, making the predictive standard error larger by a factor of √2. The importance of this diagnosis is that the repair is mechanically entailed by the quantity the interval was already supposed to represent.

The original score is not rewritten; it remains 9/16. But neither is an identified arithmetic error left unrepaired merely because the held-out data have now been seen. A separately registered corrected-instrument diagnostic applied only the mechanically entailed second variance component. The construction used the same frozen visible-side inputs; the code producing the corrected intervals was barred from reading the landed holdout values; intervals were written before scoring joined them to the concealed outcomes; and the protocol required abstention wherever no non-discretionary analogue of the correction existed.

Fifteen targets were mechanically scorable. On those fifteen, the original intervals covered 8/15; the corrected instrument covered 11/15. Three intervals changed from miss to cover — largest-province share, top-five provincial share and bracket/damage-markup density — and none moved from cover to miss. The century-by-century total-variation target abstained because its registered form was a one-sided ceiling against a fixed anchor rather than a point estimate with a mechanically transformable width. Inventing a new construction after reveal would have violated the purpose of the diagnostic.

The corrected result does not establish that the interval procedure is calibrated. Eleven of fifteen remains below nominal 90% coverage, and the repaired procedure requires testing on fresh populations before it can be certified. What the diagnostic does establish is that the omitted future-set variance component was not merely a theoretical objection: correcting it moved three real verdicts in the predicted direction without rescuing the remaining substantive failures. Four misses remained — the large-province count, range-dating precision and both marble strata — and are therefore not explained by the identified omitted-variance defect. They remain available to diagnose model error, structural mismatch, residual interval inadequacy or some combination of these.

A first registered certification-point run of the repaired interval method has already taken place on synthetic populations. It issued no certificate. Post-result assessment then exposed a defect in the certification rule itself: the registered one-sided criterion tested whether coverage was conservative, not whether it was equivalent to nominal coverage, and therefore could not establish the claim that a calibration certificate was meant to establish. A two-sided coverage-equivalence rule has since been registered for fresh synthetic populations; the earlier results cannot be reused to validate it. 36 QUINCUNX therefore encountered a second-order failure here: after repairing an interval instrument, it also discovered that the instrument designed to certify that repair was inadequate to do so.

This distinction is central to QUINCUNX's treatment of data and instrument cleaning. A frozen result records what happened; a registered repair records what happens under an objectively corrected instrument; fresh evidence determines whether the repair generalizes. All three belong in the scientific record.

13. Did QUINCUNX add anything beyond a naive predictor?

Accurate-looking predictions are not enough. A corpus divided into two large pieces may be sufficiently homogeneous that "the other half will look like this half" already predicts many aggregate quantities well. The relevant question is therefore not merely whether QUINCUNX landed near the concealed value, but whether structural differentiation improved prediction over a simpler rule.

That comparison had not been part of the original registered test. A separate addendum was therefore registered after reveal against four already frozen stratified point predictions; the QUINCUNX values could no longer move. A structure-blind null assigned both members of each pair the visible half's aggregate rate, and descriptive skill was measured as:

Skill = 1 − |QUINCUNX error| / |null error|

TargetQUINCUNX predictionStructure-blind nullLanded valueSkill
Non-Latin, Latin-West0.63%2.42%0.84%+0.87
Non-Latin, Greek-East24.63%2.42%19.03%+0.66
Marble, Mediterranean37.61%10.63%22.55%−0.26
Marble, continental0.95%10.63%12.85%−4.35

The contrast is more informative than an average score. For linguistic geography, structural differentiation was highly useful: distinguishing Latin-West from Greek-East produced substantially smaller prediction errors than assigning both groups the aggregate visible-side rate. For marble, the structural model made prediction worse. That is not an embarrassment to be averaged away; it is one of the most important results of the experiment.

QUINCUNX added useful structure where the distinction travelled and subtracted information where the chosen structure — Mediterranean coastline as a proxy for marble access — was wrong. The test did not merely reward sophistication; it rewarded the right sophistication and penalized the wrong one.

The baseline comparison therefore permits a stronger claim than "some predictions contained signal". For at least some registered structural differentiations, QUINCUNX extracted predictive information about concealed data beyond that supplied by a fixed naive baseline. That is an inferential proof of concept. It is not yet reliable reconstruction.

14. What died

The graveyard makes the character of the result clearer than the survival count alone.

14.1 A trade-geography mechanism killed by a quarry

The marble conjecture predicted higher marble use in Mediterranean-coast provinces than in continental provinces lacking direct access to Mediterranean marble trade. The claim was plausible and badly wrong. Noricum arrived from the holdout at roughly 36% marble despite being landlocked and despite visible-side continental analogues below 3%. The reason is obvious after the result is known: Noricum had access to major Alpine marble quarries.

But the registered observable had been coastline. What the historical claim was really about was access to marble; the dataset supplied coastline as the operational proxy. The holdout produced a province in which proxy and construct separated, and the conjecture died. The baseline test later showed the cost numerically: the structural marble differentiations performed worse than simply using the aggregate visible-half marble rate.

The graveyard entry therefore prices a broader family of historical claims. Trade geography is not geology.

14.2 A direction that reversed

A second conjecture predicted that frontier military provinces would contain shorter and more formulaic inscriptions than Mediterranean administrative provinces. The visible data supported the direction; the holdout reversed it. The useful lesson is not that the visible-side pattern had been imaginary, but that it did not travel.

14.3 A regional persistence pattern that stopped

A visible pattern in Africa Proconsularis suggested a relatively gentle third-century decline in epigraphic activity and motivated a directional prediction that similar persistence would extend west into Numidia and the Mauretanias. It did not. Again, the failure concerns scope: a local historical pattern had been promoted too quickly into a regional expectation.

14.4 A religious-geography reversal that failed to replicate

The visible data suggested a counterintuitive Western concentration of Jewish and Christian Late Antique epigraphy. Four Western holdout provinces were named in advance as replications. The result was zero for four, and the conjecture died.

14.5 A death at the boundary

One dating-convention claim landed at 73.33% against a registered 75% floor. It died. Because the result lay close to the boundary, it was independently recomputed before the verdict was countersigned. This is precisely the kind of unexciting death a falsification engine needs to preserve: a threshold that becomes negotiable when an attractive claim misses narrowly is not a threshold.

14.6 Undecidable

One conjecture concerning the chronological profile of Mesopotamian inscriptions encountered only four eligible dated records. The conjecture had registered a minimum-support rule, so the evidence was insufficient for adjudication. It was neither counted as a survivor nor killed; it was recorded as undecidable. "The evidence cannot decide this claim" is a legitimate machine verdict.

15. What survived

Nineteen Heidelberg conjectures survived their registered Select tests. The heavy-tailed province-size structure recurred in the concealed half, although one component survived exactly at its threshold and is recorded as a knife-edge result. The Latin-West/Greek-East linguistic differentiation survived in both registered groups, again with one boundary result at exactly its registered floor. A Cisalpine funerary-stele conjecture, explicitly discounted at Mint because it rested on a single visible exemplar, survived in Transpadana. No concealed province approached monopoly share: the largest represented only about twelve per cent of the holdout.

These are results, but they are not inference instruments. The distinction matters because the next stage was deliberately hostile to them.

16. Grade: 208 rivals

All survivors from the first Heidelberg and census streams entered an adversarial Grade round. 37 The original Mint family was not permitted to grade its own mechanisms. Two disjoint model families generated 208 rival mechanisms against the surviving conjectures.

The effect was substantial. After the first rival family alone, eleven claims appeared capable of strengthening; after the second, only five did. No Heidelberg conjecture was strengthened, and most survivors remained interpolation-only. This is an important early result because a single frontier model's inability to imagine a good alternative explanation is weak evidence that no alternative exists. Adding another model family materially changed the epistemic status of the survivor population.

The five strengthened claims also reveal the current limits of the selection environment. They concerned mainly properties of catalogues and metadata themselves rather than substantive mechanisms in the historical world. That is unsurprising: a claim that a field is sparsely populated can often be tested directly and rivalled cleanly, whereas a claim about why Roman linguistic geography or manuscript survival has a particular form may require independent historical comparanda the engine does not yet possess.

At present Grade is therefore better at certifying facts about the record than mechanisms behind the world recorded. That is a limitation, but it is precisely the kind of limitation the architecture is intended to expose rather than conceal behind a model-generated explanation.

17. The prospective census stream

The Heidelberg test offers a known answer but carries the public-corpus memory caveat. A second stream approaches the problem from another direction. The Graphosphere census is assembling a catalogue-of-catalogues of surviving manuscript records at a scale of millions of rows, and during the first QUINCUNX programme parts of that census had not yet arrived on project systems.

A forty-conjecture batch was frozen before a set of Indian-state records was fetched. As the relevant records subsequently arrived, they became deciding evidence. Of the forty, sixteen died — two at triage, fourteen against the arriving records — and twenty-four remain alive: seventeen as holdout survivors, one resolved as supported, and six as registered-open stream conjectures whose deciding evidence has not yet arrived.

These results are prospective with respect to the project, but they are not classified as pristine discoveries because the source portal already existed publicly and prior model exposure cannot be excluded. The distinction matters: absent from our systems is not absent from the world.

Nevertheless, the census stream points towards one of QUINCUNX's strongest future selection environments. The surviving pre-print record is only partially digitized. Newly scanned collections, newly described manuscripts and newly structured catalogues continually move evidence from outside the machine-readable world into it. Conjectures frozen before that movement can eventually be adjudicated against evidence that genuinely lay beyond relevant model and project access. At sufficient scale, digitization itself becomes a rolling prospective test environment.

18. Independent databases, rival mechanisms and the dependence gate

A further registered exercise challenged surviving mechanisms against a second Roman epigraphic corpus, the Epigraphic Database Roma (EDR). Its purpose was not to produce another attractive aggregate score, but to ask whether apparent confirmation or refutation could survive contact with a genuinely separate instrument.

This immediately raised the independence problem. A planned flagship triangulation would have used several databases to arbitrate a large discrepancy in their treatment of Roman linguistic evidence. The numerical comparison was never allowed to become the headline. Before the relevant outcome was accessed, a registered construction-path audit identified enough overlap in personnel and pipeline history to fail the required independence gate, and the comparison parked.

This is an important success of procedure. A results-oriented system would have calculated the dramatic number first and attached the dependence qualification afterwards; QUINCUNX refused to create the result.

Other registered rival tests did run. Some favoured rivals, some refuted them within the available scope, and some became indeterminate or parked because the required evidence or validated instrument was absent. The distribution of verdicts carries no formal batch-level error interpretation. Its value lies elsewhere: rivals could gain ground as well as lose it, and the same frozen rules applied in both directions. The registered battery, its frozen bars and both executions:

Conjecture (claim)Rival (mechanism, one clause)Verdictd_dep gradeCaveat
-02 (no single province ≥30%)02-codex-r2 — differential completeness by provinceINDETERMINATE (mixed)0.85one of two named FAVORS conditions held, not both — the bundle's own pre-registered corner case
-0202-grok-r2 — multi-contributor cataloguing structurally prevents dominanceINDETERMINATE (mixed)0.85Rome clears 30% on both sides at once, so neither the FAVORS nor the REFUTES reading applies
-03 (heavy-tailed province sizes)03-grok-r2 — split is a cataloguing artifact, not recovery unevennessFAVORS (PARTIALLY)0.85clean under the contamination sweep, Roma-robust; scope 14 of 66 provinces, catalogue-relative
-06 (exact dating stays ≤40% everywhere)06-codex-r3 — coverage artifact of what got cataloguedREFUTES (within scope)0.467 (alt. 0.400)CONTAMINATION-SENSITIVE — the registered envelope contains both verdicts; inferentially unresolved (below)
-09 (funerary/votive share ≥85%)09-grok-r1 — stone media survive/are recovered better than perishable mediaFAVORS (PARTIALLY) → corrected: none (instrument failed validation gate)0.467 (alt. 0.400)clean; subject to the conditional missed-positive caveat below, which can weaken this FAVORS but cannot strengthen it
-0909-grok-r2 — modern collecting historically prioritized funerary/votive monumentsREFUTES (within scope) → corrected: none (instrument failed validation gate)0.467 (alt. 0.400)clean, Roma-robust
-14 (habit peaks in 1st/2nd c. AD everywhere)14-codex-r2 — catalogue-entry filter favours well-studied high-imperial materialFAVORS (PARTIALLY)0.6clean; a reading note below limits what this FAVORS can support
-18 (≥35% of text marked damaged/lost)18-codex-r2 — publications select fragmentary, editing-worthy inscriptionsFAVORS (PARTIALLY, proxy sense) → corrected: REFUTES (within scope)0.533frozen verdict stands; post-result construct assessment: cross-corpus measurement invariance failed — see below
-1818-grok-r2 — bracket use is editorial house style, loosely tied to real damageFAVORS (PARTIALLY, methodological-analogy-only) → corrected: none (proxy measured invalid)0.533capped by design; post-result construct assessment: proxy semantically unvalidated — see below
-0909-codex-r3 — funerary monuments are durable; perishable media are lostMECHANICAL PARK (pre-registered)needs a prospective, effort-weighted, total-recovery archaeological survey; a curated catalogue is not that instrument
-1818-grok-r1 — publication favours fragmentary/restored inscriptionsMECHANICAL PARK (pre-registered)needs a total-recovery sample; identical reasoning
-11 (frontier instrumentum share)11-grok-r1 — high shares reflect modern recording intensity, not ancient concentrationPARKSRaetia/Noricum return zero rows in EDR's region-label inventory
-12 (frontier diploma count)12-codex-r1 — diplomas track retirement zones, not garrisonsPARKSsame region-inventory gap (Raetia/Noricum, 0 rows)
-1212-grok-r1 — high counts reflect modern excavation intensityPARKSsame region-inventory gap (Raetia/Noricum, 0 rows)

Verdicts are identical under the blind and the corrected-instrument executions (§19) except the four marked rows.

TestRegistered bar (frozen)Landed (blind run)Blind-run verdictCorrected-instrument verdict (current best)
02-codex-r2Spearman ρ ≤ −0.5 AND ≥1 reweighted share >30%ρ = +0.26; max reweighted share 40.47%INDETERMINATE (one condition of two)identical — INDETERMINATE (mixed)
02-grok-r2EDR share >30% where EDH’s <30%, same provinceRome: EDR 40.47%, EDH 36.87% — both aboveINDETERMINATE (pre-registered corner)identical — INDETERMINATE (mixed)
03-grok-r2small/large split differs from EDH’s by >2 provinces0 small under EDR vs EDH’s 7 — differs by 7FAVORS (PARTIALLY)identical — FAVORS (PARTIALLY)
06-codex-r3EDR exact-date rate >40% in ≥1 province where EDH <40%EDR maximum anywhere: 5.00% (Liguria)REFUTES (within scope)identical — REFUTES (within scope)
09-grok-r1share drops below 85% in a majority of provincesbelow 85% in 13 of 14 (pooled 75.6%)FAVORS (PARTIALLY)none — instrument failed its validation gate; parked
09-grok-r2plurality-province fraction ≥80% refutes; ≤70% favors100% — strict majority in 14 of 14REFUTES (within scope)none — instrument failed its validation gate; parked
14-codex-r2coverage-corrected peak leaves 1st/2nd c. in a majorityoutside the window in 13 of 14FAVORS (PARTIALLY)identical — FAVORS (PARTIALLY)
18-codex-r2EDR bracket density <35% where EDH’s ≥35%EDR 0.42–2.54% everywhere; EDH 50.07–68.42%FAVORS (PARTIALLY, proxy)REFUTES (within scope) — corrected density 48.98–78.31%, above 35% in all 14 provinces (serialization repair L9); contamination-sensitive
18-grok-r2≥30% of bracketed spans classify formulaic-leaning48.65% pooledFAVORS (PARTIALLY, analogy)none — proxy measured invalid; parked
Rome Pattern-3 armall pairwise contrasts within ±2.0pp after standardizationno pair within margin under either estimandNO CALLABLE PATTERNidentical — NO CALLABLE PATTERN THIS RUN

Per the corrected-instrument replication (§19), the right-hand column is the current best measurement of each quantity. The blind run retains permanent priority as the registered blind test; where its instruments were subsequently repaired under the closed fix ledger, the corrected execution supersedes it as measurement — without being a new blind confirmation. Six of the nine analyses (and the Pattern-3 arm) are computationally identical in both runs; one verdict flipped under the serialization repair; three have no corrected verdict because their measuring instruments failed registered validation gates, and for those the blind-run verdicts stand as historical record with their published assessments, not as current claims.

ArmVerdictd_dep grade
Headline (patterns 1/2/4 — which database sits closer to which)PARKED — not computed this run at all0.250 (below the 0.4 floor)
Pattern-3 (does chronology/scope composition explain the gap?)NO CALLABLE PATTERN THIS RUN0.55

19. Corrected instruments and data cleaning as engine behaviour

The cross-database work exposed a second class of failure. Sometimes the conjecture was not the problem; the measuring instrument was.

A serialization pipeline could strip meaningful markup, a classification table could misrepresent the external vocabulary it was meant to implement, a semantic proxy could fail validation, or executable code could diverge from a frozen rule stated correctly in prose. These are not peripheral software annoyances in a found-data engine. They determine what the engine thinks the evidence says.

QUINCUNX therefore registered a corrected-instrument replication. The governing principle was deliberately narrow: repairs determined mechanically by an external specification or by the already frozen semantics of the test could be applied; judgment-dependent substitutions could not. Semantic proxies were subjected to separate blinded validation gates, and where a proxy failed its registered gate, the analysis consuming it was parked rather than repaired until it produced a result.

Proxy and registered statisticFrozen gate resultPost-gate disposition
Crude bracket detector, EDR arm: sensitivity 0.986 (99% LCB 0.964); specificity 1.000 (0.988)PASSEligible for re-execution
Crude bracket detector, EDH arm: sensitivity 0.990 (99% LCB 0.970); specificity 1.000 (0.988)PASSEligible for re-execution
Cross-corpus check: observed absolute accuracy difference 0.0179; registered descriptive ceiling 0.05PASSRegistered descriptive check passed; no invariance inference
Genre, funerary-or-votive: sensitivity 0.986 (99% LCB 0.964); specificity 0.978873 (unrounded 99% LCB 0.9494647)FAILPARKED
Material, stone: sensitivity 0.921 (99% LCB 0.885); specificity 1.000 (0.972961); negative-class denominator 168 < 250FAILPARKED
Gap/supplied as loss-versus-restoration: sensitivity 0.602 (99% LCB 0.520738); specificity 0.703 (0.499886); negative-class denominator 37 < 250FAILPARKED

Eligible analyses were then re-executed once under the closed fix ledger. Most did not move; one registered verdict reversed following a specification-determined reconstruction of stripped serialization. That reversal is not evidence that the corrected verdict is "the truth" while the first is an embarrassment to be deleted. It is evidence that the original instrument was sensitive to a data-processing defect and that an objectively specified correction changed the result. Both belong in the record.

There is a useful inversion here. An anomaly discovered by a frozen instrument can carry greater procedural warrant than the same anomaly noticed during exploratory browsing, because the record establishes that the measuring rule existed before contact with the result. Blindness does not establish what caused the anomaly: the diagnosis still requires independent checking. But when that diagnosis survives checking, the spent blindness leaves behind an auditable finding about the record or the measuring instrument itself. A failed test can therefore produce knowledge about the ground on which it failed.

The same philosophy governs the corrected EDH interval diagnostic. Found-data research cannot treat the first representation of a corpus as sacred. The requirement is not never correct, but rather: never tune silently, never erase the original contact, never confuse corrected re-execution with fresh independent confirmation, and never certify a repair solely on the evidence that selected it.

That is a stronger discipline than methodological purity imagined as immobility.

20. What the first runs taught the engine

The first programme changed QUINCUNX, as an instrument-development programme should. Among the lessons now incorporated are that prediction intervals must represent the quantity they claim to predict; declared partitions must be characterized rather than assumed exchangeable; kill-conditions require severity; minimum-support rules need an undecidable state; evaluation populations belong inside frozen test definitions; construction-path independence is an empirical property; semantic proxies require validation; data cleaning and repair require their own provenance; repaired instruments may be measured on the corpus that exposed their defect but require fresh evidence for certification; multiple rival model families materially improve Grade; and the breadth-not-height firewall must remain absolute.

Almost none of these lessons required an exotic new statistical method. Most came from having to specify exactly what was meant before contact with deciding evidence, and then allowing the resulting machinery to fail in public. Registration has therefore been more than protection against hindsight. It has become one of the engine's principal diagnostic instruments.

QUINCUNX has been pointed at itself.

21. What has now been demonstrated

The first registered programme supports a stronger claim than merely that QUINCUNX can make auditable bets. It has demonstrated both experimental and inferential proof of concept.

The engine’s five stages shaded by what has run: Mint, Select and Grade have; Promote and Infer never have. Products: a published graveyard of 22, no instruments yet.
The engine at this version. Mint and Select have run twice; Grade once, mostly withholding credit; Promote and Infer have never run. The rival battery, the corrected-instrument replication and the corrected-interval diagnostic reported in Part II postdate the counts and changed none of them. Counts countersigned 7 August 2026.

Experimentally, the machinery can freeze conjectures and predictions, preserve a holdout, reconcile deciding evidence, apply mechanical verdicts, preserve deaths, abstain where support is inadequate, generate adversarial rival mechanisms, enforce park gates, detect defects in its own data and statistical instruments, register repairs, and distinguish corrected re-execution from fresh validation.

Automation itself was part of that proof of concept. Across the Heidelberg and census batches, human involvement in adjudication had already contracted largely to review, countersignature and exceptional rulings. In the census stream, a validation runner was admitted only after reproducing the hand-adjudicated events exactly; in Heidelberg, the complete scoring pass — twenty-five conjectures, sixteen quantitative predictions, eligibility rules, minimum-support floors and near-threshold flags — was executed mechanically after the holdout arrived, with near-threshold results independently re-derived before verdict. The point is not merely convenience. Registration makes mechanical adjudication possible because the rules the machine applies have been fixed before it encounters the deciding evidence. Human attention can therefore become primarily a per-batch cost, concentrated on the quality of conjectures, instrument design and genuine exceptions, rather than a per-result cost that grows linearly with every conjecture Mint produces.

Inferentially, the Heidelberg test showed that information extracted from one part of a historical dataset can predict pre-specified properties of a second part concealed from the active run. Some point predictions were close; one headline aggregate landed exactly on the frozen estimate. More importantly, a separately registered comparison against simple structure-blind predictors showed that QUINCUNX's structural differentiation substantially improved prediction for the linguistic-geography targets, while the corresponding marble differentiation performed substantially worse than the naive rule. The system therefore did not merely generate plausible structure. It produced structure that could be scored against concealed reality, and the score distinguished useful structure from harmful structure.

The strongest concise claim we can presently make is therefore:

In its first registered masked-ground-truth test, QUINCUNX demonstrated that it can infer pre-specified properties of a concealed portion of a historical dataset from the portion available to it, with measurable predictive skill beyond naive baselines in some registered comparisons; where its structural assumptions were wrong, the same frozen comparison penalized them.

That is a proof of concept for inference from incomplete found data.

22. What has not been demonstrated

QUINCUNX has not reconstructed a genuinely lost historical world. The Heidelberg answer existed: it was concealed, not destroyed. The experiment therefore supplies an artificial lacuna, epistemically absent to the active run but ontologically available to the experimenter for scoring.

No inference instrument has yet been promoted, and no promoted instrument has yet estimated a genuinely unknown historical quantity. The interval machinery has improved under a mechanically entailed correction but is not yet certified as calibrated for general use. The present conjecture batches are too small, too heterogeneous and too dependent for their raw survivor proportions to carry a general false-discovery interpretation, and the public status of the principal corpora means prior model exposure cannot be excluded.

Nor should the engine as a population process be judged from one twenty-five-conjecture batch, one forty-conjecture census batch, or sixteen quantitative predictions. At scale, the important questions will concern the ecology of the system: how rapidly different classes of conjecture die; which mechanisms repeatedly survive Grade; how often structural differentiation beats naive prediction and how often it harms it; what fraction of promoted instruments remain calibrated on genuinely fresh material; which domains repeatedly fail to provide enough independent evidence for promotion; whether the same failure modes recur across apparently unrelated corpora; and how rapidly the graveyard grows relative to the instrument catalogue.

Those are population questions. They require a population.

23. The first construction site: the pre-print world

The written world before print is an unusually severe construction site. Its record combines catastrophic survival loss, inherited catalogues, uncertain units, uneven digitization, non-random preservation and scholarly quantities that can become canonical through repeated citation without independent recounting.

It also contains questions for which the missing denominator is central. How much was written rather than merely how much survives? How many works, textual traditions or copies have vanished? How widely distributed was a practice represented today by scattered witnesses? Which apparent absences belong to the historical world and which to later survival, discovery or cataloguing?

Many of these regions of historical terra incognita may never be directly recovered; the longer-term wager is that sufficiently independent and tested instruments can nevertheless constrain their outlines from what survives around them.

Existing scholarship already contains manually constructed inference instruments for parts of this problem: survival-factor models 38, unseen-species estimators 39, manuscript-demographic approaches 40, birth–death transmission models 41 and comparative methods in which shared loss processes cancel from ratios. QUINCUNX does not replace those methods. Its proposal is to create a common foundry in which candidate instruments can be generated or imported, subjected to a shared selection discipline, calibrated where truth is known, tested against rivals, scoped explicitly and either promoted or buried.

The aim is not to automate scholarship by replacing expert judgement with model confidence. It is to turn more of the inferential process into objects whose claims, assumptions, dependencies, failures and permissions are inspectable.

24. The wager

A falsification engine should state what would falsify its own central thesis. The thesis is not merely that language models are good at generating hypotheses; that claim may remain true even if QUINCUNX fails completely. The stronger thesis is that a general conjecture-and-inference engine can be built over found data in which the additional discipline — registration, severe selection, independence auditing, provenance, rival grading, calibration and promotion — produces better inference instruments than a simpler generate-and-filter architecture.

Two failure conditions are therefore central.

24.1 Calibration failure

On synthetic written worlds with known production and loss, promoted instruments must recover known truth at the registered minimum rate under intervals that are not merely broad enough to cover everything. Coverage must be considered together with sharpness. 42 If QUINCUNX cannot recover truth when the truth is known, its estimates where the truth has vanished have no demonstrated warrant.

24.2 The discipline stack does not pay

On the same known worlds, the full QUINCUNX stack must materially reduce the false fraction of promoted instruments relative to a naive Mint–Select baseline under a comparison frozen in advance. If independence auditing, provenance controls, Grade and promotion gates do not improve the quality of the resulting instrument population, they are ornament and QUINCUNX becomes an expensive generator of plausible claims.

Either outcome would refute the engine's central architectural thesis while leaving intact the much weaker claim that language models can generate interesting hypotheses. Mint may turn out never to have been the difficult part.

25. Conclusion

QUINCUNX began from a simple change in technological circumstance. For most of the history of automated discovery, hypotheses were expensive enough to generate that machinery was built one domain at a time around carefully specified search spaces. Large language models alter that scarcity: they can mint conjectures across almost any documented domain, operationalize them, write tests and move ideas between literatures at a scale no individual research group can match. That abundance is dangerous unless selection becomes correspondingly severe.

QUINCUNX is an attempt to build that selection machinery for the domains in which it is hardest: not clean laboratories or formal spaces with cheap oracles, but found records whose incompleteness, dependence and construction histories are themselves part of the problem. Its first registered experiment created a controlled absence in a known historical corpus and asked the machine to bet on what lay behind it.

The answer was not a triumphal sixteen-for-sixteen, and that is one reason the test is useful. Some predictions landed remarkably well; some structural distinctions materially beat naive prediction; another structurally plausible distinction performed much worse than the naive baseline and died against a landlocked province sitting on its own marble. The confidence machinery itself failed, audit found a mechanical variance omission, and correcting that defect improved measured interval performance without rescuing the remaining substantive misses. A second family of rival models stripped mechanism credit from claims a first family would have strengthened. A planned cross-database headline parked because the supposed independent witness was not independent enough. Semantic proxies failed their gates, and a corrected data instrument reversed one verdict.

None of this is peripheral to the result. It is the result. The test of an inference engine is not whether its first batch looks clever, but whether the machinery can make risky claims, expose them to evidence it cannot subsequently move, distinguish useful structure from harmful structure, identify when its own measurements are defective, preserve the correction history, refuse knowledge where the evidence is inadequate, and publish the dead beside the living.

The first QUINCUNX programme has now done enough to support a genuine proof-of-concept claim: a general LLM-based engine can extract predictive information about deliberately concealed historical data beyond naive baselines in at least some registered cases, while the same machinery exposes and preserves the cases in which its structure, mechanisms and confidence estimates fail. It has not yet earned the right to infer the genuinely lost past.

That comes next, but it will not be decided by one spectacular result. QUINCUNX is a population machine. Its power, if it has any, will appear only through repeated selection across large bodies of heterogeneous evidence: many conjectures, many deaths, many rivals, many datasets and eventually a small population of instruments that continue to work when they encounter material nobody used to build them.

Only then does the engine earn the right to point into the dark.

Data, code and audit record

The evidentiary substrate of this paper is deposited as a single versioned machine-readable package, so that no reader has to reconstruct the experiment from prose (DOI 10.5281/zenodo.21878845). The package is also mirrored here: audit package v4 (zip, ~1.2 MB) · the build-a-QUINCUNX prompt (markdown).

The package contains the registered conjectures and kill-conditions; frozen evaluation populations; timestamps, batch identifiers and commit hashes; frozen point predictions and interval definitions; holdout manifests, file hashes and reconciliation records; per-conjecture verdicts and independent re-derivations; corrected-instrument protocols and outputs; baseline definitions and scores; Grade rivals and grading decisions; construction-path dependence audits and park decisions; proxy-validation frames and results; amendment and fix ledgers; model-family and execution metadata; the frozen scoring scripts as executed (archival records of the runs; the interval-diagnostic checks and the battery’s synthetic test suite re-run standalone); machine-readable versions of the quantitative tables reported in the paper; and a reusable LLM prompt for constructing a QUINCUNX-class engine over the reader’s own found-data domain. The chronological record belongs in those artifacts. The argument belongs in the paper.

Authorship and model assistance

The central QUINCUNX concept is Stephen Pink's: mass-minting falsifiable conjectures, testing them against known evidence, subjecting survivors to further selection, and using sufficiently tested survivors as instruments for inference over incomplete records. So is the central claim that large language models may make a previously domain-specific architecture general.

The architecture and this paper have been developed through sustained interaction with several frontier language-model families, including Claude, GPT and Grok. Their part was compositional, not peripheral: in the direct sense, most of the words in this paper were composed by language models, and the byline records that fact as authorship. They also supplied statistical criticism, rival generation, prior-art review, test design and adversarial reading. The human author set the questions, supplied and checked the evidence, directed and selected at every stage, and answers for the result. Within the Claude family, Sonnet 5 did much of the implementation and mechanical scoring and Opus 5 adjudicated, mostly under Fable 5’s direction; the data package records each artifact’s acting model. That arrangement creates an obvious methodological problem for a paper concerned with model memory and correlated judgement.

The answer cannot be that the prose is persuasive. It must be procedural. A reader should be able to assign no evidential weight at all to the language models' opinions and still inspect the claims that matter: registered predictions, frozen thresholds, held-out evidence, model-family separation, audit trails, repair ledgers, baseline scores, published deaths and explicit limits on what the results are allowed to mean.

The same standard applies to QUINCUNX.

Correspondence: Stephen Pink — stephenpink@gmail.com.

References

  1. Campbell, D. T. "Blind variation and selective retention in creative thought as in other knowledge processes." Psychological Review 67:6 (1960), 380–400. (The printed title reads "retention"; the publisher's own metadata gives "retentions.") https://doi.org/10.1037/h0040373
  2. Popper, K. Conjectures and Refutations (London, 1963); Objective Knowledge: An Evolutionary Approach (Oxford, 1972). Conjectures and Refutations: https://www.routledge.com/Conjectures-and-Refutations-The-Growth-of-Scientific-Knowledge/Popper/p/book/9780415285940
  3. Lenat, D. B. AM: An Artificial Intelligence Approach to Discovery in Mathematics as Heuristic Search (PhD thesis, Stanford, 1976); "Eurisko: a program that learns new heuristics and domain concepts." Artificial Intelligence 21 (1983), 61–98. AM https://searchworks.stanford.edu/view/933075 · EURISKO https://doi.org/10.1016/S0004-3702%2883%2980005-8
  4. Fajtlowicz, S. "On conjectures of Graffiti." Discrete Mathematics 72 (1988), 113–118. https://doi.org/10.1016/0012-365X%2888%2990199-9
  5. Langley, P., H. A. Simon, G. L. Bradshaw, and J. M. Zytkow. Scientific Discovery: Computational Explorations of the Creative Processes (MIT Press, 1987). (BACON.) https://doi.org/10.7551/mitpress/6090.001.0001
  6. King, R. D., et al. "Functional genomic hypothesis generation and experimentation by a robot scientist." Nature 427 (2004), 247–252. https://www.nature.com/articles/nature02236
  7. Schmidt, M., and H. Lipson. "Distilling free-form natural laws from experimental data." Science 324 (2009), 81–85. (Eureqa.) https://www.science.org/doi/10.1126/science.1165893
  8. Romera-Paredes, B., et al. "Mathematical discoveries from program search with large language models." Nature 625 (2024). (FunSearch.) https://www.nature.com/articles/s41586-023-06924-6
  9. Novikov, A., et al. "AlphaEvolve: a coding agent for scientific and algorithmic discovery" (DeepMind, 2025); arXiv:2506.13131; open successors OpenEvolve (2025) and ShinkaEvolve (Sakana AI, 2025). https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/
  10. American Numismatic Society. "RRC 97/11." Coinage of the Roman Republic Online. http://numismatics.org/crro/id/rrc-97.11
  11. National Aeronautics and Space Administration, Marshall Space Flight Center and Kennedy Space Center, with contractors. Saturn V News Reference (August 1967, revised December 1968). https://archive.org/details/saturn-v-news-reference
  12. Galton, F. Natural Inheritance. Macmillan (1889), 64. https://archive.org/details/naturalinherita03galtgoog
  13. Browne, T. Hydriotaphia, Urne-Buriall … Together with The Garden of Cyrus. Henry Brome (1658). https://archive.org/details/hydriotaphiaurne00browuoft
  14. Chauhan, K. "Dead science walking: publication bias and the AI scientist pipeline." Proceedings of the 43rd International Conference on Machine Learning (PMLR 306, 2026); arXiv:2606.04220. https://arxiv.org/abs/2606.04220
  15. Pearl, J., and E. Bareinboim. "External validity: From do-calculus to transportability across populations." Statistical Science 29:4 (2014), 579–595. https://doi.org/10.1214/14-STS486
  16. Simonsohn, U., J. P. Simmons, and L. D. Nelson. "Specification curve analysis." Nature Human Behaviour 4:11 (2020), 1208–1214. https://doi.org/10.1038/s41562-020-0912-z
  17. Rix, H.-W., D. W. Hogg, D. Boubert, et al. "Selection functions in astronomical data modeling, with the space density of white dwarfs as worked example." The Astronomical Journal 162 (2021), 142. https://doi.org/10.3847/1538-3881/ac0c13 · injection–recovery: Suchyta, E., et al., "No galaxy left behind: accurate measurements with the faintest objects in the Dark Energy Survey" (the Balrog method), Monthly Notices of the Royal Astronomical Society 457 (2016), 786–808. https://doi.org/10.1093/mnras/stv2953
  18. Malmquist, K. G. "On some relations in stellar statistics." Meddelanden från Lunds Astronomiska Observatorium Ser. I, No. 100 / Arkiv för Matematik, Astronomi och Fysik 16:23 (1922); with "A study of the stars of spectral type A," Meddelanden Ser. II, No. 22 (1920).
  19. Hanna, R., III. William Langland. Authors of the Middle Ages 3: English Writers of the Late Middle Ages (Aldershot: Variorum, 1993). https://openlibrary.org/books/OL1398020M/William_Langland
  20. Warner, L. The Myth of "Piers Plowman": Constructing a Medieval Literary Archive (Cambridge Studies in Medieval Literature 89; Cambridge, 2014). https://doi.org/10.1017/CBO9781107338821
  21. Shumailov, I., Z. Shumaylov, Y. Zhao, N. Papernot, R. Anderson, and Y. Gal. "AI models collapse when trained on recursively generated data." Nature 631:8022 (2024), 755–759. https://doi.org/10.1038/s41586-024-07566-y
  22. Pink, S., and A. J. Lappin. “The Emergence of the Medieval Graphosphere: Voyages into the Unread and Unreadable at the Dark Archives Conferences, 2019–21.” In Dark Archives Volume I: Voyages into the Medieval Unread and Unreadable, 2019–2021, ed. S. Pink and A. J. Lappin, Medium Ævum Monographs XLIII (Oxford: The Society for the Study of Medieval Languages and Literature, 2022), 1–24. https://aevum.space/monographs
  23. Mayo, D. G. Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (Cambridge University Press, 2018). (Severity distinguished from power at p. 343.) https://doi.org/10.1017/9781107286184
  24. Benjamini, Y., and Y. Hochberg. "Controlling the false discovery rate: a practical and powerful approach to multiple testing." Journal of the Royal Statistical Society: Series B 57:1 (1995), 289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x · dependency-robust variant: Benjamini, Y., and D. Yekutieli, "The control of the false discovery rate in multiple testing under dependency," Annals of Statistics 29:4 (2001), 1165–1188. https://doi.org/10.1214/aos/1013699998
  25. Kramer, S., M. Cerrato, J. Brugger, S. Džeroski, and R. D. King. "Automated scientific discovery: from equation discovery to autonomous discovery systems." Machine Learning 115:5, article 109 (2026). https://doi.org/10.1007/s10994-025-06955-2 · preprint arXiv:2305.02251 (3 May 2023; revised 26 May 2025). https://arxiv.org/abs/2305.02251
  26. Koza, J. R. Genetic Programming: On the Programming of Computers by Means of Natural Selection (MIT Press, 1992). https://mitpress.mit.edu/9780262527910/genetic-programming/
  27. Cropper, A., and R. Morel. "Learning programs by learning from failures." Machine Learning 110:4 (2021), 801–856 · arXiv:2005.02259 (May 2020). https://doi.org/10.1007/s10994-020-05934-z · https://github.com/logic-and-learning-lab/Popper
  28. Lu, C., C. Lu, R. T. Lange, Y. Yamada, S. Hu, J. Foerster, D. Ha, and J. Clune. "Towards end-to-end automation of AI research." Nature 651:8107 (March 2026), 914–919. https://doi.org/10.1038/s41586-026-10265-5 · code https://github.com/SakanaAI/AI-Scientist-v2
  29. Ghareeb, A. E., et al. "A multi-agent system for automating scientific discovery." Nature 655:8122 (2026), 497–505. (Robin, FutureHouse; fourteen authors, led by A. E. Ghareeb and S. G. Rodriques.) https://doi.org/10.1038/s41586-026-10652-y · preprint arXiv:2505.13400. https://arxiv.org/abs/2505.13400 · code and data https://github.com/Future-House/robin
  30. Gottweis, J., et al. "Accelerating scientific discovery with Co-Scientist." Nature 655:8122 (2026), 487–496. (Google; fifty-one authors, led by J. Gottweis and V. Natarajan.) https://doi.org/10.1038/s41586-026-10644-y · preprint arXiv:2502.18864 (26 February 2025), titled "Towards an AI co-scientist." https://arxiv.org/abs/2502.18864
  31. Huang, K., Y. Jin, R. Li, M. Y. Li, E. Candès, and J. Leskovec. "Automated hypothesis validation with agentic sequential falsifications." ICML 2025; arXiv:2502.09858. (The POPPER framework.) https://arxiv.org/abs/2502.09858 · code https://github.com/snap-stanford/POPPER
  32. "Enhancing Seshat with large language models." Seshat project page (peterturchin.com), describing work by R. M. del Rio-Chanona (UCL) and J. Hauser (Complexity Science Hub Vienna) on LLM-assisted extraction into the databank; no formal publication located as of 3 August 2026. Databank https://www.seshatdatabank.info
  33. Stefanski, Radoslaw (Radek). "Cultural capital and the productivity of ideas: evidence from historical texts." University of St Andrews Economics Discussion Paper No. 2602, 23 July 2026 (ISSN 2978-4026). https://www.st-andrews.ac.uk/~wwwecon/repecfiles/econdp/2602.pdf
  34. Gill, D. J., M. Trachtenberg, M. J. Gill, P. E. Tetlock, T. K. Robb, M. E. W. Varnum, C. A. Hutcherson, I. Grossmann, and Z. Trodd. "Predicting the past: testing expert historical judgement." American Historical Review 130:4 (December 2025), 1615–1630. https://doi.org/10.1093/ahr/rhaf590
  35. "Interval-construction defect in the EDH recovery test (r2-edh-b1, Block B)." Technical record, 7 August 2026. Derivation, code loci, simulation and measurement confirming the √2 width factor, and the specified repair. docs/generated/popper_edh_interval_defect_20260807.md, deposited with this paper.
  36. “Coverage certification, measurement 2: registered binding legs and outcome (interval protocol v2.2).” Technical record, 10 August 2026. The certification-point run, the no-certificate outcome, the one-sided-criterion defect, and the registration of the two-sided successor rule; in the data and audit deposit, interval_protocol/.
  37. Fa, D., and M. Culjak. "Sound agentic science requires adversarial experiments." ICLR 2026 Workshop on Agents in the Wild; arXiv:2604.22080. https://arxiv.org/abs/2604.22080
  38. Buringh, E. Medieval Manuscript Production in the Latin West: Explorations with a Global Database (Global Economic History Series 6; Brill, 2011). https://doi.org/10.1163/9789047428640
  39. Kestemont, M., F. Karsdorp, E. de Bruijn, M. Driscoll, K. A. Kapitan, A. Chao, et al. "Forgotten books: the application of unseen species models to the survival of culture." Science 375 (2022); extension in Evolutionary Human Sciences (February 2026). https://www.science.org/doi/10.1126/science.abl7655 · extension https://doi.org/10.1017/ehs.2026.10036 · software: the `copia` package.
  40. Cisne, J. L. "How science survived: medieval manuscripts' 'demography' and classic texts' extinction." Science 307:5713 (2005), 1305–1307. https://doi.org/10.1126/science.1104718
  41. Camps, J.-B., J. Randon-Furling, and U. Godreau. "On the transmission of texts: written cultures as complex systems." PNAS Nexus (2026). https://doi.org/10.1093/pnasnexus/pgag207 · preprint arXiv:2505.19246 https://arxiv.org/abs/2505.19246
  42. Gneiting, T., F. Balabdaoui, and A. E. Raftery. "Probabilistic forecasts, calibration and sharpness." Journal of the Royal Statistical Society: Series B 69:2 (2007), 243–268. https://doi.org/10.1111/j.1467-9868.2007.00587.x

© 2026 Stephen Pink. The text and original figures of this article are released under a Creative Commons Attribution 4.0 licence (CC BY 4.0), matching the article’s Zenodo deposits. The licence does not extend to third-party images, which are credited where they appear.