Cascade: Multi-Agent Causal Simulation
Case study 01A forecasting system whose retrieval layer is physically unable to read the future, and a test suite that proves it.
Milestones M0 through M9 complete · headline study results await an API credential
Every backtested forecasting system has the same credibility problem: you cannot tell, from the outside, whether the model was clever or whether it simply saw the answer. Leakage does not announce itself. It arrives through a default parameter, a filter someone removed during debugging, a document whose crawl date is not its publication date. So Cascade is built around the assumption that I will eventually make that mistake, and the job of the architecture is to make it impossible rather than unlikely.
Retrieval runs through a function called Chronofence, which takes an as-of timestamp and has no default for it. A missing cutoff is a type error in Python and a signature error in SQL, not a silent full-corpus search. The outcome labels live in a separate table that the simulation's database role has no grant on at all, so an attempt to read them raises a Postgres permission error rather than returning a row. And the whole thing is checked adversarially: 500 synthetic post-resolution documents were planted in the corpus and queried at all 180 scenario cutoffs using the poison's own text, the most favourable query it could possibly get. Zero came back. A positive control confirms they are retrievable the moment the time filter is lifted, which is what makes the zero mean something.
Above that sits the actual system. A question becomes a typed causal graph through a draft, critique, repair loop with six validator rules. A 24-step kernel runs it, with each actor seeing a partial, noisy, lagged view of the world decided by their position in the graph, and a pure-Python arbiter with no model call in it folding their actions into the next state. Two hundred seeded replicates per scenario are batched one submission per step across the whole wave, which turns 864,000 calls into 24. Scoring is written from the definitions rather than imported, because the dependency set has no scipy: Brier, Murphy decomposition, mid-rank AUC, ten-bin calibration, a paired bootstrap at B = 10,000 with Holm-Bonferroni across a 12-cell ablation grid.
Underneath is ordinary engineering held to an unusual standard. The corpus is 1,950,912 chunks over 492,270 documents, verified to contain no null, future, or timezone-naive dates anywhere in the table, checked with unqualified aggregates rather than a sample. Recall@20 is 0.9675 against exhaustive search. Twenty-five stored runs re-run in fresh interpreters under a different hash seed and reproduce their event-log hash byte for byte, and any outcome walks back to the exogenous shock that caused it in 0.27 seconds. There is exactly one place in the codebase that calls the model API, and CI fails if a second appears.
What the repository will not do is quote a result it does not have. Retrieval p95 is 90.92 ms against a 15 ms budget, the benchmark still exits non-zero, and both numbers are printed rather than softened. No causal graph has been compiled by a model yet, because that needs 540 API calls and a credential the build environment does not have, so there is no Brier score and no ablation delta on this page either. Runs made with the deterministic stand-in are stamped as such by a database column, not a footnote, and a static check fails the build if any target value is ever written into a code path that produces a report.
Architecture
- Scenario ledger
- 180 resolved binary questions loaded from Polymarket, Manifold, and a curated file, selected under hard rules (a 0.50 base rate, no domain above 25%) and sealed behind a SHA-256 manifest that fails verification if a single label moves. The simulation role has no SELECT grant on the labels table at all.
- Evidence corpus
- 1,950,912 embedded chunks from 492,270 documents across CC-NEWS, Wikipedia, government press releases, and SEC EDGAR, stored in quarterly pgvector partitions with 100% embedding coverage and not one null, future, or timezone-naive publication date.
- Chronofence retrieval
- A SECURITY DEFINER Postgres function that takes an as-of timestamp and cannot be called without one. Nothing published on or after the cutoff can be returned, which is enforced by the query plan rather than by a filter someone might forget.
- Lathe and Loom
- A compiler that turns a free-text question into a typed causal graph through draft, critique, repair, and six validator rules; then a 24-step kernel where actors see a partial, noisy, lagged view of the world and a pure-Python arbiter with no model call folds their actions into the next state.
- Chorus and Assay
- 200 seeded replicates per scenario, batched one submission per step across the whole wave so the study makes 24 calls instead of 864,000, then scored on Brier, Murphy decomposition, calibration, and mid-rank AUC with a paired bootstrap over a 12-cell ablation grid.
- Strata determinism
- One LLM call site, content-addressed record and replay, exact-Decimal cost metering against a hard ceiling, and one RNG per run seeded from a keyed hash. 25 stored runs re-run in fresh interpreters reproduce their event-log hash byte for byte, and any outcome walks back to its root cause in under a second.