Skip to main content
Harshwardhan Patil

Projects

Systems that show their working.

Four builds where the interesting part isn't the model. It's everything around it that makes the output checkable. Numbers below are measured, and where something is built but unrun, or specified but unbuilt, it says so.

Cascade: Multi-Agent Causal Simulation

Case study 01

A forecasting system whose retrieval layer is physically unable to read the future, and a test suite that proves it.

Milestones M0 through M9 complete · headline study results await an API credential

Python 3.12PostgreSQL 16pgvector (HNSW)LangGraphAnthropic APIbge-small-en-v1.5LangfuseDuckDBDocker Composepytest + Hypothesis

Every backtested forecasting system has the same credibility problem: you cannot tell, from the outside, whether the model was clever or whether it simply saw the answer. Leakage does not announce itself. It arrives through a default parameter, a filter someone removed during debugging, a document whose crawl date is not its publication date. So Cascade is built around the assumption that I will eventually make that mistake, and the job of the architecture is to make it impossible rather than unlikely.

Retrieval runs through a function called Chronofence, which takes an as-of timestamp and has no default for it. A missing cutoff is a type error in Python and a signature error in SQL, not a silent full-corpus search. The outcome labels live in a separate table that the simulation's database role has no grant on at all, so an attempt to read them raises a Postgres permission error rather than returning a row. And the whole thing is checked adversarially: 500 synthetic post-resolution documents were planted in the corpus and queried at all 180 scenario cutoffs using the poison's own text, the most favourable query it could possibly get. Zero came back. A positive control confirms they are retrievable the moment the time filter is lifted, which is what makes the zero mean something.

Above that sits the actual system. A question becomes a typed causal graph through a draft, critique, repair loop with six validator rules. A 24-step kernel runs it, with each actor seeing a partial, noisy, lagged view of the world decided by their position in the graph, and a pure-Python arbiter with no model call in it folding their actions into the next state. Two hundred seeded replicates per scenario are batched one submission per step across the whole wave, which turns 864,000 calls into 24. Scoring is written from the definitions rather than imported, because the dependency set has no scipy: Brier, Murphy decomposition, mid-rank AUC, ten-bin calibration, a paired bootstrap at B = 10,000 with Holm-Bonferroni across a 12-cell ablation grid.

Underneath is ordinary engineering held to an unusual standard. The corpus is 1,950,912 chunks over 492,270 documents, verified to contain no null, future, or timezone-naive dates anywhere in the table, checked with unqualified aggregates rather than a sample. Recall@20 is 0.9675 against exhaustive search. Twenty-five stored runs re-run in fresh interpreters under a different hash seed and reproduce their event-log hash byte for byte, and any outcome walks back to the exogenous shock that caused it in 0.27 seconds. There is exactly one place in the codebase that calls the model API, and CI fails if a second appears.

What the repository will not do is quote a result it does not have. Retrieval p95 is 90.92 ms against a 15 ms budget, the benchmark still exits non-zero, and both numbers are printed rather than softened. No causal graph has been compiled by a model yet, because that needs 540 API calls and a credential the build environment does not have, so there is no Brier score and no ablation delta on this page either. Runs made with the deterministic stand-in are stamped as such by a database column, not a footnote, and a static check fails the build if any target value is ever written into a code path that produces a report.

1,331tests passing1.95Membedded chunks over 492K documents0.9675recall@20 vs exhaustive search0 / 500poison documents retrieved25 / 25runs replay byte-identically180sealed backtest scenarios
Cascade simulation architectureA scenario passes through four stages in sequence: a Compiler that turns free text into a typed causal graph, Chronofence for time-locked retrieval, per-actor Agents operating under information asymmetry, and a pure-Python Arbiter that resolves each step. Chronofence is fed by a PostgreSQL and pgvector corpus holding only pre-cutoff evidence, and the Arbiter returns control to the Agents for each of the 24 simulation steps.Compilerscenario → causal graphChronofencetime-locked retrievalAgentsact under asymmetryArbiterresolves, no LLMPostgreSQL + pgvectorpre-cutoff evidence only24-step loop
Every stage here is built and tested. What has not run is the study itself: compiling the 180 causal graphs needs an API credential, so there is no Brier score on this page yet. The arbiter is deliberately the one stage with no model call in it.

Architecture

Scenario ledger
180 resolved binary questions loaded from Polymarket, Manifold, and a curated file, selected under hard rules (a 0.50 base rate, no domain above 25%) and sealed behind a SHA-256 manifest that fails verification if a single label moves. The simulation role has no SELECT grant on the labels table at all.
Evidence corpus
1,950,912 embedded chunks from 492,270 documents across CC-NEWS, Wikipedia, government press releases, and SEC EDGAR, stored in quarterly pgvector partitions with 100% embedding coverage and not one null, future, or timezone-naive publication date.
Chronofence retrieval
A SECURITY DEFINER Postgres function that takes an as-of timestamp and cannot be called without one. Nothing published on or after the cutoff can be returned, which is enforced by the query plan rather than by a filter someone might forget.
Lathe and Loom
A compiler that turns a free-text question into a typed causal graph through draft, critique, repair, and six validator rules; then a 24-step kernel where actors see a partial, noisy, lagged view of the world and a pure-Python arbiter with no model call folds their actions into the next state.
Chorus and Assay
200 seeded replicates per scenario, batched one submission per step across the whole wave so the study makes 24 calls instead of 864,000, then scored on Brier, Murphy decomposition, calibration, and mid-rank AUC with a paired bootstrap over a 12-cell ablation grid.
Strata determinism
One LLM call site, content-addressed record and replay, exact-Decimal cost metering against a hard ceiling, and one RNG per run seeded from a keyed hash. 25 stored runs re-run in fresh interpreters reproduce their event-log hash byte for byte, and any outcome walks back to its root cause in under a second.

Pipewright: Data Pipeline Platform

Case study 02

The operational half of data work, from connector to lineage to write-back, as one product instead of nine scripts.

Sixteen phases shipped across two roadmaps · running locally end to end

Next.js 16React 19TypeScriptFastAPISQLAlchemy 2.0AlembicPostgreSQLpandas + pyarrowDocker ComposepytestVitestGitHub Actions

Most of what makes data work hard is not the transformation. It is everything around it: where did this column come from, which run produced this table, did the schema change under us last Tuesday, who gets told when the publish fails. That work usually lives in a drawer of disconnected scripts and a spreadsheet nobody trusts. Pipewright is an attempt to put it in one place and treat it as product surface rather than plumbing.

The bet the platform is built on is that analysts export to Excel because Excel lets them see and touch the data, and engineers hate the export because it forks the truth. So the grid edits like a spreadsheet, on millions of rows, but every cell edit is a step in a recipe rather than a mutation. Underneath it is a canonical type lattice and a relational IR with pandas and SQL backends, proven identical over a differential corpus, which is what lets a spreadsheet formula get type inference, column lineage, and pushdown to the source without being a special case. The planner splits a query tree into what the source can run and what the platform runs, and refuses to approximate that boundary.

Getting data in is its own problem, and the part I am most attached to. File reading is a pipeline where every stage reports what it decided, how confident it is, and the evidence behind it; a stage that genuinely cannot tell whether 03/04/2025 is March or April returns ambiguous and blocks instead of picking one. Each decision is stored against the file's column fingerprint, so next month's version of that file reads the same way rather than being inferred again from different rows. A 72-file corpus of deliberately awful spreadsheets backs all of it, and it found seven real bugs, including every import silently truncating at 5,000 rows and identifiers losing their leading zeros.

The same honesty rule runs through the connector catalogue. There are 211 connectors, which is a meaningless number on its own, so every one carries a verification tier, its verified_by field cites a test file, and a test checks that the citation resolves and that the named file actually drives that connector. Fifty are earned against real systems. The other 161 were written from vendor documentation and never executed here, and they say so on the card, in the config form, in the connection test, and in the warnings attached to a run. Nine engines the platform genuinely cannot drive each name the interface it can read instead.

Structurally it is a modular monolith: one FastAPI gateway is the only public deployable, and 24 domain packages sit behind it with real boundaries, cross-service calls going through registered hooks rather than imports. Permission is decided in exactly one place for the whole API and fails closed on an unrecognised write path. The repository is deliberate about its edges, and the handoff document keeps a list of them: write-back has only been run against SQLite because there is no server here for the Postgres and MySQL paths, notebook Python cells are disabled on macOS because the sandbox probe cannot enforce a memory limit, and streaming and CDC are a different execution model rather than another line in the catalogue.

24bounded service packages211connectors, 50 verified against real systems173transformation tools in the library5,943Python tests652web tests30forward-only migrations

Architecture

Studio
A Next.js App Router workspace where the grid edits like a spreadsheet: two-axis virtualisation, Excel keybindings, clipboard, column profiling, and a formula engine that compiles to the same intermediate representation everything else runs on, so a typed formula gets lineage and pushdown for free.
Type lattice and IR
A canonical type system and a relational IR with pandas and SQL backends, proven identical over a differential corpus, plus a pushdown planner that splits a query tree into what the source can run and what the platform runs, and refuses to approximate the boundary.
Ingestion intelligence
File reading as a pipeline where every stage reports what it decided, how confident it is, and the evidence behind it. A stage that genuinely cannot decide returns ambiguous and blocks rather than defaulting. Each decision is saved against the file's column fingerprint, so next month's file reads the same way instead of being re-inferred from different data.
Connector catalogue
211 connectors from four generators, each carrying a verification tier where verified_by cites a test file and a test checks that the citation resolves and that the file actually drives that connector. 50 are earned against real systems; the other 161 say so on the card, in the config form, in the connection test, and in the run's warnings.
Domain services
24 Python packages with real boundaries behind a single FastAPI gateway: workflows, lineage, observability, access, governance, connectors, reporting, intelligence, enterprise, quality, write-back and the rest. Cross-service calls go through registered hooks rather than imports, and permission is decided in exactly one place for the whole API.
Write-back and operations
Row identity resolution, staged change sets, a dry run inside a rolled-back transaction, a blast-radius guard, and migration export. Around it sit Postgres-leased schedules, incident grouping, notifications, a system-status surface, and a GitHub Actions pipeline that runs exactly the same verification path as the local make verify.

Metric Movement Diagnostics Platform

Case study 03

Answers the question every dashboard raises and none of them answer: why did that number move?

Shipped · validated against a known +52.1% spike

PythonPostgreSQLClaude APIBlock BootstrapStreamlitPlotlypytestPlaywright

A dashboard tells you revenue jumped 52.1% month over month. It will never tell you why, and the analyst who has to explain it by Monday spends the weekend slicing the same table by hand. This platform does that investigation.

Underneath is a PostgreSQL star schema over 112,000-plus transactions, which is what makes systematic slicing possible in the first place. On top of it runs a block bootstrap at 2,000 iterations, resampling in blocks so temporal structure survives, to establish which apparent drivers are real and which are variance wearing a convincing costume. I validated the whole approach against that +52.1% month-over-month spike specifically, because a diagnostic tool that cannot explain a movement you already understand has no business explaining one you do not.

The reasoning layer is a three-agent Claude pipeline: an Investigator that proposes candidate explanations, an Analyst that tests them against the statistical evidence, and a Critic that fact-checks every claim against the source data before a human ever sees it. Every hand-off is JSON-schema constrained, so the stages compose instead of drifting, and the agents sit behind a hard boundary where aggregates cross and raw records and PII never do. Correctness is enforced by 63 pytest tests, and the Streamlit and Plotly interface is verified end to end with Playwright.

112K+transactions in the star schema2,000bootstrap iterations per diagnosis+52.1%MoM spike used as the validation case63automated tests

Architecture

Warehouse
A PostgreSQL star schema over 112,000-plus transactions, modelled so a metric can be sliced systematically along every dimension rather than whichever ones someone thought to index.
Statistical engine
A block bootstrap at 2,000 iterations. Resampling in blocks preserves temporal structure, which is what separates a driver that is real from one that is variance in a convincing costume.
Reasoning pipeline
Three Claude agents in sequence: an Investigator proposing candidate explanations, an Analyst testing them against the statistical evidence, and a Critic that fact-checks every claim against source data before a human sees it. Hand-offs are JSON-schema constrained so stages compose instead of drifting.
Data boundary
Aggregates cross into the model context; raw records and PII never do. The boundary is a property of the pipeline, not a prompt instruction.
Interface
A Streamlit and Plotly application, verified end to end with Playwright and backed by 63 pytest tests.

Credit Card Fraud Decisioning System

Case study 04

Fraud modelling stops at a score. Fraud operations start there, so this one is built around the queue rather than the AUC.

Complete · eight-stage pipeline with a Streamlit review surface

PythonScikit-learnPandasJupyterStreamlitMatplotlib

A fraud model that is 99% accurate is not an achievement, it is a description of the class balance. What a fraud team actually needs to know is different: how many transactions land in the review queue tomorrow, whether the analysts can clear them, and which mistakes cost the most. So this project treats fraud as a decisioning problem and the classifier as one component inside it.

The pipeline runs in eight stages, each with a job it does not share. A data audit validates schemas and compares train/test structure before any feature exists, because a structural mismatch discovered during modelling gets absorbed instead of fixed. Feature engineering builds transaction, customer, merchant, temporal, and behavioural signals with preprocessing kept deliberately separate from the model fit, which is the boundary that keeps leakage out. Modelling compares supervised baselines against stronger approaches and anomaly-based components in a champion/challenger frame.

Then it stops being a modelling project. Scores are converted into a three-band operating policy of approve, review, and high-risk, chosen by evaluating review volume against capture rate, which is a capacity question rather than a statistical one. Error analysis looks specifically at where the system fails and which segments generate operational drag. The Streamlit dashboard is built for the person running the queue: threshold grids and policy comparison, analyst capacity simulation, alert prioritisation, false-positive and false-negative breakdowns, and diagnostics on the transactions sitting right at the threshold edge, where the policy is actually decided.

Architecture

Data audit
Schema validation, missingness inspection, and train/test structure comparison before a single feature is built. This is the stage that catches the problems modelling would otherwise absorb.
Feature engineering
Transaction, customer, merchant, temporal, and behavioural signals, kept leak-safe by holding preprocessing and modelling responsibilities apart rather than blurring them into one fit.
Modelling
Supervised baselines against more advanced approaches plus anomaly-based components, framed as champion/challenger so a replacement has to earn the slot.
Threshold strategy
Scores converted into an operating policy: approve, review, and high-risk bands chosen against review volume, capture trade-offs, and analyst capacity rather than a single F1-optimal cut.
Error analysis
Where the system fails and who pays for it: false positives, false negatives, edge cases, and the segments that quietly create operational drag.
Review surface
A Streamlit dashboard for KPI and trend monitoring, threshold grids, analyst capacity simulation, model comparison, alert-queue prioritisation, and error concentration.

More on GitHub

The repositories behind the case studies above, plus smaller experiments and analysis notebooks, live on GitHub.

github.com/harshkvpatil98