Home/Validation
The Evidence

A sealed engine, tested cold on a domain it had never seen.

The first completed preregistered validation applied the Time Engine to historical FDIC banking data — with no banking equations, no banking training, and no domain tuning. It measures the temporal condition that precedes and shapes outcomes, and here those measurements carried meaningful forward-looking signal concerning later failures.

Preregistered · Verdict: Partial

Evaluated on FDIC Call Report data — raw filings spanning 1999Q1–2012Q4, with a 2000Q1–2012Q4 evaluation window — across a full credit cycle, the 2007–2009 crisis, and the resolution period that followed.

0.891
AUROC discriminating eventual failures (95% CI 0.869–0.910)
+0.438
AUROC advantage over the naive capital baseline (0.453)
8,775
community-commercial banks · 166,968 scored bank-quarters
No banking equationsNo domain retrainingNo hand-tuningHorizon: 4 quarters (1 year)
Why it counts

Frozen first. Scored second.

The result is credible because of what happened before any score was computed. The validation protocol, the engine, the domain mapping, and the pass/fail criteria were all fixed and cryptographically hashed in advance. The engine had no access to the raw financial variables — only an abstract, pre-specified canonical representation. Nothing was adjusted after scoring began.

Frozen

Sealed framework

The engine and its logic were fixed and hashed before the banking data was ever scored.

Blind

Abstract inputs only

The engine saw canonical temporal state — not banks, not balance sheets, not what any variable meant.

Preregistered

Criteria set in advance

Success, non-inferiority, and kill criteria were defined before results existed.

What it did — and didn't — show

A partial result, reported honestly.

The engine cleared its discrimination threshold decisively (AUROC 0.891, well above the 0.75 bar; no kill criterion triggered). It did not meet preregistered non-inferiority against a purpose-built banking model (CAMELS logistic, AUROC 0.964, ΔAUROC −0.073). Under the protocol, that records the outcome as Partial — and that honesty is the point: a domain-specific model tuned for banks still edges out a general engine that knew nothing about banking. That the general engine came this close, cold, is the signal.

This is the first completed formal validation of the Time Engine. It documents the platform's universal capability empirically in one domain; it does not define the platform's scope. Additional cross-domain validations are underway.

The numbers, defined

Every figure, reconciled to the report.

Two counts are easy to confuse: 574 is the number of failed institutions; 737 is the number of failure-labeled bank-quarters in the scored set. A failing bank contributes several failure-labeled quarters within the one-year horizon, so the two counts differ by design.

SourceChicago Federal Reserve — FDIC Call Report (MDRM) extract
Raw filings466,585 bank-quarters (1999Q1–2012Q4)
Community-commercial cohort343,998 bank-quarters across 8,775 unique banks
Evaluation window2000Q1–2012Q4
Prediction horizon4 quarters (1 year ahead)
Scored evaluation set166,968 bank-quarters
Documented bank failures574 institutions (FDIC seizure events)
Failure-labeled bank-quarters737 positive observations (0.44% base rate)
AUROC0.891 (95% CI 0.869–0.910)
Baseline — naive capital0.453
Baseline — CAMELS logistic0.964
Non-inferiority vs CAMELSΔAUROC −0.073 (CI −0.089 to −0.054) — not met
AUPRC0.043 (≈ 9.8× the 0.44% base rate)
Preregistered verdictPartial