Validation

SI vs predictive uncertainty

The claim under test: a learned representation of model-relative surprise contains information about future regime change and model failure that is not captured by predictive uncertainty alone.

Demo modelSynthetic dataResearch preview

Protocol — for every domain and every step of the synthetic scenario clock, the SI head P(model failure at t+h) is scored against five conventional detectors computed from the same references. Steps at or after t=9 are labelled consequential (structural break, reference re-ordering). Metrics: ranking AUC, mean lead time before onset at a self-calibrated alarm threshold (mean + 1.5σ over pre-onset steps, sustained two steps), Brier score and precision@8.

synthetic corpusfrozen checkpoint0 scored steps0 positives
SI ranking AUC
0.000
vs labelled consequential steps
Lead time advantage
0.0 steps
best baseline: —
SI Brier
0.000
lower is better

Method comparison

Sorted by ranking AUC across all three domains.

methodAUClead time (steps)Brierprecision@8

Score trajectories

Where each detector first departs from baseline relative to the labelled onset.

SI (z^S head)Ensemble variancePredictive entropyOOD scoreAnomaly scoreConformal width

What this does and does not establish

Established here
  • — The output contract and evaluation harness exist end-to-end.
  • — SI is comparable against five uncertainty baselines on identical inputs.
  • — Lead time, not just separation, is a first-class metric.
Not yet established
  • — Results are on a synthetic, deterministic corpus with a frozen checkpoint.
  • — No trained weights, no held-out real events, no cross-domain transfer test.
  • — The model family claim requires generalization of z^S across situations and then domains.