SI vs predictive uncertainty
The claim under test: a learned representation of model-relative surprise contains information about future regime change and model failure that is not captured by predictive uncertainty alone.
Protocol — for every domain and every step of the synthetic scenario clock, the SI head P(model failure at t+h) is scored against five conventional detectors computed from the same references. Steps at or after t=9 are labelled consequential (structural break, reference re-ordering). Metrics: ranking AUC, mean lead time before onset at a self-calibrated alarm threshold (mean + 1.5σ over pre-onset steps, sustained two steps), Brier score and precision@8.
Method comparison
Sorted by ranking AUC across all three domains.
| method | AUC | lead time (steps) | Brier | precision@8 |
|---|
Score trajectories
Where each detector first departs from baseline relative to the labelled onset.
What this does and does not establish
- — The output contract and evaluation harness exist end-to-end.
- — SI is comparable against five uncertainty baselines on identical inputs.
- — Lead time, not just separation, is a first-class metric.
- — Results are on a synthetic, deterministic corpus with a frozen checkpoint.
- — No trained weights, no held-out real events, no cross-domain transfer test.
- — The model family claim requires generalization of z^S across situations and then domains.