The Surprise Intelligence Model
Which forecast earned your trust?
Zeno learns when forecasts deserve trust before a decision depends on them. It compares answers, track records and shared mistakes before it adjusts, warns or keeps.
forecasts
answers, ranges and timing
→track records
what was known before the decision
→shared mistakes
who has failed together
→trust decision
adjust, warn or keep
forecasts
answers, ranges and timing
track records
what was known before the decision
shared mistakes
who has failed together
trust decision
adjust, warn or keep
01 · distribution
What improved.
Every comparison uses later weeks and that source's named reference.
02 · warning
Confidence can be ranked.
Only ensemble weather beat its tuned warning comparison; that does not mean its ranges were perfectly calibrated.
warning test
Can Zeno rank which confident forecaster is more likely to miss?
coverage test
Did the stated 90% ranges actually contain about 90% of outcomes?
Known limit · coverage was about 81% (crypto), 85% (ensemble weather), 87–89% (others)
03 · serving
Proof controls the switch.
A shared model does not mean a universal claim.
Crypto price ranges
Zeno correction
+11.7%
Ensemble weather
Zeno correction
+5.6%
Deterministic weather
correction · no history
+6.2%
COVID-19 hospitalisations
correction · no history
+11.8%
Sea conditions
correction · no history
+3.2%
Prediction markets
reference kept
RSV
reference kept
FluSight
reference kept
ForecastBench
reference kept
Numinous AI forecasters
learn-only
04 · boundary
No future leaks backward.
Inputs are admitted only when they were knowable; missing history remains missing.
input
forecast set + available history
kernel
member and pair relations
output
distribution + warning + identity
05 · architecture
Start simple. Earn each correction.
The model begins exactly at the best simple forecast and moves only where earlier dates support it.
01 · inputs
Forecast set
Each forecaster's number, range or probability, admitted only if known before the decision.
02 · records
Track records
Per forecaster and task; pairs that miss together, kept for crowds above 64.
03 · kernel
Shared kernel
One network for every domain reads members, pairs and the task label.
04 · baseline
Best simple forecast
Chosen on earlier weeks. The untrained model must equal it exactly.
05 · correction
Correction × strength
Starts at zero; strength per source chosen on held-back newest dates.
06 · output
Answer
Corrected forecast where the source passed; otherwise the simple forecast, labelled.