the arena
Who sees tomorrow clearly?
Forecasts are everywhere and almost none of them are ever checked. The arena writes them down before the outcome is known, scores them when reality arrives, and makes the record public — so skill becomes visible, comparable and useful.
one real settled question · 2017-Q2 survey
Euro-area inflation (HICP) for calendar year 2019
41 forecasters answered by 2017-06-30. The first published figure (2020-01-22) was 1.196%.
| answer | value | miss |
|---|---|---|
| R0 equal-weight average | 1.67% | 0.474 |
| R0 median | 1.7% | 0.504 |
| closest single forecaster (007, known only afterwards) | 1.4% | 0.204 |
| Zeno (experimental) | no validated reading · v0.26 in research, not confirmed | |
limits and provenance
No Zeno reading exists for this question yet: v0.24 kept no per-case rows and v0.25 has not been scored. Shown when it exists, labelled experimental. Forecasters compete against reality; Zeno competes at judging when to trust them. Rankings only compare the same question type, horizon and outcome population.
Source: offdiagonal.space shared outcome panel 1.2.0, slice panel-ecb-v1 (seal 3b28c49b…c447f). Case pcase:1:09ce79b6de39821a9ae8107d1198c304. Decision time is a round publication bound (end of quarter), not a measured release date. Outcome: ecb_rtdb_2020-01-22.
weather · measured
Forecast models ranked on shared cases
Offdiagonal scores weather models on the same cases after the outcome is known. That ranking is measured, not illustrative.
open the measured rankingmarkets · watchlist
Open questions, before the outcome
Live Polymarket prices. Zeno shows no validated reading here yet; these become scored receipts only once Zeno records readings in advance.
loading…
four views on the same record
Being right on average is the least interesting thing about a forecaster
The useful question is narrower: who is reliable here, now, under these conditions. Each figure below illustrates one part of that question; each says plainly whether it has been measured yet.
Forecast trails
What each forecaster predicted, how they revised it as the day approached, and what reality finally did. The forecast line and the outcome are kept apart on purpose — a forecast is a claim, the outcome is a fact.
The forecast lines end at the last revision; the filled point standing alone at the right is the observed outcome: it rained 40% of the area.
illustrative example · a resolved rain question · not a measured result
Conditional skill atlas
A forecaster can be excellent in one situation and unreliable in another. Each cell is one forecaster in one kind of condition; selecting a cell shows how much evidence sits behind it.
| calm | changing | extreme | |
|---|---|---|---|
| Forecaster A | |||
| Forecaster B | |||
| Forecaster C |
Darker means stronger historical skill within this one question and horizon. Select a cell to see its supporting cases.
illustrative example · shape of the view, not measured skill
Shared-miss map
Agreement is not independence. Two impressive forecasters can be wrong in the same way, while a less celebrated one brings something the others miss. In a measured view, links would come from similar historical errors — this sketch only shows the shape of that idea.
A, B, C often miss the same rapid shifts.
D and E tend to be too cautious together.
F sits between the two groups and adds the most that is new.
illustrative example · nodes and links drawn by hand · not measured errors
Learning frontier
How a candidate compares with a strong reference as training proceeds. What matters is the difference and the uncertainty around it: a line above zero whose band still crosses zero has not shown anything yet.
Above zero is better than the reference. The dashed band is the uncertainty around that difference.
illustrative example · the live view is driven by the running experiment
written down first
A forecast counts only if it was recorded before the outcome, with question, horizon and units fixed.
scored identically
Same cases for every method, each question in its own units, uncertainty shown with every margin.
next proof
The next experiment replays decisions in time, starts from a strong anchor and asks whether learned interaction adds value across matched cases and seeds.
next