model assurance
Know when your models stop being right.
Zeno Divergent watches a live stream and warns you when the forecast, risk engine or monitor you depend on is losing its validity — before the miss shows up as a loss.
open weights on Hugging Face · open benchmark protocol
purpose
Why a model of surprise, and why now
Every consequential decision today is downstream of a model — a flood forecast, a risk engine, a grid schedule, an LLM policy. Those models are trained on the world as it was, and they fail silently when it moves. Zeno Divergent exists to make that moment predictable and legible before the loss lands.
the problem
Failure is discovered downstream
A reference model keeps emitting confident numbers well past the point where its assumptions hold. Nothing in its own output says 'I no longer apply here'. The break is found in the post-mortem, in the incident review, in the loss.
the object
Model validity as a forecastable quantity
Zeno Divergent takes the reference models themselves as inputs and learns a latent surprise state over the gap between what they expect and what arrives. From that one vector it reads the kind of divergence, what could not be observed, how long each reference stays valid, and whether to abstain.
the consequence
A decision you can defend
You get a calibrated probability with a named type and horizon, and an action — act, watch, ask for more information, or stop trusting this model. Abstention is a first-class answer, which is what separates assurance from another anomaly score.
pipeline
From a public feed to a decision you can defend
Four stages, each one inspectable in the workspace. Nothing is imputed to fill a gap: an unobserved channel stays unobserved, and the model is allowed to abstain.
Public feeds, persisted
USGS event and streamflow services, Coinbase bars, NOAA SWPC indices, FRED series and GHCN-Daily stations land in the production corpus with their own units and cadence.
corpus_series · corpus_observations
Onto one cadence grid
Entities are aligned to a shared 12-step window and joined with the reference predictors they are meant to obey. Channels a feed does not carry are marked unobserved, never imputed.
72 features · cadence, magnitude, evidence, residual
Calibrated, per domain
A per-domain adapter feeds a shared dilated causal trunk. Outputs are Platt-scaled, so a stated probability means what it says, and abstention is an available answer.
v0.7.0 · 1.87M parameters
Typed, not a single score
A SurpriseReport names the kind of divergence, what could not be seen, validity per horizon, and the next action — including request-information.
surprise latent · reliability · risk–coverage
what you get
A SurpriseReport, not a single number
Every scored entity comes back in the same typed contract, so the answer can be acted on and audited instead of interpreted.
- divergence
- which kind of break it is — magnitude, timing, structure, mechanism or coverage.
- validity horizon
- how long the reference you rely on is still expected to hold.
- blindness
- what could not be observed, kept separate from what looks calm.
- confidence
- a calibrated probability, or an explicit abstain when it cannot tell.
- next action
- act, watch, request information, or stop trusting this model.
example
entity gauge-11447650 divergence structure · cadence break validity reference holds ≤ 6h blindness 2 of 4 channels unobserved confidence 0.81 (calibrated) action request information
structure
The hierarchy
The model is the intellectual centre. Everything above it is subordinate to it.
Surprise Intelligence
the research direction
Open-world model assurance: making the edge of a model's validity a forecastable quantity rather than a post-mortem finding.
Zeno Divergent
the learned model
Models are first-class inputs: it attends over expectations, not features. It predicts future model validity, and can abstain instead of scoring.
SurpriseBench
the benchmark
Held-out episodes, locked thresholds, degenerate controls, label-shuffled copies and leave-one-mechanism-out transfer. Built to return a negative verdict — and it has.
SurpriseReport
the interface
One typed contract: kind of divergence, blindness, validity per horizon, a surprise-potential ranking, and an action — including request-information.
kernel
Six primitives, one kernel
Every domain pack shares the same reasoning machinery: six primitives that turn a set of reference expectations into a typed account of where they are breaking.
Reference graphs
Several independent expectations per question, with disagreement kept explicit instead of collapsed into consensus.
Typed surprise
Surprise is classified — magnitude, timing, structure, mechanism, coverage — never reduced to one anomaly score.
Unknown / blindness
A first-class channel for what could not be observed, separating low risk from low visibility.
Credible tail search
Bounded search over admissible futures that the current references do not carry.
Model validity
Admission control decides whether a model may still speak: ignore, remember, open a regime, ask, rebuild.
Surprise memory
Events are stored with magnitude and novelty apart, so recurrence is never mistaken for novelty.
domains
Three verticals, one cross-cutting capability
Earth Observation, Finance & Web3 and Critical Infrastructure & Defence are the worlds the model learns in, each keeping its own units, references and unit of consequence. AI Meta-Evaluation & Safety is not a fourth industry: it is the same machinery turned on models themselves.
nothing here is live in production — maturity is stated per sub-area below.
vertical product families
near = implemented · build = needs a tenant and labels · research = weak labels or physics
Three industries where the consequence of a stale model is measured in a real unit — hectares, money, or service-hours. Each family owns its references, its constraints and its own unit; they are never blended into one score.
Earth Observation
unit · physical magnitude × exposure
Tells you when the physical forecasts you rely on — fire, flood, heat, drought — have drifted outside the conditions they were fitted on, and where to look next.
- wildfire & fuel statebuild
- extreme weatherbuild
- flood & hydrologybuild
- drought & vegetationbuild
- ocean & sea-surface temperatureresearch
- land-use & urban changeresearch
- sensor tasking by expected surprise reductionresearch
refs: physics model · ensemble · persistence · satellite retrieval
Finance & Web3
unit · money
Flags the moment pricing, credit and execution models stop describing the market they are trading in — before the loss, while there is still time to widen, hedge or stand down.
- HFT microstructure fragilitybuild
- onchain protocol & liquidity riskbuild
- uncollateralized credit (lender-side)build
- treasury & liquiditybuild
- desk / market risk overlaybuild
- payroll & payout railsnear
- mempool & flash-loan interceptionresearch
- execution intent protectionresearch
- insurance & catastrophe pricingresearch
refs: originator model · peer cohort · cash-vs-tape · venue order book · latency / fill model · yesterday's model
Critical Infrastructure & Defence
unit · safety & continuity
Watches the detectors, plant models and fused pictures that keep essential systems running, and separates a quiet estate from an unobserved one.
- pre-zero-day threat huntingbuild
- identity-model validity / UEBAbuild
- OT & industrial controlbuild
- grid & energy cascadesbuild
- detection-model monitoringnear
- water & municipal systemsresearch
- defence sensor-fusion picture assuranceresearch
- autonomous swarm coordinationresearch
refs: role cohort · change calendar · plant grammar · endpoint telemetry health · sensor fusion
cross-cutting capability
One capability that sits across every vertical instead of beside them: judging models themselves. Its unit is trust, not consequence, and its subject can be any predictor — ours or someone else's.
AI Meta-Evaluation & Safety
unit · model trust
Not an industry — a capability that applies to all of them. Point it at any model, in any of the verticals above or outside them, and it grades how far that model can still be trusted and when it should abstain.
- third-party fragility audit (overconfidence index)near
- detection / forecast model passportsnear
- OOD guardrails on production telemetrybuild
- continuous production monitoring loopbuild
- agent behaviour & tool-use driftresearch
refs: subject model logits · held-out episodes · production telemetry · degenerate controls
SurpriseBench — the shared protocol
All three verticals and the cross-cutting pack are graded under the same frozen protocol, and it is the part that exists today: grading other models' stale worlds. Any external predictor can be run as the subject — typed miss, lead time at a fixed false-alert rate, calibration and degeneracy guards, with mandatory constant, random and shuffled-label controls. It is domain-independent and it publishes its losses.
open SurpriseBench →coupling edges — unit conversion, not a universal risk score
fire / drought Ψ → catastrophe price and borrower cash-flow
km² → USD
heat + drought → grid reserve margin and trip potential
°C → service-hours
outage → payroll, working capital and insured loss
service-hours → USD
detector degradation → downstream model trust and abstention rate
coverage → trust
proof — the falsifiable claim
A learned representation of model-relative surprise contains information about future regime change and model failure that predictive uncertainty alone does not capture.
Stated so it can fail. The surprise index and eight conventional detectors are scored on identical inputs and labels, with thresholds locked before the run; if a baseline matches it, the harness says the claim is unsupported.
- latest release
- v0.7.0 · v0.7.0
- promoted domains
- hydrology, climate
- promotion gate
- promoted
- parameters
- ≈1.87M
- in-browser scoring
- v0.7.0 + v0.5.0 ONNX heads
- not promoted
- seismic · markets · spaceweather · macro · wildfire · chaos
compounding
How the understanding compounds
One loop, repeated per release: add real corpora, declare the gate before training, promote only what clears it, freeze everything so the next comparison is honest.
01 · widen
Each release adds real public corpora that share no physics. A latent that survives seismicity, streamflow, markets, space weather and climate at once is encoding departure itself, not one domain's habits.
02 · gate
Thresholds, degenerate controls and leave-one-domain-out transfer are fixed before the run. v0.7.0 promoted 2 of 8 domains; v0.4.0 promoted none and is published as a failure.
03 · learn from failures
v0.7.0 promotes fewer domains than v0.6.0 because the gate now demands a win over a plain logistic regression on the same features. hydrology and climate clear it; the rest ship as candidates with their blocking reason attached.
04 · freeze
Every release is an immutable Hugging Face revision carrying checksums, calibration parameters and the full gate report — so improvement is checkable by anyone, release over release.
limits
What it cannot do yet
This is research in the open, not a finished product. The current release clears its gate on hydrology, climate and does not on seismic, markets, spaceweather, macro, wildfire, chaos; no vertical is running in production, and the claim above is not yet demonstrated beyond those two domains — everywhere else a plain logistic baseline is still the honest choice. Every failure is published alongside the wins rather than removed from the table.
read the full evidence page, failures included →get started
Point it at a stream you care about
Upload a CSV or connect a feed, and the workspace returns a scored SurpriseReport for every entity in it.