model assurance

Know when your models stop being right.

Zeno Divergent watches a live stream and warns you when the forecast, risk engine or monitor you depend on is losing its validity — before the miss shows up as a loss.

open weights on Hugging Face · open benchmark protocol

REFERENCE ENVELOPEMODEL VALIDITY DEGRADESCREDIBLE TAILS

purpose

Why a model of surprise, and why now

Every consequential decision today is downstream of a model — a flood forecast, a risk engine, a grid schedule, an LLM policy. Those models are trained on the world as it was, and they fail silently when it moves. Zeno Divergent exists to make that moment predictable and legible before the loss lands.

the problem

Failure is discovered downstream

A reference model keeps emitting confident numbers well past the point where its assumptions hold. Nothing in its own output says 'I no longer apply here'. The break is found in the post-mortem, in the incident review, in the loss.

the object

Model validity as a forecastable quantity

Zeno Divergent takes the reference models themselves as inputs and learns a latent surprise state over the gap between what they expect and what arrives. From that one vector it reads the kind of divergence, what could not be observed, how long each reference stays valid, and whether to abstain.

the consequence

A decision you can defend

You get a calibrated probability with a named type and horizon, and an action — act, watch, ask for more information, or stop trusting this model. Abstention is a first-class answer, which is what separates assurance from another anomaly score.

pipeline

From a public feed to a decision you can defend

Four stages, each one inspectable in the workspace. Nothing is imputed to fill a gap: an unobserved channel stays unobserved, and the model is allowed to abstain.

01 · Ingest

Public feeds, persisted

USGS event and streamflow services, Coinbase bars, NOAA SWPC indices, FRED series and GHCN-Daily stations land in the production corpus with their own units and cadence.

corpus_series · corpus_observations

02 · Fuse

Onto one cadence grid

Entities are aligned to a shared 12-step window and joined with the reference predictors they are meant to obey. Channels a feed does not carry are marked unobserved, never imputed.

72 features · cadence, magnitude, evidence, residual

03 · Infer

Calibrated, per domain

A per-domain adapter feeds a shared dilated causal trunk. Outputs are Platt-scaled, so a stated probability means what it says, and abstention is an available answer.

v0.7.0 · 1.87M parameters

04 · Explain

Typed, not a single score

A SurpriseReport names the kind of divergence, what could not be seen, validity per horizon, and the next action — including request-information.

surprise latent · reliability · risk–coverage

what you get

A SurpriseReport, not a single number

Every scored entity comes back in the same typed contract, so the answer can be acted on and audited instead of interpreted.

divergence
which kind of break it is — magnitude, timing, structure, mechanism or coverage.
validity horizon
how long the reference you rely on is still expected to hold.
blindness
what could not be observed, kept separate from what looks calm.
confidence
a calibrated probability, or an explicit abstain when it cannot tell.
next action
act, watch, request information, or stop trusting this model.

example

entity        gauge-11447650
divergence    structure · cadence break
validity      reference holds ≤ 6h
blindness     2 of 4 channels unobserved
confidence    0.81  (calibrated)
action        request information

structure

The hierarchy

The model is the intellectual centre. Everything above it is subordinate to it.

01

Surprise Intelligence

the research direction

Open-world model assurance: making the edge of a model's validity a forecastable quantity rather than a post-mortem finding.

02

Zeno Divergent

the learned model

Models are first-class inputs: it attends over expectations, not features. It predicts future model validity, and can abstain instead of scoring.

03

SurpriseBench

the benchmark

Held-out episodes, locked thresholds, degenerate controls, label-shuffled copies and leave-one-mechanism-out transfer. Built to return a negative verdict — and it has.

04

SurpriseReport

the interface

One typed contract: kind of divergence, blindness, validity per horizon, a surprise-potential ranking, and an action — including request-information.

kernel

Six primitives, one kernel

Every domain pack shares the same reasoning machinery: six primitives that turn a set of reference expectations into a typed account of where they are breaking.

k01

Reference graphs

Several independent expectations per question, with disagreement kept explicit instead of collapsed into consensus.

k02

Typed surprise

Surprise is classified — magnitude, timing, structure, mechanism, coverage — never reduced to one anomaly score.

k03

Unknown / blindness

A first-class channel for what could not be observed, separating low risk from low visibility.

k04

Credible tail search

Bounded search over admissible futures that the current references do not carry.

k05

Model validity

Admission control decides whether a model may still speak: ignore, remember, open a regime, ask, rebuild.

k06

Surprise memory

Events are stored with magnitude and novelty apart, so recurrence is never mistaken for novelty.

domains

Three verticals, one cross-cutting capability

Earth Observation, Finance & Web3 and Critical Infrastructure & Defence are the worlds the model learns in, each keeping its own units, references and unit of consequence. AI Meta-Evaluation & Safety is not a fourth industry: it is the same machinery turned on models themselves.

nothing here is live in production — maturity is stated per sub-area below.

vertical product families

near = implemented · build = needs a tenant and labels · research = weak labels or physics

Three industries where the consequence of a stale model is measured in a real unit — hectares, money, or service-hours. Each family owns its references, its constraints and its own unit; they are never blended into one score.

Earth Observation

unit · physical magnitude × exposure

Tells you when the physical forecasts you rely on — fire, flood, heat, drought — have drifted outside the conditions they were fitted on, and where to look next.

  • wildfire & fuel statebuild
  • extreme weatherbuild
  • flood & hydrologybuild
  • drought & vegetationbuild
  • ocean & sea-surface temperatureresearch
  • land-use & urban changeresearch
  • sensor tasking by expected surprise reductionresearch

refs: physics model · ensemble · persistence · satellite retrieval

Finance & Web3

unit · money

Flags the moment pricing, credit and execution models stop describing the market they are trading in — before the loss, while there is still time to widen, hedge or stand down.

  • HFT microstructure fragilitybuild
  • onchain protocol & liquidity riskbuild
  • uncollateralized credit (lender-side)build
  • treasury & liquiditybuild
  • desk / market risk overlaybuild
  • payroll & payout railsnear
  • mempool & flash-loan interceptionresearch
  • execution intent protectionresearch
  • insurance & catastrophe pricingresearch

refs: originator model · peer cohort · cash-vs-tape · venue order book · latency / fill model · yesterday's model

Critical Infrastructure & Defence

unit · safety & continuity

Watches the detectors, plant models and fused pictures that keep essential systems running, and separates a quiet estate from an unobserved one.

  • pre-zero-day threat huntingbuild
  • identity-model validity / UEBAbuild
  • OT & industrial controlbuild
  • grid & energy cascadesbuild
  • detection-model monitoringnear
  • water & municipal systemsresearch
  • defence sensor-fusion picture assuranceresearch
  • autonomous swarm coordinationresearch

refs: role cohort · change calendar · plant grammar · endpoint telemetry health · sensor fusion

cross-cutting capability

One capability that sits across every vertical instead of beside them: judging models themselves. Its unit is trust, not consequence, and its subject can be any predictor — ours or someone else's.

AI Meta-Evaluation & Safety

unit · model trust

Not an industry — a capability that applies to all of them. Point it at any model, in any of the verticals above or outside them, and it grades how far that model can still be trusted and when it should abstain.

  • third-party fragility audit (overconfidence index)near
  • detection / forecast model passportsnear
  • OOD guardrails on production telemetrybuild
  • continuous production monitoring loopbuild
  • agent behaviour & tool-use driftresearch

refs: subject model logits · held-out episodes · production telemetry · degenerate controls

SurpriseBench — the shared protocol

All three verticals and the cross-cutting pack are graded under the same frozen protocol, and it is the part that exists today: grading other models' stale worlds. Any external predictor can be run as the subject — typed miss, lead time at a fixed false-alert rate, calibration and degeneracy guards, with mandatory constant, random and shuffled-label controls. It is domain-independent and it publishes its losses.

open SurpriseBench →

coupling edges — unit conversion, not a universal risk score

EO → FI

fire / drought Ψ → catastrophe price and borrower cash-flow

km² → USD

EO → CI

heat + drought → grid reserve margin and trip potential

°C → service-hours

CI → FI

outage → payroll, working capital and insured loss

service-hours → USD

CI → AI

detector degradation → downstream model trust and abstention rate

coverage → trust

proof — the falsifiable claim

A learned representation of model-relative surprise contains information about future regime change and model failure that predictive uncertainty alone does not capture.

Stated so it can fail. The surprise index and eight conventional detectors are scored on identical inputs and labels, with thresholds locked before the run; if a baseline matches it, the harness says the claim is unsupported.

latest release
v0.7.0 · v0.7.0
promoted domains
hydrology, climate
promotion gate
promoted
parameters
≈1.87M
in-browser scoring
v0.7.0 + v0.5.0 ONNX heads
not promoted
seismic · markets · spaceweather · macro · wildfire · chaos

compounding

How the understanding compounds

One loop, repeated per release: add real corpora, declare the gate before training, promote only what clears it, freeze everything so the next comparison is honest.

  1. 01 · widen

    Each release adds real public corpora that share no physics. A latent that survives seismicity, streamflow, markets, space weather and climate at once is encoding departure itself, not one domain's habits.

  2. 02 · gate

    Thresholds, degenerate controls and leave-one-domain-out transfer are fixed before the run. v0.7.0 promoted 2 of 8 domains; v0.4.0 promoted none and is published as a failure.

  3. 03 · learn from failures

    v0.7.0 promotes fewer domains than v0.6.0 because the gate now demands a win over a plain logistic regression on the same features. hydrology and climate clear it; the rest ship as candidates with their blocking reason attached.

  4. 04 · freeze

    Every release is an immutable Hugging Face revision carrying checksums, calibration parameters and the full gate report — so improvement is checkable by anyone, release over release.

limits

What it cannot do yet

This is research in the open, not a finished product. The current release clears its gate on hydrology, climate and does not on seismic, markets, spaceweather, macro, wildfire, chaos; no vertical is running in production, and the claim above is not yet demonstrated beyond those two domains — everywhere else a plain logistic baseline is still the honest choice. Every failure is published alongside the wins rather than removed from the table.

read the full evidence page, failures included →

get started

Point it at a stream you care about

Upload a CSV or connect a feed, and the workspace returns a scored SurpriseReport for every entity in it.