manifesto

A model should be able to tell you when it has stopped being right.

Almost every model in production answers one question: what happens next. Almost none answers the question an operator actually asks at 3 a.m. — is the thing I am relying on still valid, and if not, what should I look at instead. Zeno Divergent exists because that second question deserves a learned model of its own.

01

The interesting object is the gap, not the forecast.

Every deployed model already produces an expectation. What almost nothing produces is a learned account of how reality has been departing from that expectation. That residual structure is a first-class object with its own dynamics, and it deserves its own model.

02

Uncertainty is not surprise.

A wide interval means the model knows it is unsure. Surprise means the model was sure and wrong, in a way that repeats. Confusing the two is why teams get calibrated confidence right up to the moment the regime changes.

03

One anomaly score destroys the only information you needed.

Magnitude, timing, structure, mechanism and coverage fail differently and demand different responses. Collapsing them into a single scalar produces a number that fires, and an operator with nothing to do about it.

04

Not-observed is not the same as not-happening.

Blindness gets a dedicated channel. A quiet sensor, a stale feed and a genuinely calm system produce the same low reading and mean opposite things; a system that cannot tell them apart is confidently reporting on a region it cannot see.

05

Models must be able to lose the right to speak.

Admission control is explicit: ignore, remember, open a regime, request information, rebuild. A reference that has been wrong in a structured way for weeks should not keep contributing to a consensus as if nothing happened.

06

Never blend models into a single risk number.

Disagreement between references is signal, not noise to be averaged away. Coupling across domains converts one domain's output into the receiving domain's own unit — there is no universal risk score, because there is no universal unit.

07

State the claim so it can fail.

The claim is a single sentence, scored against eight conventional detectors, two degenerate controls and a label-shuffled copy of itself, on identical inputs. If a baseline matches it, the harness prints not supported. Everything published here is synthetic until it is reproduced on real historical events, and the site says so on every page that shows a number.

the vocabulary this rests on

Every position above depends on four ideas: , , and the they are decoded from. Hover any of them, or read the full definitions below.

definitions

Glossary

The terms used across the manifesto and how-it-works pages, defined once and cross-referenced.

reference graph
The set of existing models that claimed to cover a question, kept as nodes with declared coverage and validity rather than averaged into one consensus. Every forecaster, simulator or risk engine that made a prediction for the same step becomes a node carrying its own coverage, assumptions and validity conditions. Edges record where two references overlap. Zeno reads the pattern of agreement and disagreement across that graph instead of collapsing it into an ensemble mean, because the shape of who disagrees is the signal.

see also admission control, model-relative surprise

model-relative surprise
Departure measured against what the reference models expected, not against a historical average. The learned object is the difference between reality and the expectations of the models currently in production. The same observation can be unremarkable for one reference set and a severe departure for another; surprise is therefore always relative to a stated set of expectations.

see also surprise latent

surprise latent
A ten-dimensional learned state that encodes the structure of departure — magnitude, persistence, structure, mechanism and coverage. Written z^S_t, the latent is trained under several objectives at once: reconstructing the departure, contrasting regime change, predicting reference validity and an information bottleneck that stops it from memorising the outcome. It is the piece the falsifiable claim is about — that it carries information about future regime change which predictive uncertainty alone does not.

see also typed surprise channels, predictive uncertainty

typed surprise channels
The latent is decoded into separate channels — disagreement, structural departure and blindness — that are never blended into one score. A single risk number cannot tell you whether to act or to look. Typed surprise splits the reading into channels so the response can differ: reference disagreement calls for arbitration between models, structural departure calls for a change of policy, and blindness calls for observation. Each channel carries its own magnitude, timing and consequence estimate in the receiving domain's unit.

see also blindness (unknown channel), information request

blindness (unknown channel)
How much of a low reading is genuine calm and how much is simply not observed. Missing data is treated as a first-class input rather than imputed away. The blindness channel reports the share of the current window that no reference could see — a cloud-covered tile, a halted instrument, a venue that stopped reporting. A quiet surprise score with high blindness is not reassurance; it is an instruction to restore visibility before trusting the reading.

see also information request, typed surprise channels

admission control
A per-reference decision about whether a model still has the right to speak about the current situation. Each reference declares the conditions under which it is valid. Admission control tracks whether the observed regime still sits inside those conditions and marks a reference as admitted, provisional or withdrawn. Withdrawn references keep producing numbers — they just stop being counted as expectation, and the fact of withdrawal is itself reported as evidence that the world moved outside the modelled envelope.

see also model validity, reference graph

model validity
Whether the reference set as a whole is still operating inside its assumptions. An aggregate state — valid, strained or invalid — derived from admission control across references, calibration drift and coverage. It answers the operator's question directly: is the thing I am relying on still right, and if not, since when.

see also admission control

credible tail
An unexpected but constraint-consistent future found by search, not by random sampling. Credible tail search looks for trajectories that are far from reference expectation yet still satisfy domain constraints and the current surprise state. Trajectories that violate a hard constraint are discarded rather than down-weighted, so what survives is a small set of counter-expectations that are hard to dismiss on physical or institutional grounds.

see also typed surprise channels

information request
The model's statement of what should be observed next to resolve the current ambiguity. Rather than only reporting a score, the model names the observation that would most reduce its own uncertainty about whether a departure is real — a re-tasked sensor, a second venue, an instrument check. It is the actionable half of the blindness channel.

see also blindness (unknown channel)

systemic coupling
Explicit edges that carry surprise from one domain into another in the receiving domain's own unit. A departure in Earth Observation may imply exposure in Financial Services or load risk in Critical Infrastructure. Coupling edges declare the mechanism, and propagated surprise is converted into the receiving domain's native unit and attributed back to its source and path. There is no universal risk score.
predictive uncertainty
A model's own stated confidence — the baseline the surprise latent must beat to be worth anything. Variance, entropy or conformal width reported by an existing model. It is the strongest conventional rival: if predictive uncertainty alone were enough to anticipate regime change and model failure, no separate surprise representation would be needed. SurpriseBench runs it as a baseline on identical inputs.

see also surprise latent, lead@FAR

lead@FAR
Average warning time before a break, measured only while the false-alert rate is held at a fixed budget. Lead time is trivial to inflate by alarming constantly, so it is only reported with the alarm threshold locked to a false-positive rate on clean steps (10% in the published harness). A method that cannot hold the budget reports no lead at all.

see also predictive uncertainty

what this is not

  • Not a replacement for your forecasters, simulators or risk engines. It reads them.
  • Not an early-warning oracle. It reports where warning is possible and where visibility is missing.
  • Not validated on real-world data. The current checkpoint is trained on a synthetic corpus and every published number is a property of that harness.