This is a position paper from the POE 1 research programme. It sets out how we think about the hardest problem in applied forecasting — keeping stated confidence honest when the world stops resembling its own history — and how that thinking is embedded in the engine's architecture. Where questions remain empirically open, we say so.
1. The problem, stated plainly
A forecasting system is calibrated when its stated confidence matches its realised accuracy. If a system says "70%" across a thousand predictions, roughly seven hundred of them should come true. This is the minimal, non-negotiable condition for a probability to mean anything at all. A probability that does not survive this test is not information; it is decoration.
Calibration is measurable. The Brier score penalises the squared distance between stated probability and realised outcome. Expected Calibration Error (ECE) buckets predictions by stated confidence and measures the gap between each bucket's average confidence and its average accuracy. These are old, unglamorous, well-understood instruments, and they have one decisive property: they cannot be argued with. A system either is calibrated on a body of resolved predictions or it is not.
So far, so tractable. The difficulty — the one this paper is about — is that calibration is always calibration with respect to a history. Every calibrated system earned its calibration on some distribution of past cases. The implicit promise it makes is that the future will be drawn, more or less, from the same distribution. And the most consequential moments in any domain — the ones where a forecast is worth the most — are precisely the moments when that promise breaks. Regime change is the name for that break: the point at which the generating process behind outcomes shifts, and the past stops being a reliable sample of the future.
The cruel structure of the problem is this: a system's confidence is most likely to be wrong exactly when its confidence matters most. A model calibrated on a stable regime will sail into a regime break carrying confidence it earned somewhere else. Its stated 80% is an artefact of a world that no longer exists. And nothing inside a conventional model flags this, because the model's confidence machinery has no concept of the conditions under which it was earned.
2. Why regime change defeats conventional calibration
It is worth being precise about the failure mechanism, because the obvious fixes do not work.
Recalibration lags by construction. The standard remedy for drifting calibration is to recalibrate on recent outcomes. But recalibration requires resolved outcomes from the new regime, and regime breaks are most dangerous in their opening phase, before any outcomes from the new regime have resolved. A system recalibrating on a trailing window is always calibrated to the regime that is ending. The faster the break, the worse the lag.
Ensembles average the problem; they do not detect it. Combining models trained on different histories widens the spread of opinion, which helps, but an ensemble of systems that all assume distributional stability fails together when stability fails. Disagreement among ensemble members is a weak and noisy regime signal — it can mean a break, or it can mean ordinary epistemic noise — and the ensemble itself has no way to tell.
Fatter uncertainty bands are a tax, not a solution. One can respond to regime risk by permanently widening every interval. This protects against the break at the cost of degrading every forecast issued during stability — which is most of them. A system that is vague always, in case the world changes sometime, has solved overconfidence by abandoning informativeness. Calibration was supposed to make probabilities meaningful; uniform vagueness makes them meaningless from the other direction.
The conclusion we draw is structural: regime robustness cannot be achieved inside the confidence number. It has to be achieved by an architecture that treats "what regime am I in, and how sure am I about that?" as a first-class measured quantity — separate from, and prior to, the forecast itself.
3. Decomposing confidence: the nine dimensions
The first architectural commitment in POE 1 follows directly from the analysis above: a single scalar confidence is not enough, because a scalar cannot distinguish why it is what it is. Two predictions can both deserve "60%" for completely different reasons — one because the evidence is rich but genuinely conflicting, the other because the evidence is thin. Those two 60%s should behave completely differently under regime stress, and a system that collapses them into one number has thrown away exactly the information regime detection needs.
POE 1 therefore constructs confidence from nine measured dimensions, each scored on every prediction:
Novelty — how far the present case sits from the body of cases the system has resolved before. High novelty is the leading edge of regime risk: the system is being asked about territory its history does not cover.
Contradiction — the degree of genuine conflict within the evidence pool. Contradiction is informative: regime breaks frequently announce themselves as a rise in credible sources disagreeing about what was previously settled.
Instability — the volatility of the relevant signals over recent windows. Stable regimes produce stable signal statistics; rising instability is a forward indicator that the generating process is moving.
Support — the depth and directness of evidence bearing on this specific question, as opposed to evidence that is merely adjacent.
Missing data — an explicit accounting of what the system knows it does not have. Absence of evidence is measured, not ignored.
Regime mismatch — the dimension aimed squarely at this paper's problem: a direct assessment of whether the conditions under which the relevant historical patterns were learned still hold in the present case.
Sparsity — how thin the case base is in this region of the problem space, independent of how clean it is.
Trust — the assessed reliability of the sources contributing evidence, screened and scored rather than assumed.
Historical reliability — the system's own track record on comparable predictions, measured with Brier scores and ECE. This is the dimension that closes the loop: it makes the system's past calibration an input to its present confidence.
The decomposition matters because it changes what a confidence figure is. In POE 1, a confidence is not a feeling rendered as a percentage; it is the governed aggregation of nine audited measurements, each of which is stored, inspectable, and attached to the prediction's lineage. And — critically for the regime problem — the dimensions that lead regime breaks (novelty, contradiction, instability, regime mismatch) are visible individually, before they have dragged the aggregate down. The system can see the storm forming in the instruments while the headline number still looks calm.
4. Evidence hygiene as a calibration problem
There is a second, less discussed route by which regime change corrupts confidence: it pollutes the evidence pool. When a domain shifts, the retrieval problem shifts with it. Sources that were authoritative become stale. Commentary written under the old regime continues to circulate as if current. Material from adjacent domains and regions — irrelevant but superficially similar — floods into any retrieval process tuned on the old regime's structure. A forecasting system that ingests this pool uncritically is not just working with weaker evidence; it is working with evidence that systematically points backwards, towards the regime that ended.
POE 1 treats this as a screening problem with teeth. Every evidence pool passes a contamination guard that tests six pollution categories: wrong domain, wrong sub-domain, wrong region, wrong source type, generic filler, and broken pages. The guard's verdicts are graduated and consequential. A clean pool passes. A contaminated-but-usable pool forces the prediction into scenario mode: the system is downgraded from issuing a single calibrated estimate to describing structured alternative futures, because the evidence will not support more. A dominantly contaminated pool triggers the system's most important behaviour: it fails closed. No prediction is issued.
Fail-closed deserves a moment of emphasis, because it is the behaviour most foreign to the current culture of AI systems, which fail open by default — when evidence is bad, they answer anyway, fluently. Failing closed is not an error state. It is a calibration statement: the honest probability assignment here is no probability at all. A system that cannot abstain cannot be calibrated under regime stress, because regime stress is precisely the condition under which abstention is sometimes the only honest output.
5. Abstention and downgrade as calibrated behaviours
This reframing — abstention as a calibration act — is, we think, the conceptually important move, so it is worth making carefully.
Conventional calibration thinking evaluates a system on the predictions it makes. But a system that chooses when to predict is making a second, prior decision, and that decision has its own calibration question: does the system abstain in the cases where its predictions would have been bad, and predict in the cases where they would have been good? A system with excellent Brier scores achieved by answering everything is strictly worse, under regime risk, than a system with the same Brier scores achieved while correctly routing its hardest cases into scenario mode or abstention — because the second system's silence is informative. Its refusal to quote a number is a forecast: a forecast that the question, right now, does not admit a justified probability.
POE 1's graduated output ladder — calibrated estimate, scenario mode, abstention — is therefore not a safety feature bolted onto a forecasting system. It is the forecasting system, extended to be honest about its own boundary. The nine dimensions decide where on the ladder a given case lands; the contamination guard can force a case down the ladder regardless of how the dimensions score; and the platform's launch gate sits above all of it, issuing green, amber, red, or do-not-claim verdicts with an asymmetry that encodes the philosophy: polluted evidence can always force do-not-claim, and missing governance can never produce green.
6. Learning across regimes: Case → Cluster → Law
A final architectural question remains: what should a system retain across a regime break? Discard everything and you forfeit the genuine regularities that survive regimes. Retain everything and you carry the dead regime's patterns into the new world with unearned authority.
POE 1's knowledge lifecycle is built to make this trade-off explicit rather than implicit. Resolved predictions accumulate as cases — the atomic, regime-stamped record of what was forecast, on what evidence, with what confidence, and what happened. Recurring structure across cases forms clusters. And only clusters that demonstrate promotion readiness — recurrence robust enough, across conditions varied enough — are elevated into laws: explicit, versioned statements of pattern that future predictions may draw on.
Two properties of this lifecycle matter for the regime problem. First, promotion is conservative and tracked: a pattern that has only ever been observed inside one regime carries that fact in its record, and the regime-mismatch dimension can discount it the moment present conditions diverge from the conditions of its formation. Second, laws are versioned, never silently mutated: when a regime break invalidates a law, the invalidation is itself an event in the record — auditable, dated, and informative. The knowledge base does not merely survive regime change; it documents it. Over time, the record of which laws broke under which breaks becomes some of the most valuable evidence the system holds, because the next break will rhyme with the last ones.
7. What we claim, and what remains open
It is important, in a paper like this, to separate the architectural argument from the empirical one.
The architectural argument we are prepared to defend in full: scalar confidence cannot be regime-robust; confidence must be decomposed into measured dimensions that include regime mismatch directly; evidence hygiene is a calibration problem and must be enforced with graduated, fail-closed screening; abstention and downgrade are calibrated behaviours and must be first-class outputs; and cross-regime learning requires a governed, versioned promotion lifecycle rather than undifferentiated accumulation. Each of these is implemented in POE 1 as described, and each is inspectable in any prediction's lineage.
The empirical questions are open, and we hold them as obligations rather than embarrassments. How early do the leading dimensions — novelty, contradiction, instability, regime mismatch — actually flag historical regime breaks, and at what false-positive cost? What is the realised calibration of scenario-mode outputs, which require their own evaluation methodology? Where exactly should the abstention threshold sit in each domain, given that abstention has a price too? These questions can only be answered the slow way: by running the engine against live, resolving outcomes and publishing the Brier scores and ECE that result. That programme is under way, beginning in domains where outcomes resolve in hours rather than years, because fast resolution is fast evidence — about the engine itself.
That, ultimately, is the discipline regime change imposes. It cannot be solved in advance, only instrumented for. A forecasting system's claim to be trusted through a regime break rests on three things: that it measured the break's approach, that it downgraded itself honestly as the evidence thinned, and that its record afterwards shows the downgrades were right. We have built the instruments. The record is being written now.
IP Factory HQ builds POE 1, the Predictive Outcome Engine — governed intelligence systems that identify patterns, quantify probabilities, and inform decisions before outcomes occur.