Calibrated conviction beats raw confidence
“The system is highly confident.” Fine — but out of all the times it said that, how often was it actually right? That second question is the whole game. Without it the first number is decoration. With it, you have something you can size a position against.
Ch. 01 — Doctrine
What calibration means
A system is calibrated when its stated confidence matches its observed hit rate. If it expresses a given level of confidence across a large number of cases, roughly that proportion should resolve the way it predicted. If it consistently states more confidence than its record supports, it is overconfident — and any position sizing tied to that stated confidence is broken at the root.
This matters because raw model scores are generally not calibrated out of the box. They are produced by systems optimised to pick the most likely answer, not to be honest about how likely that answer is. Treating a raw score as if it were a probability is a quiet, consequential mistake.
How we handle it
We do not surface raw model output. Every confidence the system shows you has passed through a correction layer fitted on held-out outcomes. We are not going to detail the shape of that layer here — it would hand a competitor the interesting part without the unglamorous work that makes it hold. The part we will commit to in public is the discipline:
- Calibration is measured on data the system was not developed against, never on the data it learned from.
- We track the gap between stated confidence and observed outcome across the whole range, not just at the top end where the flattering cases live.
- When that gap is materially wide anywhere in the range, the component does not ship. It goes back, or it is retired.
An uncalibrated confidence number is not a weak signal. It is a broken input to every sizing decision downstream of it.
On why calibration is upstream of everythingWhy this connects directly to position size
Honest confidence is what makes sizing a solvable problem rather than a guess. If a stated level of confidence genuinely corresponds to how often that level is right, then the amount of capital that level justifies has a defensible answer. If it does not, then any sizing rule built on top of it is arithmetic performed on a fiction.
Qovaryx scales exposure with calibrated confidence, and it scales it conservatively — deliberately below what the arithmetic alone would permit, because the arithmetic assumes the calibration is perfect and it never is. Below a floor level of confidence, the system does not size down. It declines entirely.
That floor is not a round number picked for how it reads. It is set where the calibration record stops being reliable enough to size against — the point below which the expected value of acting, after the cost of transacting, stops being defensible.
We publish no performance figures, and no parameter values. This page describes how sizing is derived, not the numbers it produces. Any return, win-rate or accuracy figure attributed to Qovaryx did not originate from us.
If a tool tells you how confident it is but not how often that confidence was right, your position size is being chosen by a marketing department.
On confidence as a sales figure