The data moat
People ask how a small operation competes with frontier labs on model quality. The honest answer is that model architecture is not where this fight is won. It is won on data nobody else bothered to collect, curate and grade — and on the discipline applied to it afterwards.
Ch. 01 — Architecture
What sits underneath the system
Different parts of the system draw on different task-appropriate slices of the material assembled for Qovaryx. In broad strokes:
What we will not publish: exact snapshot dates, the specific list of symbols, the functions that produce the labels, or the code that validates the bridging between real and synthetic. Those are the moat.
The grading discipline
Volume is not a moat on its own. Data is only a moat if it is used properly, and two rules govern every development run:
- Leakage scrubbing. Every dataset is audited for information that would let a component see the answer before it has earned it — before a single development run starts. When we discovered that a portion of earlier synthetic labels had been contaminated by adjustment errors in the underlying price history, we retracted everything that depended on them and re-labelled from scratch. What survived the re-audit is what runs today.
- Held-out means held out. Slices reserved by time, by symbol and by market character are set aside for final evaluation and are not touched before it. Records document which split was used and which consumer it was evaluated through. A promising development metric is never accepted as a substitute for the held-out result.
A curated window is not a shortcut to the same answer. It is a different answer, to a question nobody asked.
On why the bad years stay inWhy it compounds
Every week spent adding labelled cases, unfamiliar market conditions and new failure modes to this material is a week a competitor fine-tuning a general-purpose model on public data cannot recover. Public options data ends where somebody put it behind a paywall. Ours begins with what is public, adds what we have paid for, and adds what we have labelled ourselves — and the last of those is not for sale at any price.
This is also why we take the no-recipe-leakage policy seriously in these write-ups. We publish the framing: what we tried, what we retired, and why. We do not publish the labelling functions, the specific evaluation slices, or the exact inputs that survived the process. That is the recipe, and the recipe is the moat.
We publish no performance figures. This page describes what the source material is and how it is graded, never what it produced. Any return, win-rate or accuracy number attributed to Qovaryx did not originate from us.
Model architecture is a photocopiable commodity. A well-graded record of things that actually happened is not.
On where the durable advantage sitsCompanion piece: how the specialist cluster evolved on top of this foundation.