What changed in the specialists, and what it cost
Between the public beta and the current milestone, the cluster went through a restructuring we had not attempted before. The count went up. The authority went down. Both were deliberate, and the second one mattered more.
Ch. 01 — Architecture
The core change: authority tiers
Previously, any specialist that scored a situation was implicitly allowed to influence what the system did about it. That was tolerable when there were a handful of them. It stopped being tolerable as the cluster grew, because a specialist with flattering development metrics and weak transfer to unseen conditions could quietly distort the outcome without anyone being able to name it as the cause.
The current build introduces an explicit four-tier taxonomy of authority. Every specialist is classified, and the tier — not the metric, not the enthusiasm of whoever built it — determines whether it can influence a live action, block one, or only be recorded for observation:
Nothing changes an action without a declared contract naming the specialist, the consumer it may affect, and the specific change it is permitted to make. That contract is versioned in code and enforced by the consumer and the shared risk layer — not by convention, and not by whoever is paying attention that day.
A specialist with good numbers and no declared authority is a research result. Treating it as anything more is how quiet distortion gets in.
On why the tiers existThe funnel every candidate goes through
Through this cycle we put a large number of candidate specialists through the same batch-graded process. The common path:
- Cheap screen first. A low-cost pass on a fixed snapshot of inputs. Weak ideas are rejected before anything expensive is spent on them.
- Paired comparison against whatever currently holds that job. A replacement has to add something useful and non-duplicative. Being merely as good is not a reason to change anything.
- Replay through the real consumer — the actual strategy and action path that would use it, including applicable costs and the abstention behaviour it would inherit. A result that only exists in isolation does not count.
- Stress across market characters. Does it hold up in rising, falling, calm and violent conditions? A specialist that is exceptional in one and neutral-to-harmful in the rest gets scoped so it only fires in the one where it works.
What a candidate ends up allowed to do depends entirely on the evidence it earned: scoped acting authority, observation, context, or research only. Rejected work is archived off the active path with its failure reason attached, which is what stops a retired claim from quietly reappearing later.
Rehabilitation, done deliberately
A subtler change: failed ideas are now formally rehabilitated rather than either abandoned or quietly re-run until they pass. The method is fixed. Identify the specific failure mode — missing inputs, overfitting to one market character, information leakage, or a mismatch with the consumer that would use it. Rebuild only when there is a specific corrective hypothesis to test. Then re-run the entire funnel from the start. A rebuilt candidate inherits none of the previous version's authority.
The retirement discipline
This cycle also retired specialists. Specifically: ones whose advantage on unseen data collapsed once a proper leakage audit was finished; ones whose development data turned out to be contaminated by adjustment errors in the underlying price history; and ones whose apparent signal existed only in synthetic labels and evaporated when replayed against real chains. Every retirement is dated and reasoned in the internal registry.
What ships is the survivor set, not a museum of everything that once posted a promising number. Keeping the rejected set, with reasons, is what prevents old claims from creeping back onto the active path.
The wiring around them
The surrounding safety system was connected in the same cycle: checks for company earnings and scheduled events, handling for disagreement in options activity, market-character context, verification that each specialist actually has the inputs it needs, checks on what the connected brokerage can actually do, activity and gap prefilters, and symbol exclusions. Packaging now enforces bundled assets and per-user runtime state, so a development fallback cannot impersonate a clean release.
We publish no performance figures. This page describes what changed structurally and why, never what any component scored. Any return, win-rate or accuracy number attributed to Qovaryx did not originate from us.
The best architecture is not the one with the most specialists. It is the one with the most specialists that earned their place — and the most retirements whose reasons you can still point to.
On what a good cluster looks likeCompanion piece: the data underneath all of this, and how it is graded.