AI Mini-Series Part X: Taking the Win Apart: What the Control Group Found in Our Best Result

Part 10 of a series on building and testing an AI layer for our systematic framework. This one is about the first two things our control taught us once it was running, and about the tension underneath both: our effort to make a reasoning layer measured and repeatable, and what that effort costs the reasoning.
When we built a control group to measure the Council (the ninth piece), we ended on a rule: findings are cheap, and what matters is what we allow them to change. The first test of that rule turned out to be our own best result. Our map of where the machine adds value (the eighth piece) pointed at one book in particular, the one we run to trade leadership rotating from one part of the market to another, and it was the one book where our Council had come out ahead of its twin. The Heretic we built to argue against ourselves (the fifth piece) taught us that the scorecard is the last thing we doubt, after the read and after the method. This piece is what that looked like when the scorecard happened to be flattering us.
A word about where these findings come from, because it shapes how to read them. The first piece said the burden runs one way: our process has to make the AI systematic, and systematic means measured and repeatable. We have spent most of the past two months making the reasoning layer exactly that, pinning what it sees, scoring what it says, and testing what it does against a control. This is a learning process, and nearly everything in this piece and the next is what our measuring turned up. Some of it is about the model. More of it is about a tension we did not fully appreciate when the series began: every measurement we add is structure, structure is deterministic, and it is possible to make a reasoning layer so measured that there is less and less reasoning left inside it. Finding the balance is our work.
A shape is not yet a reason
Our rotation book runs on a ranking. Each night the Council orders the sectors of the market from most favored to least, and our code turns that ordering into positions, sized and neutralized with no further say from the model. The ranking is the model’s entire contribution, so when the book came out ahead of its twin, we read it the natural way: the ranking was carrying something the twin’s did not. That was the shape our map predicts, and it was a fair first reading.
A working journal has to ask the next question, which is what the ranking was made of. So we audited the ranking itself, across every graded sector call our Council had made, several hundred of them. We found that the rankings tracked each sector’s trailing few weeks of price closely, whichever Force the model said it was reading. Between half and two thirds of the ranking, depending on how we measured it, was trailing momentum, and once we set that momentum aside, the remainder carried no forward signal we could detect.
The reason is the lesson, and every builder of these systems will meet it. We had been sending all five Forces to our sector agent, but we sent them as descriptions of the market as a whole: national macro, index technicals, aggregate fundamentals, market-wide sentiment. The only inputs we gave it that told one sector from another were relative-strength figures, every one a transformation of price. When we asked it to rank technology against healthcare, it could only rank on what distinguished them, and what distinguished them was price. It did not go wrong. It went the only way our inputs allowed. A reasoning layer can only reason about differences it has been shown, and the first question we now ask of any ranking it produces is what, exactly, told the candidates apart.
There is a second, quieter lesson inside the first. We noticed the model had attributed many of those calls to the macro Force, and we had given it no sector-level macro that could have driven them. A model asked to explain its call will supply an explanation whether or not the cause was in front of it. That is not dishonesty and it is not a defect. It is what fluency does when the form we built demands a reason. We said in the fourth piece that the wrongness gets quieter and better dressed as you climb; this is the best-dressed instance we have met, and our control is what found it.
Our fix was not a better prompt, because a prompt cannot conjure information that was never in the room. We gave the sector agent, for each sector, the parts of the other Forces that actually differ between sectors: how each one is positioned, what its earnings are doing, and where its money is flowing. For the first time our ranking has inputs that are not price, and so for the first time it can disagree with price. Whether it does, and whether that helps, is a question our record can now answer. Because it is a question and not an answer, we wrote the test down before the first result existed: what counts as evidence, what counts as nothing, and one rule above the rest. A ranking that beats the market by riding momentum in a trending tape has demonstrated momentum, not our Council, and we declared that outcome in advance to be nothing. A trending market would hand the old ranking a flattering score for free. We wrote the test so that flattery cannot be mistaken for skill.
What the arithmetic protects us from
Before we rebuilt anything, we did the arithmetic the first piece only gestured at, and it protects us from a mistake in either direction.
The margin by which our rotation book had beaten its twin was small and well inside noise. When we split the record at the boundary between replayed history and the live period, the sign of the gap reversed on most of the books we run this way. And at the daily noise we observe, our current record could only detect an edge of a third to a half of the book per year; a real edge of the size a serious strategy might have would take four to five years of this record to establish. The honest word for our scorecard is unfalsified, and it is a useful word, because it refuses both temptations at once: to declare our apparatus proven on a good month, and to declare it a failure on a bad one.
That arithmetic is also why we will not be quoting returns as evidence in this series for a long time. Returns arrive once per book per day, and years of them are needed. The mechanism we can measure every day, and months of it will do: whether our ranking has stopped being a shadow of price is a question the record answers an order of magnitude sooner than whether the book made money. So we measure the mechanism first, and the returns wait.
Our map from the eighth piece stands. The record that seemed to confirm it was a first look rather than a verdict, and what we did was hold our best result to the same bar we would put in front of a stranger’s strategy, which we promised in the fourth piece we would. It handed us something better than a confirmation: a clear statement of what our instrument had been missing, and an instrument we can now test.
Certainty is not a measurement
The second finding is the one we think is more valuable, because it is about the technology rather than about us.
With every read, our Council states its confidence, a number we designed to swell and shrink with the quality of the evidence, and a severity, for how extreme conditions are. In the control group piece we noticed that the confidence had not moved across the sample. When we went back and checked every decision the Council has ever recorded, nearly two hundred of them, the confidence was the same number on every single one. Its severity had never once reached the top of its scale. Then we looked at our twin, which has none of our apparatus, and its confidence was likewise a single number on every decision it has ever made. That is what turned a finding about our system into a finding about the model. Asked to rate its own certainty, a language model does not use the range it is given, and giving it a wider range does not help. Self-reported certainty is not a measurement, and we now key nothing that matters to it.
We found two places in our own wiring where we had. Our guardian over the deterministic book (the sixth piece) could take its fullest action only when the Council’s severity reached the top of its scale, and we had set a specialist consultation to be called when its confidence fell. Both waited on levels the model does not produce, so neither had ever fired, and both had read on the page as strict. Our audit found them, and the remedy became a standing rule for our platform. We let the model write the thesis. We let arithmetic decide when something fires. Quantities we compose from the model’s read, or compute from the data, do reach their extremes: over the same runs on which the confidence never moved, a score we assemble from the five Forces’ directions and convictions used its whole range. Composition restores what self-report cannot supply.
We now compute a conviction from things our system already measures, the completeness of the evidence, the agreement among the Forces, and how their recent calls have graded, and we keep it beside the model’s own number rather than in place of it. That choice matters. Deterministic code operating on a model’s number is not an independent check, and two readers that can disagree have to be shown side by side, each attributed to its author, rather than resolved into one word. Overwriting the model’s number would also have erased our only evidence of what that number does. We label the computed score unproven until the record says otherwise, and we run the guardian’s new trigger, keyed to a measured quantity, in shadow first, logging what it would have done each night without doing it.
Where the line moved
It is worth noticing what kind of change each of these was. The computed conviction, the measured trigger, the shadow log, the frozen test: every one is structure we added around the model, and every one makes our layer more deterministic than it was. The rebuilt sector inputs are the counterweight, and we think the more important half. They give the model more to reason about, not less: real differences between sectors that a rule would struggle to weigh and that a reader of language can. That is the balance we are looking for, structure where the job is to measure and fire, judgment where the job is to read and weigh, and the honest report is that we are still finding the line by moving it.
What changed in the argument of this series is small and worth stating exactly. Our map survives. The record that first seemed to confirm it was a shape, and a shape is not yet a reason; the instrument that could confirm it now exists, and we froze its test. The twin we said in the eighth piece the long and short book still owed now runs beside it every night. And the dial we thought was failing turned out to be telling us something true about the model, which is that certainty is one more thing it cannot report about itself, and that measuring it is our arithmetic’s job.
The next piece is about our operating record: what our independent manager did with the slack we gave it in the sixth piece, and what that taught us about structure and judgment.
Blackworks Capital LLC manages funds through Blackworks Capital Management LLC, an Exempt Reporting Adviser. Nothing here is an offer or solicitation or investment advice. The systems described are in research and paper-testing and do not manage client capital. Past performance does not guarantee future results.
