AI Mini-Series Part IX: The Control Group: Proving an AI Investment Process Is Real, Not Just Plausible

Part 9 of a series on building and testing an AI layer for our systematic framework. This one is about the hardest question in the project: how do you know an AI investment process is actually working, when it cannot be backtested? Our answer runs every night, and it was built to be impossible to argue with.
The first piece of this series made a promise I have been carrying ever since: the reasoning layer cannot be validated on the past, so it would be validated forward, in paper, one real day at a time. That was honest, and it left a problem quietly unsolved. Forward against what? A paper track alone tells you what the system did, not what it earned. If the Council’s book rises for a month, I still do not know whether the framework, the memory, the calibration, and the adversary produced that result, or whether any capable model reading the same data would have done as well, or whether a coin would have.
The question has teeth because the apparatus is now substantial. Everything this series has described was added on an argument: the framework to bound the read, the wall to keep fluency away from the numbers, the Heretic to manufacture doubt, the graded memory, the calibration feedback. Every piece is reasonable. None of it was proven in advance, because none of it could be. And from the inside, a plausible system and a working system look identical. The prose is confident either way. The decisions arrive on time either way. Medicine met this exact problem long ago, and its answer was not better arguments. Its answer was the control group.
So we built one. We call it the Null Council.
A twin with everything removed
The Null Council is the same model that runs the real one, reading the same data, making the same kind of decisions for the same paper books, under the same caps and the same checks. Everything else is removed. It has no framework. It carries no graded memory of its own record. It receives no calibration feedback. No adversary cross-examines it. It keeps yesterday’s book and nothing more, because even a control needs continuity to trade. It is, quite deliberately, the thing this series has argued against from the beginning: a capable model pointed at the market and asked what to do. If that alone is enough, I want to know, because it would mean the apparatus is decoration.
The comparison is built to be clean. Both versions’ books are computed under identical conventions and labeled as computed rather than executed, even where that honesty costs us a flattering chart. The design is paired: because the twin runs the same books, the difference between the two moves only on the days they disagree, so evidence accrues with every divergence rather than with every year. This series has already admitted we cannot wait years to learn nothing.
A control is also only as good as the discipline around it. The twin’s instructions are frozen, their version recorded, and any change starts a new era in the record, because the subtle way to cheat here is to keep adjusting your own control until it loses. When we extended its record backward by replaying it over the archived inputs the Council had actually received on each past night, the instructions were frozen before anyone saw a result, because hindsight contaminates experimenters at least as well as it contaminates models. And nothing the twin produces is ever shown to the Council, its agents, or its gates, in either direction. A control that leaks into the experiment stops being a control.
The disease we screened for was not the one it found
I should admit what I expected it to catch. After this much manufactured doubt, an adversary trained to object, a memory that grades every call, feedback that recites the misses each night, the failure I feared was timidity. A system carrying that much suspicion might compress into permanent, hedged neutrality: never wrong, never useful. The guardian’s silence in the sixth piece was audited and found mostly honest, but the wider worry stood. Was the whole apparatus quietly sitting on the model’s chest? The Null Council was built, first, to answer that.
It answered by finding the opposite. The caveats come first because they are load-bearing: the live sample is weeks, in one kind of market, and the performance differences sit inside noise. We treat the shapes as the findings and the scorecard as unfinished. But the shapes were consistent. The Council’s conviction setting, the one that is supposed to swell and shrink with the quality of the evidence, sat pinned at its ceiling for the entire sample. A dial that never moves carries no information, whichever end it sticks on. The guardian’s stillness was the design working; the dial’s stillness is the design failing. The shortest-horizon calls were sized near the full daily move the options market was pricing, which is not what suppressed conviction looks like, and they leaned the way the tape had just leaned, into a market that kept snapping back. Graded coldly, a standing forecast of nothing beat them. Not timid. Unmodulated. We built the control expecting a system afraid to speak and found one that would not stop.
And the twin? It barely played. Mild optimism or neutrality, never once a bearish day, and ahead on most of the books over the sample, precisely by not playing, in a tape that drifted upward and rewarded staying put. I owe the control the same honesty I owe the Council: that is a fair-weather victory. A twin that never says bearish has not yet met a market that punishes optimism, and a control that wins one regime settles nothing, which is exactly why it runs forever rather than for a month. The one book where the Council came out ahead of its twin was the one the last piece described, where its rotation diagnosis had a vehicle built to act on it. The control did not just score the apparatus. It showed where the apparatus pays.
The schema was writing the script
The deepest finding was not about either contestant. It was about the form they are made to answer in.
At the short view, the Council’s record skewed bearish, roughly three calls down for every one up, against a market that rose on more days than it fell. At its longest view it has never once made a bearish call. Neither shape made sense until we read the drivers each call cited. The bearish days overwhelmingly pointed at volatility and dealer positioning, the fragility of the tape, and almost never at weak fundamentals. The model was diagnosing fragility, which is a claim about how wide the outcomes might swing. Our schema demanded a signed direction with a citable driver. There was no box for wide. So fragility came out dressed as bearishness, night after night, because the form we built could not say the thing the model actually knew.
The third piece argued that a reasoning read must become structure to be usable, and it must. The control taught the next lesson, and it is sharper: the structure is itself a hypothesis about what the machine knows, and the hypothesis can be wrong. A form that must produce a direction will manufacture directions, the same way a model asked for lessons will produce lessons whether or not any exist, the finding the seventh piece was built around. So the vocabulary is being rebuilt around the answers the machine has earned the right to give: whether there is an edge at all, which way it points, and how wide the outcomes look, with no edge as a first-class answer and fragility feeding the width of the claim rather than its sign. The grading moves with it. Judged in buckets against a do-nothing baseline, the week-out view already clears a bar the day-ahead view fails, which is the Heretic’s lesson (the fifth piece) in new clothes: what you choose to measure decides what you are able to see.
The adversary that does not argue
The piece on the two roles and the leash that holds them (the sixth) ended on a question I refused to dress up. The most consequential act in the whole layer, the final allocation, was the one place no adversary stood, and my guess on the record was that we would end up building one. We did. It is not the one I expected to build.
The skeptic I imagined was another agent, an arguer standing over the allocation the way the Heretic stands over the read. But an arguer at that spot would re-run a contest of prose at exactly the point where prose should have no authority left, and arguments can be lost by the right side. The Null Council is a different kind of adversary. It does not read the Council’s work. It does not respond to it. It simply does the same job without the reasoning, under the same rules, and the daily gap between the two is the reasoning’s paycheck, computed in the open, updated forever. You cannot out-argue a control. You can only out-perform it.
The same pattern now runs down through the layer. The long and short book from the last piece owes its own twin, a mechanical ordering pushed through the identical machinery, so that what its reasoning adds becomes a number rather than a claim. Doubt at this firm started as an agent, deepened into a battery of tests, and has settled into a standing population of controls. Each step argues less and measures more.
We looked for precedent before writing this piece and could not find another fund that publishes a running control against its own AI. If one exists, I would genuinely like to compare notes, because the method deserves to be ordinary.
What the control changes
Findings are cheap. What matters is what they are allowed to change, and under what rules. The revisions the control has demanded, a conviction setting that must earn its movement, a vocabulary rebuilt around edge and width, calibration that pushes against recency instead of reciting it, will ship as one marked revision, a new era in the record, with its success criteria written down before it runs. The alternative is the loop the seventh piece refused, one level up: a system quietly tuned, change by change, until it beats its own control, having learned the test instead of the market.
And one rule outlives every finding: whatever the Council becomes, the twin runs beside it. The adversary we built to argue that we are wrong (the fifth piece) set the rule: an AI layer earns a larger role here in proportion to how thoroughly it survives doubt. The Null Council is doubt in its most durable form: never tired, never persuaded, never impressed. What remains ahead of this series is the rest of the concrete half: where the candidate strategies come from in the first place, what it takes to run a machine like this honestly, and what the record of the book we built finally says. The conviction from the first piece has not moved. The direction is certain, the timeline is not, and the distance between those two beliefs is crossed exactly this way, one control at a time.
Blackworks Capital LLC manages funds through Blackworks Capital Management LLC, an Exempt Reporting Adviser. Nothing here is an offer or solicitation or investment advice. The systems described are in research and paper-testing and do not manage client capital. Past performance does not guarantee future results.
