AI Mini-Series Part V: The Heretic: An Argument We Build to Lose

Part 5 of a series on building and testing an AI layer for our systematic framework. This one is about the doubt we build into the system on purpose, and how far we let that doubt reach: from the read, to the methods behind the read, to the scorecard we grade everything against.
The last piece of this series ended on the deepest refusal of all, the one that points the suspicion back at us rather than at the model. A disciplined team can talk itself into a confident answer that everyone agrees with and that happens to be wrong, and the defense I said we needed was something whose whole purpose is to argue we are wrong. This piece is about that something.
Here is the spine of it. Strength does not come from a system agreeing with itself, any more than it comes from a team agreeing in a room. It comes from disagreement, and disagreement is fragile when left to chance, because a system, like a team, settles toward consensus on its own. So we manufacture it, at more than one level, because the most expensive errors tend to hide one level deeper than wherever we last aimed the doubt. At the surface is the doubt that runs: a live agent whose only assignment, each time the system forms a view, is to argue that view is wrong. Beneath it is the doubt that built the system, the tests we put our own research through before any of it was allowed to matter. Under that is the hardest doubt of all, aimed past any single answer and past any method, at the scorecard we grade everything against. Only the first is an agent. The other two are how we work, and they are the reason the agent exists at all.
We call that agent the Heretic. The name is still provisional, but the role is fixed. The Heretic has one temperament and one job. The job is to take the read the rest of the system has produced and make the strongest available case that it is mistaken. The temperament is that it is never satisfied. It is there to lose an argument we have set up so that the only way it can win is by finding something real, and to make us pay attention when it does.
The reasoning adversary doubts the read
When the system has formed its read of the Forces, the Heretic does not see a draft it is asked to improve. It sees a position it is asked to break. It cross-examines the read the way a good opposing counsel cross-examines a witness who sounds a little too composed. Where is this leaning on a single input that would collapse the whole thing if it turned. What would the same evidence look like if you assumed the opposite and tried to make that fit. Which part of this is signal and which part is the system telling itself a story it finds comfortable. The Heretic takes the strongest objection it can find, deepens it rather than listing ten weak ones, and puts its dissent on the record next to the read.
People ask why this needs to be a dedicated agent at all. Why not simply instruct the whole system to be skeptical. The answer is the same reason a team of thoughtful people will still walk out of a meeting having agreed to a bad idea. When skepticism is everyone’s job, it is no one’s. Each participant assumes someone else will raise the hard objection, the objection feels rude against a read taking shape nicely, and the path of least resistance is assent. So we do not share the doubt. We concentrate it in one place and make it the entire reason that place exists. The Heretic is rude by construction, and it cannot defer to anyone, because there is no one else in the room whose job is the same.
The same principle runs one level down, into each individual strategy, which is where the next piece picks up. Doubt has to be owned by something, or it evaporates.
Now the part that keeps this from becoming theater. An adversary you build can fail in a quiet and flattering way, by finding nothing, and the temptation when the Heretic reports no serious objection is to feel reassured. We treat that clean bill of health as a warning instead, because in markets this uncertain there is almost always a real case against any confident read, the way a smoke detector that never chirps is more likely broken than proof the house will not burn. So the incentive is set against sign-off: the Heretic earns its keep by landing objections that survive a look, and it earns nothing for waving the read through. An adversary paid to agree is just a second narrator, and a narrator that only nods adds nothing to a read we already believe.
That is the Heretic, the doubt you can watch happen live each time a read is formed. It is also the shallowest layer. A conclusion is the easiest thing to argue with, and by the time there is a read to attack, the riskier mistakes are often already baked into the methods that produced it. Those were doubted earlier, somewhere the Heretic never sees, while we were still building.
The statistical adversary doubts the method
Doubt does not only arrive as an argument. It also arrives as a test built to kill a result, pointed at our own work before anyone else has to. None of it runs in the moment a decision is made; it happens earlier, in research, where a result can still be thrown out at no cost. An argument can be answered with a better argument, which is a contest the more fluent side tends to win. A test does not care how well you argue. It returns a verdict, and you accept it. This is the part where the technical reader and I can speak plainly; I will gloss each method as I go so the rest of the room stays with us.
The sharpest of these tests is the permutation null. The idea is almost cruel in its simplicity. Take the thing that is supposed to carry the edge, the relationship between a signal and what happens next, and break it. Shuffle which outcome is paired with which signal reading, so the signal still looks like itself but its link to what happened next is destroyed by construction, then measure again, many times over, and see how good the results look when you know nothing real is underneath. If your true result does not stand clearly above the cloud of scrambled ones, you did not have an edge. You had a pattern that randomness reproduces about as well, noise wearing the costume of signal, and the permutation null is how you make it take the costume off.
The permutation null travels with companions, because each closes a different way of fooling yourself.
The first is walk-forward out-of-sample testing, the discipline of holding back stretches of history the method was never allowed to see while it was built, then asking how it performs on those untouched stretches. Anything can be made to look brilliant on the data it was fitted to. The only performance that means anything comes from data the method never had a chance to memorize, walked forward through time the way the future actually arrives.
The second is correction for multiple testing, which handles the most ordinary sin in all of quantitative work. Test enough ideas and some of them will look excellent purely by luck, the way someone in a large enough crowd will flip ten heads in a row and feel briefly gifted. Report the handful that shone and you have not found edges. You have found the luckiest coincidences and dressed them as discoveries. The fix raises the bar a result must clear in light of how many chances it had to look good by accident. One version of it, controlling the false discovery rate, holds down the share of your reported winners that are flukes, still paying for all the rolls of the dice you took.
There is a third companion, and it guards a quieter mistake. Because market days arrive in streaks rather than as independent rolls of a die, we resample history in connected blocks, not scattered single days. This is the block bootstrap, and it matters because pretending each day stands alone makes any edge look far more certain than it is. Keeping the real clumpiness of returns intact gives you a range of uncertainty that is honest instead of flattering.
None of this is exotic; it is the standard equipment of careful empirical work, and naming it is not the point. The point is where we aim it: first and hardest at the ideas we want to be true, including the ones our own research produces. We ran a panel of independent adversarial reviewers, deliberately separated, over our own signal research, with a permutation null underneath the whole exercise. The reviewers were not there to polish the work. They were there to find the reason it should be thrown out, and the permutation null was there to catch the case where the work survived every human objection and was still, underneath, indistinguishable from a scramble of itself.
This is the same standard the fourth piece described, when our research produced a clean method for combining the Forces and it did not pass. The statistical adversary exists so that the grading is done by something that does not care whether we are right and gains nothing from sparing our feelings. Doubt aimed at the read questions a conclusion. Doubt aimed at the method questions the people who liked the conclusion, and that is a harder thing to build, because those people are us.
There is one doubt left, and it is larger than either of these, because it does not question an answer or the way we reached it. It questions whether we have been keeping score correctly the entire time.
The deepest heresy doubts the scorecard
The deepest doubt is one the Heretic never raises, and no runtime agent should. It aims past any single read, and past the methods too, at the ruler itself, a question we had to keep asking while we built: is right even the thing we should be counting?
I have always built our models around convexity, by instinct more than by formula. It is how the firm has thought about its own book for as long as we have run one. The gap worth describing here sat between that intuition and how little of it we had ever made explicit.
When we set out to build the AI’s scorecard, we reached for win-rate first. It was the easy answer, and it seemed like the right one: the model’s job is to make a call about what happens next in the market, and the obvious way to grade a call is to ask how often it proves correct.
The score it produced and the results it implied did not match the models we actually run. Ranked by how often they were right, our signals came out in an order that disagreed with everything the live book had taught us about which of them we valued and why. The number was clean and the ranking was wrong, and the distance between those two facts is what sent me back to the question more than once. A simple win-rate was missing the thing I had always understood about how we build, which is that some of our best edges are correct less than half the time.
That sentence sounds like an indictment until you look at what those edges actually do. They are wrong often, in small and survivable ways, and then occasionally they are right in a way that pays for all of the small losses and a great deal more. What drives the money is the shape of that distribution rather than how often the signal is right. A bet like that is convex. You lose a little a lot of the time and win a lot a little of the time, and over a full cycle the arithmetic is firmly in your favor even though a hit-rate scorecard says the signal barely works.
Now hold a pure hit-rate test up against a bet like that. It penalizes the signal for being wrong most of the time, the very property that comes attached to the convex payoff, and it rewards the opposite kind of signal, right far more often, with small wins and rare severe losses. A test that asks only how often you are correct throws away the very bets we most want to own. The scorecard only looked neutral. Underneath, it was selecting against convexity the whole time, silently, while wearing the face of an objective measure of quality.
So we rebuilt it. We are grading the AI’s contributions on expected value and the shape of the payoff: whether the wins, sized properly and weighted by how often they come, outrun the losses sized and weighted the same way, regardless of which column has more entries in it, and whether the distribution of outcomes is convex enough to be worth holding through the long run of small losses.
What we understood about convexity had been there all along. The work was to make it mechanical, to turn a judgment I had made by feel for years into a rule a machine can apply the same way every run, and the only reason we got there was that the first scorecard was doubted until it broke. It matters most for the AI, because a language model graded on hit rate would learn to avoid the asymmetric reads that are most valuable, while looking more accurate.
There is a fair objection worth meeting head on. This series has insisted the reasoning read cannot be backtested, and convexity is something you measure across a long run of outcomes. Both are true, and they do not collide, because the same standard can be met with very different evidence. On the signals, which carry a payoff history, we measure the shape of the distribution directly, as we always have on the deterministic book. On the reasoning read we cannot, for the reason the first piece gave, that there is no honest way to put a language model back into the past without hindsight leaking in, so it is held to the same bar but earns it only forward, in paper trading, one real outcome at a time. The bar is identical. The evidence for clearing it accrues at two very different speeds.
Grading this way also forces a distinction we now make deliberately, and I will keep it conceptual, because with none of this trading live capital, numbers would be premature and meaningless. A contribution can earn its place in two ways. It can be a risk-shaper, something that does not add directional edge but improves the shape of the payoff, trimming the worst outcomes or smoothing the path. Or it can be an edge-enhancer, something that actually lifts the expected value of the decisions themselves. These are different jobs with different bars, and a hit-rate scorecard cannot tell them apart. Once you grade on expected value and convexity, the two roles separate cleanly, and a contribution has to declare which job it is doing and clear the bar for that job. That deepest doubt, the one about the scorecard, is what made the distinction visible in the first place.
The same discipline, pushed to its limit
Our Five Forces material has always held a line that sounds almost contrarian until you have lived through enough cycles to feel why it is true. For a portfolio of systems to be robust, conflicting signals are not a problem to be engineered away. They are a benefit of holding a diverse, uncorrelated set of strategies, and the framework reads the conflict itself as information. When macro is loud and bearish while sentiment reads constructive, the disagreement across uncorrelated Forces is the framework’s protection against overconfidence, and alignment across them is its highest-conviction signal.
The Heretic is that conviction made into a role you can point at, the part of the doubt that runs live. The rest of it works earlier and deeper, in how we build and what we choose to measure, so the read can never coast on the comfort of internal agreement. And doubting our own scorecard is the same discipline pushed to the edge of where it can go, because there is nothing left to doubt after you have agreed to question the ruler you measure everything with. We did not build the Heretic to make the AI look rigorous. We built it because the willingness to argue this hard against ourselves, against the read, against our methods, and against our own definition of success, is the precondition for ever trusting the system with more than a paper environment. An AI layer earns a larger role here in proportion to how thoroughly it can survive being doubted, and not one step before.
That larger role, and how we actually intend to use this, is the subject of the next piece. Once a read has been formed, doubted, tested, and graded against the right scorecard, something still has to turn it into positioning inside a portfolio. That is where the PMs come in, the per-strategy managers that close the loop between a trustworthy read and an actual investment decision.
Blackworks Capital LLC manages funds through Blackworks Capital Management LLC, an Exempt Reporting Adviser. Nothing here is an offer or solicitation or investment advice. The systems described are in research and paper-testing and do not manage client capital. Past performance does not guarantee future results.
