AI Mini-Series Part XI: Obedient by Design: Why a Model That Takes Guidance Well Needs Room to Disagree

Part 11 of a series on building and testing an AI layer for our systematic framework. This one is about our operating record: what our independent manager did with the slack we gave it in the sixth piece, how far we swung toward structure in response, and where we now draw the line between structure and judgment.
In the last piece we graded the read. In this one we grade the operation. When we described the two roles and the leash that holds them (the sixth piece), we made a choice we still stand behind: we gave our independent manager some slack, because a role pinned at every joint teaches you nothing about how it would manage capital, and beta is when you most want to watch a thing behave. We have now run it in paper across four books for most of a summer, with a twin beside it. The record taught us where slack goes. It also taught us something about our own response, which is the more useful lesson. Most of what we changed this summer was an effort to make the independent manager’s behavior measured and repeatable, which is what systematic means, and in places we swung too far, toward an independent manager so structured that the judgment we built it for had little room to act. Finding that balance is the hardest part of our project, and this piece is our honest account of one swing of the pendulum.
A view is not a position, until a rule makes it one
Start with our Chair, the agent we built to turn the five Forces into a daily view. Over a run of consecutive nights we measured its shortest-horizon call against what the market had done that day and what it did the next. On days it called bearish, the index had already fallen; on days it called neutral, it had already risen; and its call for the following day had no value, because bullish and bearish calls preceded the same small average move. We said in the control group piece that our Council chases at the shortest horizon. This is what that looks like measured, and it is consistent with everything we have said in this series about where the machine adds least. The structured view we ask it for moved more than the world did, for the same reason its stated confidence never moves at all. Neither is a measurement.
On its own a wobbling read is harmless. A view is not a position. What turns a view into a position is a rule, and we had written one. We had written a rule that required our independent manager to act whenever the Council’s read changed. We wrote it so the manager would not sit through a real change in the market out of timidity, which is a reasonable thing to want. In practice, every wobble in the Chair’s read became mandatory turnover across our books on the same afternoon. One example stands for the pattern: our rotation book shorted the energy sector after oil had already fallen, because the read had changed, and unwound it when the read changed back. The lesson generalizes. A mandate to act has to be keyed to something that moves only when the world does. We had keyed ours to a number the model reports about itself, and a mandate for decisiveness on a noisy number manufactures decisiveness.
The model took our guidance well
The deeper pattern took us longer to see, and it is the one we would most want another builder to hear.
We wanted our Council to take a holistic, context-driven view of the market, the kind a rule cannot take, and to disagree with our rules when the context warranted it. What our record showed was a model that took guidance extremely well, and guidance that had gradually become our own rules’ answers. We precede the macro Force with a deterministic detector that scores how much macro matters that day, and on days it read dormant we had instructed the analyst to keep its report short and defer to the other Forces. It complied on most of the nights we examined, which meant that in a tape where rates were repricing at record levels, the one Force we built to read that had been told to stand down by a score we had calibrated to crisis. Our sector agent, handed only price, produced price. Our Chair, shown the day’s volatility move, restated it.
None of this was the model circumventing our intent. It was the model following our frame with great fidelity, and our frame had become rule-shaped. A language model anchors on the most explicit thing in front of it, and over the summer the most explicit things we put in front of it were our own computed scores. Our desire to instill structure had made the layer more deterministic and less AI-driven, one reasonable rule at a time.
Three design consequences follow, and they are our way back. We feed the model what rules cannot read: the language of statements, calls and filings, the real differences between sectors, the context a number compresses away. We withhold our rules’ answers until it has committed: we let it form its own view of the Forces first, then show it what our deterministic reads say and ask it to reconcile, and we record both, because a reader that has seen the answer will echo it, and an echo is not corroboration. And we let our detectors inform rather than govern: a dominance score should tell the Chair how loud a Force is, not tell the analyst to be quiet. We said in the first piece that the whole value of a second reader is its independence. We have learned that independence has to be protected in the wiring, not just intended.
Tightening, and the review we turned on ourselves
The rest of our response was structure, and we want to be plain about both halves of that.
We changed what our independent manager is allowed to act on. Its mandate no longer answers to the Council’s daily read but to measured change in the world: it holds until something it can name has moved, and when it moves, it has to say what moved. We also handed it its own graded record against the twin, which our memory design (the seventh piece) had always intended and which, we discovered, we had never actually put in front of it. Nothing errors on a missing paragraph. The discipline that found it was reading what we actually show the model rather than what our design says we show it.
Then we put our own changes through the doubt our Heretic prescribes: independent review, assigned to disagree with us, before anything shipped. The best catch was a rule we had written to let a book out of its leash in a crisis. We had keyed it to a condition, and a reviewer pointed out that the condition can persist for the length of an entire bear market, so our rule would have removed the leash for exactly the period we most wanted it on. We rewrote it to answer to a change in conditions rather than to the condition itself. That lesson generalizes to every trigger we write: key it to what changes, not to what is. It is the kind of rule that is obvious once stated and invisible until someone is asked to look for it, which is our whole case for reviewing our own structure the way we review the model.
Here is the other half. Every one of those changes is structure, and structure is deterministic. An independent manager we consult only when something measured has moved, hold within caps, and otherwise carry forward is closer to a rules engine than we imagined in the sixth piece. In the control group piece we worried that our apparatus might be decoration around a model doing the work. Our worry now is the mirror image: a model that has become decoration around our apparatus. We swung too far in places, and we think that is what a learning process looks like from the inside. You swing toward structure to make a thing measurable, you watch the record, and then you hand judgment back where the record shows it earns its keep.
Drawing the line on purpose
Our newest book is our first attempt to draw that line by design rather than by correction. A deterministic screen we wrote narrows the market to a short list each day, and our independent manager chooses within it. Structure decides what is eligible; judgment decides what is chosen; and we grade the choice against the screen alone, so that what the reasoning adds is a number rather than a claim. The screen is the measured part. The choice within it is the model’s. That is the balance this series has been circling, made into architecture, and it is where we expect our pendulum to come to rest: not a model fenced until it cannot act, and not a model pointed at the market and asked what to do, but a model given a real decision inside a frame that makes the decision measurable.
What our record can say
None of this was visible in the returns, and it is worth being exact about why. Over the summer our twin was ahead of the independent manager on most of our books, and our rotation book trailed a simple equal-weight benchmark. It would be easy to call that a verdict. The arithmetic from the last piece says it is not: our record could only detect an edge of a third to a half of the book per year, the gap reversed sign on most of our books when we split the sample, and pooled together the gap is indistinguishable from zero. And every behavioral measurement in our record comes from a single regime, a summer in which the index went nearly nowhere while sectors dispersed widely underneath it. A flat read in a flat market may be correct. Structural findings hold in any regime; behavioral conclusions from one do not.
So the decisions we made were refusals. We did not retire or re-weight a book on the evidence. We wrote down, before more data could arrive, what would count and what would not. And we put our search for proof where it can be found soonest, in the mechanism rather than the returns. The fairest description of the independent manager’s summer is not that it failed but that it was unfalsified, in a tape that could not have falsified much, and the fairest description of ours is that we learned where the line was by crossing it.
What it takes
In the sixth piece we asked what it takes to let a machine manage money and answered with a leash. After a summer of running it, our answer is longer and more specific. It takes a record we cannot edit and a control we cannot argue with. It takes reading what we show the model rather than what our design says we show it, because the missing joints are silent. It takes keying every action to a number that can move, and none to a number the model reports about itself. It takes reviewers we assign to disagree with us, and the habit of publishing what they found. And it takes holding two things at once that pull against each other: enough structure that our layer is measured and repeatable, and enough room that there is still a reasoning layer to measure. We have not found that line. We have found how to look for it, and where to draw it on purpose.
What remains of this series is the record of the book we rebuilt, and we will write it when our pre-registered test has the windows it needs and not before, whichever way it lands. The conviction I stated in the first piece has not moved. The direction is certain, the timeline is not, and the distance between those two beliefs is crossed exactly this way, one honest review at a time.
Blackworks Capital LLC manages funds through Blackworks Capital Management LLC, an Exempt Reporting Adviser. Nothing here is an offer or solicitation or investment advice. The systems described are in research and paper-testing and do not manage client capital. Past performance does not guarantee future results.
