AI Mini-Series Part VIII: Fluent Everywhere, Valuable Somewhere: Where AI Actually Adds Value in a Systematic Fund

Part 8 of a series on building and testing an AI layer for our systematic framework. This one is about the question the series has been walking toward from the start: which strategies are actually built to use what this technology does best, and which only look like they are.
The last piece ended with a rule about memory: it earns its way into the loop by passing the test, and not a day sooner. That standard settles one question and uncovers a larger one behind it, the one I promised in the first piece and have been deferring ever since. Where should this technology be pointed at all? I believe AI will become the dominant method for managing systematic investing. I have also said from the start that this conviction is not a claim that AI improves every strategy equally. Certain kinds of strategies are built to use what a language model does best. Others gain very little from it. Knowing which is which is not a detail along the way. It is the difference between owning an instrument and owning a gadget.
The reason the question is hard is that the model will never answer it for you. Ask a capable model about anything, a currency cross, a commodity roll, where the index closes on Friday, and the answer arrives in the same confident, fluent voice. The fourth piece was about what that fluency hides inside a single answer. Here the problem is wider. Fluency is uniform across domains, and edge is not. A tool that sounds equally good everywhere cannot be allowed to choose its own ground, because it will happily claim all of it. So we did not ask the machine where it should manage money. We drew the map ourselves, by subtraction: start from everywhere, and cut away the territory where the evidence says the reasoning does not earn its keep. What is left is small, specific, and worth more than the whole map it came from.
Where it adds the least
The first cut is the largest, and it runs straight through our own fund. Most of the deterministic book lives at a fast tempo on price, reading stretch and pressure in the tape and acting on it within a day. The inputs to those decisions are already clean numbers before any judgment is asked for. The first piece called that distillation and defended it: compressing a Force into a single value is what makes a rule fast, testable across decades of history, and immune to the mood of whoever is watching it.
Where distillation is the feature, reasoning is overhead. A language model brought to that ground has nothing to hold, because the whole point of the signal is that the context has already been squeezed out of it and what remains still works. There is no policy statement to parse inside an oversold reading. The third piece named what happens if you insist on involving the model anyway: handed the same inputs the rules already vote on, it produces the same conclusion in better prose, a second voice that agrees by construction. And the validation asymmetry from the first piece, that a rule can be tested on decades of history while a reasoning read can only be tested forward, decides the tiebreak. The rules carry decades of out-of-sample evidence at that tempo. The model’s read cannot be backtested at all. A contest between a tested instrument and an untestable one, held on the tested instrument’s home ground, is not a contest.
So the fast, price-mechanical center of systematic trading, the territory most people picture when they imagine an AI trading, is where ours adds the least, and the deterministic book that lives there stays exactly as it is.
The tempo it can defend
The obvious rejoinder is to keep the machine off the rules’ turf but let it keep the speed: give it its own inputs, the language and the context, and let it trade fast on those. This is the day-trading fantasy, and it deserved a test rather than an argument.
We built the test into the system. Alongside the Council we run a twin that is stripped of the apparatus: the same model, reading the same data, with no framework, no graded memory of its own record, no calibration, and no adversary. Its purpose is the same as the permutation null in the Heretic’s battery of tests (the fifth piece): destroy the link between signal and outcome and see how good the results look when nothing real is underneath. Before crediting the structure, measure what survives with the structure removed. And separately, we read back through the Council’s own graded record at its shortest horizon, the view it publishes every night about the next day.
I want to be careful here, because the sample is weeks and one kind of market, not a verdict. But the shape of what we found was consistent, and I noticed it in the record by eye before the measurement agreed. At the shortest view, the Council’s calls tracked the tape it had just seen. After a run of down days it kept leaning down, into a market that kept snapping back. A reasoning engine reduced to reacting is a lagging indicator with a vocabulary, and in a tape that mean-reverts, following the recent past is worse than saying nothing. A week out, the same record improves to defensible. The differences sit within noise at this sample size, so we treat the scorecard as unfinished and the shape as the finding: the shorter the horizon, the more the machine’s read collapses into a narration of what just happened.
The outside record says the same thing louder. Our research surveyed the public, live evaluations of language models trading at daily cadence, and the honest ones converge on one result: performance indistinguishable from simply holding the market, or worse. No credible group has published the opposite. The most popular picture of AI in markets, the machine that out-trades the tape day by day, is the one picture with no evidence behind it anywhere we could find.
So the tempo the machine can defend sits in a band. Below a few days, the tape moves faster than the reasons, and the model degenerates into chasing. Too far beyond a few weeks, the opposite problem arrives: honest grading needs outcomes that do not overlap, and at long horizons you would wait years for enough independent verdicts to know anything. Days to weeks is where the read is both formable and gradable. That is the band we point it at.
Diagnosis is the edge
The most valuable lesson in that record was not about tempo at all. It was about job description.
The same graded history that embarrassed the shortest-horizon calls contained one clear bright spot, and its shape has changed how I think about the entire layer. On a meaningful share of days, the Council correctly diagnosed that leadership inside the market was rotating: that the story was not the index going somewhere, but money moving from one part of the market to another. Read as an index forecast, that diagnosis earned nothing, because knowing leadership is rotating does not tell you whether the market closes higher. But one of the books it helps manage exists to trade exactly that, holding the parts of the market gaining leadership against the parts losing it, and in that vehicle the same read paid. One diagnosis, two vehicles. Where the vehicle matched what the machine knew, it added value. Where the vehicle demanded a forecast it did not have, it subtracted.
Stated plainly, the division sounds almost obvious. The machine is good at saying what is happening. It is unreliable at converting what is happening into what happens next, and worse at deciding how much to bet on it. So we split the job along that seam. The model diagnoses the state: which Force is driving, what kind of tape this is, where the pressure sits. Code converts the diagnosis into positioning, with the conversion written down, tested, and fixed where the model cannot argue with it. The third piece drew this line inside the read itself, computed values on one side and reasoning on the other. The benchmark taught us to draw the same line one level higher, when choosing strategies: point the machine at books where the diagnosis is the edge and the conversion is mechanical. Keep it away from books where the edge is the conversion.
The ground that remains
After the subtraction, the territory left is not a leftover. It has a specific shape, and every feature of it maps to something the machine has shown it can do.
It is relative rather than absolute. Asking which parts of a market are better positioned than the others suits a language model far better than asking what something will be worth. A ranking has no magnitude to overstate, which quietly removes the danger the third piece worried over, a number carrying more authority than it earned. Ranking across a cross-section also cancels the market’s common direction, which is precisely the part of the future the machine proved worst at guessing. And a ranking can be doubted cheaply: swap the model for a simple mechanical ordering, run the identical book, and the gap between them is exactly what the reasoning is worth. The doubt comes built into the design, which is the standard the Heretic (the fifth piece) set for anything that wants to be trusted: it has to survive being argued against.
It is language-rich and breadth-hungry. The advantage this series opened with, taking in more context from more sources closer to real time, compounds when a decision spans many instruments at once. A person can hold the full context of a handful of names. A rule holds none of it. The machine holds the context of an entire cross-section at roughly the cost of holding one name, and the strategies that reward this are the ones with many moving parts to compare at once: sectors against sectors, industries against industries, the long tail of instruments that no small firm could cover by hand.
And it turns over inside the defended band, days to weeks: fast enough that graded outcomes accrue and a record builds, slow enough that the read is reasoning about drivers rather than reacting to prints.
Why our size sharpens the aim
One more constraint shapes the map, and it is economic rather than technical. Edges decay in proportion to how much capital can crowd into them. The most durable inefficiencies are the ones that stay too small to matter to large funds: too niche, too capacity-constrained, too costly to find relative to the size of the prize. A firm managing billions cannot justify staffing an edge that stops absorbing capital long before it moves their number. We are not a firm managing billions, and in this design I have stopped thinking of that as a limitation. It is the aim.
What kept small edges unharvested was never that nobody suspected they existed. It was coverage. Finding them, verifying them, and watching them by hand costs more than they pay, so attention went where size could follow. A reader that holds the long tail’s context all at once collapses that cost. For a fund our size, the honest target is therefore not the crowded center, where the largest firms compete on speed and scale we will never have. It is the periphery, where the binding constraint was never intelligence but attention, and attention is the input that has just become abundant.
I want to state that claim at its honest size. This is a thesis about where to hunt, not an announcement of treasure found. We have built a research library and screens to work that periphery systematically, and everything they surface faces the same battery this series has described. The bar does not move because an idea came from a promising neighborhood.
The book this points at
Now put the criteria in one place. A tempo of days to weeks. Relative calls across a cross-section rather than absolute forecasts. Rich, many-sourced context that arrives as language. The machine confined to diagnosis, with conversion and sizing fixed in tested code. The market’s common direction, the machine’s weakest subject, removed from the bet. And doubt designed in from the first day, including the cheap swap to a mechanical ordering that gives the reasoning a number to beat.
That is not a wish list. It is the specification of a strategy we have already built and funded in our paper environment: a long and short book that holds the parts of the market the diagnosis favors against the parts it disfavors, neutralized and sized by code, the model’s authority confined to the ordering, every cap fixed where it cannot argue. It has been running for weeks, deciding daily, graded daily. A mechanical twin ordering the same book is the doubt it still owes. Its record is where this series ends: the strategy we built on the ground the evidence pointed to, reported the way this series reports everything, including whatever the record says that I would rather it did not. The pieces between here and there are the concrete half of the work: how we prove a reasoning layer is real rather than merely plausible, how strategies get researched, tested, and killed here, and what it takes to run a machine like this honestly.
I opened this series with a conviction and a caution: the direction is certain, the timeline is not. The map in this piece is how the two live together without softening each other. Not everywhere, not yet. One book, built where the machine has shown what it is good at, fenced where it has shown what it is not, and promoted only by what it earns.
Blackworks Capital LLC manages funds through Blackworks Capital Management LLC, an Exempt Reporting Adviser. Nothing here is an offer or solicitation or investment advice. The systems described are in research and paper-testing and do not manage client capital. Past performance does not guarantee future results.
