Artificial Intelligence

AI Mini-Series Part IV: Fluent Is Not Proof: Why AI Answers Can Look Just as Good Wrong as They Do Right

Blackworks Capital Team
AI Mini-Series Part IV: Fluent Is Not Proof: Why AI Answers Can Look Just as Good Wrong as They Do Right

Part 4 of a series on building and testing an AI layer for our systematic framework. This one is about the hardest problem in the project: making a fluent model’s read trustworthy enough to inform a decision about real money.

Where we are: We are building an AI layer we call the Council: a second reader of the same Five Forces our rules-based strategies already read, running in paper, not with client capital. Its inputs are held to a rule we will not bend, and its read is bounded by the framework and becomes a structured view rather than a vote. This piece is about the hardest problem in the project: making a fluent model trustworthy, by keeping its prose away from the numbers that act. New to the series? Start with the one-page guide.

The last piece of this series ended on a hard problem. The moment you ask a language model for a score instead of a paragraph, you give it the power to be confidently, structurally wrong, and a score carries an authority that prose does not.

The danger is specific, and worth stating plainly: a fluent model’s confident wrong answer looks exactly like its confident right one. There is no tell, no hesitation in the writing to warn you. So the whole task of making the AI trustworthy comes down to a single rule. Its confidence is never allowed to count as evidence. Nothing it says moves real money on the strength of how convincing it sounds.

What follows is the search for the places that rule could be broken, closed in an order. We start with the errors that any honest check would surface, the ones where the model is simply wrong about a fact you can look up. Then we climb to the errors that leave no trace, where the model is wrong about a read and the writing gives you nothing to flag. And we end at the errors that get harder to catch the more reasonable they look, the ones that flatter you into agreeing because they wear the shape of progress. The wrongness does not get louder as you climb. It gets quieter, and better dressed.

It does not get to own the facts

The lowest rung is the easiest to defend, because most of reading a Force comes down to arithmetic that has to come out the same way every time. There is little judgment in it.

Take the macroeconomic Force. Part of reading it is deciding whether the market is under genuine structural stress, the kind that surfaces in credit and funding conditions before it ever reaches price. Those conditions are computed by deterministic code. Each one either meets its threshold or it does not, the same way every day. The model receives the result as a fixed fact and writes the interpretation on top of it. It can explain what a firing condition means in context. It cannot change whether the condition is firing.

We did not start there, and the reason we moved is the whole argument in miniature. An earlier version let the model judge some of those conditions itself. It leaned one way, systematically, and it did so in a voice that read as careful and sure. A person skimming the output quickly would not have caught it, because nothing on the face of the writing announced that the read was tilted. The prose was as clean as ever. That is the danger at its smallest and most instructive: a confident interpretation sitting on top of a judgment that should never have been the model’s to make. We took the judgment away. It now lives in code that behaves identically every run, and the model does the part it is built for, which is saying what the facts mean rather than deciding what they are.

The same discipline applies to any number the model puts in its writing. Every figure has to trace to something we can point to, and every figure is checked. The failure worth guarding against here is not an invented number; an invented number is loud. It is the model lifting a real figure out of something it read, a figure that happens to be wrong, so that a check comparing the model against its own reading material waves it through, because the wrong number really did appear in the text. So we check the other way, against the source the number is supposed to come from, not the article the model happened to find it in. The question is never whether it quoted faithfully. The question is whether the thing it quoted was true.

So far the wrongness has been the kind a careful audit surfaces, because in every case the model was reasoning over a fact, and a fact can be checked. The next rung offers nothing to check against.

Its fluency never reaches the decision

Take the facts away from the model and a harder problem remains, because the model’s real work is not reporting facts. It is reading context and forming a view, and a view has no source document to test it against. When the model interprets a policy statement or weighs a surprise, there is no column of true values to hold the interpretation up to. This is where the danger from the opening bites hardest, because now the confident-wrong and the confident-right are not facts that disagree. They are two reads, written in the same assured prose, and nothing on the page separates them. This is the heart of the matter, and it is where most of the structure in the system is spent.

The most important piece of that structure is a wall. The scored read the rest of the machinery acts on is computed in code, from validated signals, before the model is asked to say a word. The model’s reasoning still matters. Its job is to read context and to propose, not to set the number. When it reads a statement and concludes something useful, that conclusion is a candidate, and a candidate does not move the acted-on score until it has been turned into a defined signal and put through the same testing as everything else we run. Until then it is narrative, and the narrative is walled off from the dial.

We do not just trust that the wall holds. We test it. The model is shown some context purely for color, real things that have not earned a place in the decision, and we have flipped what that context says and confirmed the acted-on score does not move. If it ever did, the model’s prose would have found a quiet path into the outcome, which is the one thing the wall exists to prevent. A model that can describe the weather is useful. A model that can change the forecast by describing it differently is a liability.

The wall has a companion, and it guards the same invisible failure from the other side. The model has to be willing to say nothing. Silence is the hardest behavior to get from a fluent system, because producing fluent text is the one thing it is built to do, and left alone it will always produce something. So we spend real effort making it comfortable reporting that it does not have an answer. When an input is missing or cannot be trusted, the model is required to say so plainly, and the system shows the gap rather than smoothing it over. An honest “unavailable” is worth more to me than a confident guess, because I can act on a known gap, and I cannot act on a guess I have mistaken for a fact. The wall keeps the model’s prose from moving the number. The willingness to be silent keeps it from filling a hole the same way.

Both of these defend against an error you cannot see by reading the output. They also share an assumption worth examining, which is that the suspect read comes from the model and not from us. The last rung is harder still, because the error arrives looking like a good idea, and it is our own.

Its own ideas face the same bar

The most seductive error in the whole project came from the model itself, and it nearly worked on us. Much of our research is AI-assisted now, and at some point our research produced a candidate method for combining the Forces into the regime score, with the model putting it forward. The proposal was clean. It was internally consistent, it read as the kind of thing a thoughtful quant would build, and it would have been easy to start trusting. That is the trap, and it is a different trap from the first two. There was no fact to check, because it was a method. There was no prose sneaking past the wall, because it was a real, structured idea. It simply looked right, and looking right is the most dangerous property an idea can have when the people evaluating it would be flattered to be correct.

There is a wider version of this trap worth discussing, because it has become common enough to shape how people invest. A great deal of what gets built and sold right now rests on a quiet assumption: that if an AI produced an idea, or an AI signed off on it, the idea must carry an edge. It does not follow. An edge is what survives honest testing, and what proposed the idea has no bearing on whether the edge is there. A fluent, AI-authored proposal is, if anything, the kind most likely to be mistaken for a tested one, because it arrives sounding like the conclusion of work that was never actually done. We said in the first piece that the reasoning layer cannot be backtested the way a rule can. That asymmetry cuts both ways. It is why the AI is worth having, and it is why an idea it hands us earns no credit for having come from it. I believe AI will take on a large role at this firm. It will take that role the same way every strategy does, by passing the test, and not a day sooner.

So we held it to the same bar we would put in front of a strategy a stranger was trying to sell us. It did not pass. The problems were ordinary, which is the point, because they are the ones that flatter quantitative work everywhere. The result was measured on overlapping samples that make an edge look more certain than it is. It was never held back from a stretch of history it was not allowed to see while it was being built. And it leaned too heavily on a single input, so that one source was quietly carrying the conclusion.

The idea did not become part of the read. It went back to be rebuilt and retested, like anything else that has not yet earned its place. The standard applies to the machine’s own proposals first, and it applies hardest there, because that is where the temptation to relax it is strongest. The model does not get to grade its own homework, and neither do we when we are the ones who liked the answer.

Trust is what we refuse to let it do

None of this makes the model smarter, and that is the point. Each refusal removes a path by which it could be confidently wrong in a way I would not catch in time. It does not get to own the facts; the arithmetic of reading a Force is decided in code, and every number it cites is checked against the source rather than the article. Its fluency never reaches the decision; the acted-on score is set before it speaks, the wall is tested, and an honest gap beats a fluent guess. Its own ideas face the same bar as everyone else’s, and the better they look, the harder we hold the line. What is left, after all of those refusals, is a model doing the one job it is suited for, reasoning in language over information it is allowed to trust, with no way to let its confidence outrun its evidence.

There is one refusal left, and it is the deepest, because it points the suspicion at us rather than at the model. A disciplined team can still talk itself into a confident answer that everyone agrees with and that happens to be wrong, and the only defense is to build something whose whole purpose is to argue we are wrong: wrong about the read, and wrong about whether we are even measuring the right thing in the first place. That is the next piece.

Blackworks Capital LLC manages funds through Blackworks Capital Management LLC, an Exempt Reporting Adviser. Nothing here is an offer or solicitation or investment advice. The systems described are in research and paper-testing and do not manage client capital. Past performance does not guarantee future results.

Start a Conversation

Speak directly with the Founder and Fund Manager about the BWC Founders Fund, our systematic approach, and whether it aligns with your investment objectives.