I’m completely done with LLMs in the enterprise

Step 39 of a 40 step workflow

My writing is not for everyone. I am not going for major views and clickthroughs because I’m writing about running agentic AI at scale, for people running agentic AI at scale, or trying to anyway. And that what sets us all apart is that we are all working at the frontier of unsexy AI.

Why do I call it that?

Well, if you have the chance, as an AI fetishist, to work for a company that develops a 1.6 trillion parameter model with Fable-like capabilities, or you can work on the next Hermes or Openclaw, or when Yann LeCun, Demis Wasabi (I think his real name is Hassabis, but I’m not really sure), or Fei-Fei Li – the godmother of AI (LOL) – asks you to join them in their endeavors to build the next big AI hit, the world model, then why would you even want to work on Enterprise AI technology?

Article content

It isn’t sexy, it doesn’t embrace the latest technologies, the road to production is as hard and painful as a man trying to birth a baby, and it is filled with legacy systems, compliance officers who’ve never heard of a token, and a change advisory board that meets once a month whether you’re ready or not.

An enterprise is an oil tanker. It doesn’t want to change course, hit the brakes, and make a U-turn. It wants incremental changes, and it needs the entire crew to adapt to the new course, and it is sailing within well-defined shipping lanes, and deviation from those lanes can have serious legal consequences, well, at least in Europe, where the situation is even more dire.

Article content

So yeah, why work on Enterprise AI anyway? Is it because you’re no demi-god, you flunked at uni, your math skills are poor, or you were simply born on the wrong continent?

Well, I’m going to bust your bubble.

No, it isn’t sexy at all. That one is correct. And no, we are not working with the latest and greatest stuff at all, but the thing is that the enterprise is so demanding, and the forests of rules are so thick that you need to use all of your wits to deliver a product that really works. That is why I have even dedicated a two year research program and the precious (drinking) time of students to solving a few of the most pressing problems that the sexy frontier guys find beneath them..

Yeah, the thing is that all the frontier labs want to be sexy, and then they completely forget to understand what makes their enterprise customers tick. And I don’t mean the sexy demos they churn out every month. No, I am talking about what creates ROI.

RO-what?

Businesses adopt tech to give them a competitive advantage, or it saves them time or money. In the end, only the results count. And this is what makes our collective businesses the best thermometer about the actual state the AI industry is in. And I must say, it is in dire straits.

Have you come across a frontier lab investing in making your AI project a success by helping you change the technology so you can improve your margin? No, because instead they act like fire-and-forget missiles.

Their playbook always goes like this.

ONE: First you seed fear. . .

It started with Sam the Scam fear mongering about 300 million job losses by 2030 — this was big tech’s narrative well int 2025. In 2026, their narrative shifted away from job loss to a more scaring narrative – the last stunt was somewhere mid-2025, when Mario’s brother, Dario, announced that by 2030, up to 50% of all white-collar entry-level jobs would disappear, leading to 10–20% structural unemployment.

Their tactics worked, and it got them the attention they needed to get VCs on board and invest further in developing new capabilities. They appealed to the lower emotions of people — fear and greed – but they couldn’t keep up the lie for long. Their dystopian story has been debunked by so many researchers recently, including myself, MIT, and the ILO.

In 2026 their narrative is no longer job loss.

No, it is now the fear of AI hacking its way into governments and companies. We used to talk about the military-industrial complex, but now it’s the AI-Trump complex, where government and Big Tech mutually support each other, simply because AI is the biggest driver of the US economy and they frequently support a president with donations.

So yeah, it didn’t come as a surprise to me when the US government (Trump, I mean) “banned” Anthropic from public use of Fable. It was just for a few weeks, but it got the world talking about their fears of its capabilities. Then Zhipu AI released GLM 5.2, OpenAI came out with 5.6, and the government just watched in silence. Shortly after, Fable returned, but infused with a government-approved seal of power.

Article content

TWO: Then you start giving impressive demos. . .

Every week I read up and watch the latest capabilities, and though they’re very impressive, hyped-up influencers cannot stop talking in superlatives about them (I’m sick of people like Julian Goldie, who’s always like “capability X is INSANE” and “IT’S MASSIVE / UNREAL / OVER” or “The SHOCKING truth about…”). This propaganda literally triggers my gag reflex.

Julian uses these headlines because he is an SEO guy and he knows what converts. And as you know, when you look at Facebook or YouTube, these algorithms all reward the lowest of emotions. Fear. Disgust. Anger. Because this is what gets you hooked.

And this works for the frontier guys. . .

Article content

And THREE: The AI “evangelists”.

When you’re an AI “evangelist” inside your organization, you regurgitate what these guys are saying, but through your internal platforms, and you’re pushing for AI adoption, using the knowledge of your organization to put it in context so it “sells” to the board. And since in the land of the blind one eye is king, these “evangelists” quickly get promoted to AI Officer, because to the board they are the interface between their view of the world and the scary world model of AI. This explains why so many nincompoops make it into higher management, in charge of huge transformational budgets, who get invited to join weekly dinner sessions sponsored by Big Tech *(yes, that happens) to chat with other ‘evangelists’, but without the necessary technical baggage to even grasp where they should start, nor what they should NOT touch, because the technology is simply not ready.

These organizations are set up for failure. And these are the ones that run AI programs that run in the red – value wise – because they simply aren’t able to convert this new discipline to their world because they’re simply lacking the required skills. The only skill they’ve mastered is the ‘corporate survival’ skill and that is what’s keeping them alive. Hiring the Big Four consulting clubs like McKinsey Quantum Black (I actually thought it was a credit card) to build them their toys, so they can blame them for when it blows up in their face.

Article content

Haven’t we learned from the Digital Transformation era?

I saw the same pattern back then, with influencers screaming at the top of their lungs that the world was going to change, people were going to be out of a job, and managers from the least successful business units were adopting this technology quickly to get some sort of advantage in the corporate career ladder (tip: look at what business units are on the right of the org chart and you can predict who has “already” adopted AI, meaning they have the shiny demos and the biggest stories of them all, and then you know that the general manager of that unit will become their first AI officer).

These freshly installed officers have great plans for setting up their AI programs, focusing on the “program governance”, “first principles”, the “intake process,” the “architecture,” “capability building”, “business case justification”, “value offices”, and a huge effort to build a security harness. At least in Europe that is the way I see these programs being run.

But what they all fail to do is that they don’t want to understand the technology. An architect is handy, but THEY need to have that knowledge in order to spot the opportunities and the pitfalls.

All these managers suffer from a bad case of Dunning-Kruger’s syndrome and a lot of them have come to believe their own personal branding too much, oversold their “technological literacy”, created huge business cases all based on Zone III (which is either too risky or simply doesn’t work from a technical perspective), and then they start running these huge programs with overstretched supply lines, and after one year in the making, they hand you a slide deck full of green KPIs and a production system that has never once completed a full run, and it’s on you to figure out how to explain that to the board.

Technically, most current AI programs are bankrupt. They do not generate value but continue anyway because “AI is all about compound multi-year effects” – regurgitating McKinsey’s latest report – or “We don’t know how to measure the results” is another one I’ve heard frequently.

Article content

And right before, while standing at the edge of the cliff, the industry saves them once again. They release the latest “major release,” promising the moon, sowing fear into management with yet another narrative, and then the evangelists translate this fear mongering into “early-adopter” opportunities, or that is at least what these internal Dunning Kruger types do, and so they are the ones contributing to the never-ending hype cycle, as I predicted would happen in five “spaces” about two years ago.

This is the death spiral we’re currently in.

Big Tech has created a tapered alpha helix that sucks in one company after the other, recombines their DNA with Big Tech’s, and churns and churns them until they have milked us all dry and the tumbleweeds roll through what used to be a market, past the husks of companies that mistook a subscription for a strategy.

But they are all leaving a crucial part untouched, what actually matters to us all, and not only to the companies out there, but to us consumers as well.

We are all recombined DNA with Big Tech’s.

We are in it for the ride, but if we don’t act, we will all suffer. What is still not proven is ROI. That one thing that matters to us all. Because if we aren’t able to prove that it can do what it has promised since 2023, we are all screwed, because the AI bubble will pop, investors will go broke, and companies that have spent millions or even billions on AI without ever understanding its intricate details will hit the brakes hard, rebalance in the best case, or completely divest their CAPEX investments in AI altogether, and since AI is the motor behind the US economy, and the US economy is, for some strange reason, still the motor behind the world economy, we will all go down the drain.

So yes, you and I will all suffer. Even if you’re a pipe-fitter or a teacher. The economy will be hit hard, people will have to sell their houses cheap (which in turn are being bought by the Warren Buffetts and Bill Gateses of this world), being in debt for the rest of their lives, and believe me, when this bubble bursts we will be worse off than in the Great Depression.

And we’re in a downward spiral even as Big Tech tries to prevent it with yet another update.

And the question is, can we turn the tide?

I think we can.

And the place to start is also the unsexy frontier. The enterprises.

The places that pay all our salaries and keep our fragile economies afloat.

We have to prove that we are able to create ROI, even though Big Tech doesn’t give a damn, and they don’t care, because they’re all about short-term gain. The billionaires wanting to cash out don’t buy bonds, they want highs and lows so they can hurly-burly their way from a billion to a trillion in valuation.

I have been writing about how to set up an agentification factory all along, and I’ve been making my research into improving the AI required for the enterprise public, and this is also why I started ATLAS – the Zone III long-horizon AI community – where AI practitioners can share their experiences, so we all can learn from it. I share it because if we don’t act, we will all suffer.

Yes, it is possible to generate ROI on your AI programs, but only when you know what NOT to do right now. I’ve written enough about it in my previous blogs to account for a whole book already, and maybe that should be my next step, but I’ve found that I am not a book-man. I’m a researcher who is working at the frontier of unsexy AI, and I eat my own dogfood – I implement the crap that I come up with – and that should be the case for most of the research that’s out there. Research for the sake of research isn’t worth diddly squat. Applied research is what we need in AI.

And that brings me to the core argument of this here blog. . .

I am done with language models.

They should be reduced to what they are. They are, in fact, language models.

Article content

But what they’re not is tools, and they aren’t workflows either. With all the apparent success of the harnesses out there, we have forgotten that what they’re inherently good at is generating text. Not operating the enterprise.

Agentic AI in its current form is going to fail.

And if you disagree, you’ve probably never tried to automate a high-stakes, 90-step business process spanning 20 different systems.

To be fair, neither have I. Because today, almost nobody has. Not successfully at scale, anyway.

Here’s why.

Assume your AI model has a 1 to 3% chance of making a mistake on each decision.

For a single-step task, that’s manageable. If something goes wrong, you detect it, roll it back, and move on. Now chain those decisions together. A 20-step workflow has only about a 54 to 82% chance of completing without a single error. That means there’s already an 18 to 46% chance that at least one step fails.

Now take a 90-step workflow.

Even with just a 1% error rate per step, the probability of at least one error rises to almost 60%. At a 3% error rate, it exceeds 93%. The problem isn’t so much that the model is terrible, but these small probabilities do compound, and every additional decision increases the chance that something eventually goes wrong.

Enterprise workflows aren’t one prediction. No, they’re hundreds of coordinated decisions, validations, API calls, approvals, and state transitions. That’s why making the model “a little smarter” won’t solve enterprise Agentic AI.

The architecture of the underlying model has to change.

We need a totally different paradigm to help the enterprise automate processes.

Remember the promise back in 2024? Generative AI doesn’t yield any ROI – I forgot the names of the research organizations that sold that to us – and then they all said in unison, “the next generation, agentic AI, will save the day.”

Well, my friend, I am here to say that it won’t.

Well, at least not when you accept that agentic ROI is fragile, and the only way you can get your program to show black digits instead of red is to be frugal (read my latest blogs on this topic) – or Dutch Tokenomics, as I’ve been calling it lately, finally putting a meme about the Dutch to good use – and also accept that Zone III agentic workflows (those where all the business cases are made) are simply impossible, and third, that we are at the end of the lifecycle of the language model, especially when it comes to enterprise operations.

There is a reason why the three gods of AI are all moving into world models.

And no, dear Dunning Kruger adept (not you), if you’re reading this, which I find highly unlikely, because you cannot handle the truth, and since your future is chained to that of Big Tech, a World Model is not going to be your savior.

A World Model I write with both capitals, because it is different from a “world model.” The latter is the “idea an AI has of the world.” In an LLM that is through text. In geometric AI that is through 3D objects, in AlphaFold that is through protein folding. An LLM doesn’t understand the world the way we do. They only see words and relationships between them. That’s what we call attention, and from there on they predict the most likely text matches. And they’re darned good at it, and they don’t know that a coffee mug has an ear, and when you rotate that… well, they don’t even know what “rotation” is all about, apart from the fact that it matches quite nicely with turning and 360 degrees and more of these synonyms.

A geometric AI, on the other hand, understands the concept of rotation. But only rotation and other 3D permutations and operations. Its world model is limited to the 3D world. AlphaFold is a combination of a few models, and it is tuned toward the folding of proteins, but it cannot write a single letter of prose. And equally so, an enterprise business process should have an AI with a model that “understands” the concepts of control and coordination like we do. And language is fine ‘n all, but it only understands the WORD business process and the WORD control and governance and adjacent concepts, but it does NOT have a concept of a coordination state.

So yeah, in this article I am going to talk about everything that is wrong with an LLM, that hinders our businesses’ ability to make money with AI. In total there are twelve reasons why a language model sucks for enterprises, and at the same time I make a case for a totally different model, geared toward enterprise workflows – the stuff that makes us all money – that can do the work repeatably, cost-effectively, with enough quality and in a controlled fashion. Those are exactly the three pillars that constitute our research program: cost, control, and quality. The trifecta. And all three influence one another.

Article content

Reason one. The math doesn’t care about your targets

Remember that little compounding math I dropped on you a few paragraphs ago? The 54 to 82% on a 20-step workflow, the near-60% failure rate at just 1% error per step on a 90-step chain? That wasn’t some rhetorical flourish I did to entertain you guys, no, that IS the whole ballgame, and almost no one who is building agentic workflows has internalized it, because it’s counterintuitive as hell. Your brain wants to average. Ninety-nine percent accurate step, that feels like it should stay in the high nineties forever. But nah, it doesn’t. Probabilities in a chain multiply, not average, and multiplication is a vicious little operator when you’re doing it two hundred times in a row. Take this hallucination rate of 2% on a single tool call, chain ten of those calls back to back the way any real agentic workflow does, database query into API call into document retrieval into decision node and around again, and by link ten you’re looking at over 1000% relative amplification. I know that number sounds like it can’t be real, and I had to check it three times before I believed it myself, and I don’t scare easy on math.

This can not be solved with a “wait for GPT-6” problem. This is an inherent problem of the architecture of the thing itself. A transformer that is generating the next token is, structurally, always rolling dice, and every additional roll in the chain is another chance for the dice to betray you. You simply cannot fine-tune your way out of this kind of multiplication. This is one of the reasons, why, in our factory, we’re calling it quits when the workflow is longer than 15 steps, after which a human is re-introduced. Zone II workflows are the only ones that can be trusted at this point in time, and as of yet, no one has convinced me of the opposite.

Article content

Reason two. I don’t know what you did last summer

Here’s what gets enterprise buyers every single time, because they’ve all used ChatGPT personally and it feels like it remembers stuff. It is actually the only reason why I keep ChatGPT on hand anyway. The harness memory is quite well developed and it offers lots of value, but the thing is that the model itself doesn’t have a recollection of what it did a nanosecond ago. A neural network starts every session from nearly zero, fills up with whatever you throw at it – usually the prompt plus thousands of lines of instructions the harness provides the model without you being aware of it (the context) – and then the whole thing evaporates the second the session closes. A neural network is stateless, and even inside one session, when the context window is technically still full of everything you fed it, the model doesn’t pay attention evenly across all of it. Stuff that you buried in the middle of a long document gets what researchers politely call “attention dilution” and what I call the model going cross-eyed. The 1M token context window you get with QWEN is a whiteboard that gets wiped clean between meetings, and even though the meeting is still going on, somebody’s erased the middle third of the whiteboard when you weren’t looking.

I once had a document pipeline fall over exactly this way. It was a beaut in the demo, because the demo used five-page contracts and the liability clause sat near the top where the model could actually see it. But once in a stress test, someone threw forty-page contracts at it, the liability clause landed on page twenty-two, dead center of the fog zone, and extraction accuracy on that specific clause dropped by more than half. These architectures just ain’t built to hold uniform attention across forty pages, and the spec sheet that says “1M tokens” doesn’t tell you that.

Article content

Reason three. It drifts in about six different ways

Let me name a few: Semantic drift. Agent drift. Goal drift. Constraint drift. Role drift. Belief drift. Six different research papers, that warned us about the doom called drift, and they are all six different flavors of the same underlying problem, that a system with no persistent anchor to its own original state will reconstruct what it thinks it’s doing based on whatever happens to be sitting in its context window at that exact moment.

When I re-read this sentence, I knew I had to provide an example.

Here you go . . .

Let’s take an agent that is processing an insurance claim. Step one, it gets its objective straight from the policy manual – “Approve claims only if all required documents are present and the payout is below €10,000. Otherwise escalate to a human reviewer”. Now, forty steps later it’s buried under thousands of tokens of emails, OCR output, customer chats, API responses, and its own previous reasoning, that original instruction is basically fossil record at this point. It is buried under so much stuff that it doesn’t know what deserves its attention and what not. What’s actually sitting front and center in the context window now is a different chorus entirely.

“Customer is waiting.” “Priority case.” “Previous claims were approved.” “Payment overdue.”

The model hasn’t gone stupid, nor has it forgotten anything in the way we forget a phone number or anything like that, no, it’s done something worse, it has rebuilt its own understanding of the job from whatever happens to be loudest in the room right now. Yes.. the loudest gets the attention. “My job is to enforce the approval policy” turns into something like “my job is to get this customer paid”. It’s a slow, unremarkable erosion of what the thing thought its own objective was.

Yup, that’s drift, the third reason on my list above, and it needs a stateless model, an ever-changing context window, and enough steps for the loudest recent signal to start outvoting the instruction it was actually given when it started. Take fifty steps, maybe seventy five, or a hundred even, and drift will bite you in the ass. It’s part of the architecture, and it’s what you get when you ask something with no persistent identity to keep reconstructing its identity from whatever word-salad is currently sitting in front of it.

One study* showed agents maintaining near-perfect goal adherence up to about 100,000 tokens before they start showing real susceptibility to pattern-matching nonsense that pulls them off course. Another one found semantic drift hitting nearly half of tested agents by 600 interactions. In multi-agent setups, constraint drift basically means that your safety rule can get lost or watered down as it passes hand to hand between agents. And to make things even worse, the rule was never technically violated in any single step. It simply stopped mattering somewhere along the chain.

And the darn thing is that you cannot patch drift. You can only build a system that doesn’t have the disease in the first place.

*_I actually found three separate papers on ATLAS that reported this. Here’s one: Goal drift, near-perfect adherence tot ~100k tokens: Rauno Arike, Elizabeth Donoway, Henning Bartsch, Marius Hobbhahn (2025), Technical Report: Evaluating Goal Drift in Language Model Agents.

Article content

Reason four. Making a plan and executing a plan are not the same skill

I get genuinely annoyed when people tell me “the model reasons through it”. No it doesn’t, not the way they think anyway. Give a state-of-the-art model a ten-city trip planning problem with realistic constraints and you can bet your ass the accuracy will fall below 5%. It can produce something plan-shaped, but holding coherence across constraints when the plan extends is a completely different muscle, and that muscle doesn’t exist yet in these architectures. Chain-of-thought prompting (yes, prompting), the whole “think step by step” magic trick everyone leans on, has been shown across sixteen frontier models to consistently underperform just answering directly once you push past the training distribution. Researchers are now calling it a mirage resembling reasoning, which is a phrase I wish I’d written myself.

There’s a structural reason to this.

Chain-of-thought prompting is simply step-by-step generation but with a fancier name, and it commits to its first move based on whatever looks plausible in that moment. And once it has made that call, it doesn’t go back and check it like a human would. It keeps building on top of that first guess, right or wrong.

Now, before the comment section explodes, yes, there are ways to make today’s agents far more reliable.

But notice where all the progress is happening.

Not inside the model.

Outside it.

The state of the art today relies on external memory, explicit state stores, symbolic constraint solvers, your search algorithms, external policy engines, deterministic validators, and continual replanning. Shall I go on? The harness is for the most part built to compensate for the weaknesses of the model. We increasingly surround the LLM with systems that remember, verify, constrain and reason structurally because the model itself doesn’t.

To be honest, these approaches work remarkably well. They make agents far more reliable than a naked LLM ever could.

But they also reveal the underlying limitation that if your solution requires half a dozen external systems to compensate for capabilities the architecture doesn’t naturally possess, you’ve improved the system.

You simply haven’t removed the bottleneck.

Article content

Reason five. Every tool call is a fresh opportunity to make something up

Tool use is where your money leaving the building, because a hallucination in a chat window is embarrassing but a hallucination inside an API call is a purchase order for the wrong currency.

Researchers building a taxonomy of tool execution hallucination found that this was happening consistently at the tool selection and calling stage and, worst of all, with errors cascading across iterations because the system has zero built-in sense of whether the task it just tried was even solvable with the tools that it had.

Oh, before I forget, here’s another one . . .

Xuannan Liu, PeiPei Li, and friends, the peeps behind the AgentHallu benchmark they built specifically to trace hallucinations back to the responsible step wrote a paper in which they stated that even the best frontier models, GPT-5, Gemini 2.5 Pro, could only correctly identify which step caused the failure 11.6% of the time. In the vast majority of failures, not anyone, not you nor the model, not eveb the logging system, could point at where it actually went wrong. You get a bad answer at the end of a ninety-step chain and something close to a coin flip on which of the ninety steps did it.

Article content

Reason six. The darn thing can’t even fix its own mistakes

The Darwin Gödel Machine machine with it’s promise of self-correction and self-optimization sounds solved because every demo now shows a model catching its own error mid-stream. But what the actual research is showing, is that this mostly requires an external prompt, meaning a human has to notice the error first and then tell it to fix itself.

Yes, mister DK, one layer deeper, your narrative just falls apart.

True autonomous error detection, the model catching itself without anyone pointing, is still a real gap. And here’s something that should worry all DK who are betting their roadmap on the next reasoning upgrade. There’s a documented, reproducible, method-agnostic relationship where boosting a model’s reasoning capability through reinforcement learning proportionally increases tool hallucination. Chenlong Yin and friends make this discovery in 2026 and published it in their paper “The reasoning trap: How enhancing LLM reasoning amplifies tool hallucination”.

I pause here for a second.

Yes. When you turn the hallucination dial down, utility goes down with it. When you turn utility up, hallucination goes up with it.

That problem is apparently baked into the training objective itself, which leads to the realization that the industry’s answer to every reliability complaint, “just make it smarter,” is fighting directly against the thing that determines whether you can trust it to run unsupervised.

When I got this paper through ATLAS, I had to sit down and sob for a while.

Article content

Reason seven. Stacking reliable pieces does not give you a reliable whole

There is this craft called systems engineering, but everyone forgets it the second the pieces are called “agents” instead of “(micro-)services.” Take five agent primitives, each individually excellent, each at 99% uptime, sounds like a rounding error you’d accept without blinking.

But remember reason one?

Stack five of them and your system reliability is not 99%, it’s 0.99 to the fifth power, roughly 95%.

This is a four-point drop, from components that each looked essentially flawless on their own, and every additional agent that bolted on in the name of clean architecture makes it worse.

Enterprise architects reach for multi-agent designs by instinct, because splitting a big task into specialized pieces feels like proper engineering way to do it and it is the same instinct that gives you clean microservices in normal software. Well, that instinct is correct in normal software, where every service behaves deterministically. But it is a true disaster here, because every additional agent is another probabilistic dice roll adding to the uncertainty.

Article content

Reason eight. More agents means more ways to be wrong together

Multi-agent gets pitched as the fix for single-agent weakness, and sure, in a curated benchmark it can even look like one.

But man, in production, coordination overhead does not scale gently.

Communication overhead in broadcast-style multi-agent setups grows quadratically with agent count. This means your infrastructure is hitting 50% overhead at just seven agents in one measured system.

Even hierarchical designs that dodge the worst of that still lose 23% task accuracy to information bottlenecks.

Mert Cemri, Melissa Pan, and about 11 more comrades give or take, did some research in their spare time (2025: “Why do multi-agent LLM systems fail”), where they built a taxonomy from over 1,600 annotated production traces, and they found fourteen distinct failure modes in multi-agent systems. Their blunt conclusion was that multi-agent performance gains on benchmarks are frequently minimal because these failures eat the theoretical benefit before it reaches anyone useful.

This even gets worse with small inaccuracies over iteration. It doesn’t stay small. It gets passed forward, then gets treated as ground truth, built on, then gets passed forward again, until it hardens into what researchers call false consensus, that means a system-wide agreement on something flatly wrong. And no single agent is ever definitively to blame because the error migrated through all of them.

I sat through a vendor pitch for exactly this setup once.

It had five agents, intent classification, retrieval, drafting, compliance review, formatting, each one pitched straight out of a software architecture slide deck. What the vendor did not mention, and what our testing told us, is that the compliance-review agent was reading input that had already passed through three probabilistic transformations, each distorting it a little, and the compliance agent had never once seen the original customer message. Just an increasingly (overconfident) summary of a summary of a summary.

Five competent agents produced a system worse than one agent doing the whole job, and all got a compliance stamp at the end that meant a lot less than it looked like.

Article content

Reason nine. The benchmarks are lying to our CAIO

This one makes me the most cynical, and I don’t scare easy on cynicism either. There are a lot of benchmarks out there. Too many actually. There’s GAIA, WebArena, most standard agentic benchmarks cap out around 20 steps, a fraction of the 50 to 200+ that a real Zone III enterprise workflow demands. A model looks like a champion on the leaderboard, when it is being tested on a task an order of magnitude easier than what it’ll face in your production environment. When SWE-Bench Pro raised the bar to cut down on contamination, frontier models that scored over 70% on the original SWE-bench Verified dropped to roughly 23%.

Yes. I kid you not.

GPT-5. Claude Opus 4. Both of them. The benchmark exposed that most of the original score was measuring benchmark familiarity of the model, and not actual capability.

Tanzila Kehkashan and friends, somewhere in 2026, did a review of fifteen major agent benchmarks (From benchmarks to deployment: a comprehensive review of agentic AI evaluation) and they found that zero of them scoring for security, zero scoring for cost efficiency, and thirteen relying purely on binary pass-fail, which throws away every bit of nuance about how close or how expensive or how risky an attempt actually was.

And on top of that, measurement noise itself, GAIA’s intraclass correlation coefficient swings between 0.304 and 0.774 depending on task structure, meaning a real chunk of what looks like a capability score is just luck of the sample.

So, next time your DK Chief AI Officer or your procurement team points at a leaderboard number to justify a Zone III business case, they are pointing at a number that was never designed to predict what happens at step 91 of your actual ERP integration.

Article content

Reason ten. It’s friggin’ expensive in a way you literally cannot budget for

Agentic tasks burn roughly a thousand times more tokens than a comparable code chat.

I could’ve stopped here.

By the way, it’s the input tokens doing the damage, not the output, because every step re-feeds the accumulated context back into the model. Fine, that would be tolerable if the cost were at least predictable. Buuuuuut, it’s not. The exact same task, same model, same prompt, has been measured to swing up to 30x in total token consumption run to run. Try building a business case around a cost that moves 30x between identical attempts.

And here’s the kick in the teeth, my smart friend. More tokens does not buy you more accuracy, so you can’t even reliably pay your way out of the reliability problem. A study on agentic coding reported that the iterative code review stage alone was eating 59.4% of total token consumption. That is a stage whose whole job is catching and fixing mistakes the rest of the pipeline made.

And that, my friend, tells you a lot about the quality of what the rest of the pipeline is producing in the first place.

Article content

Reason eleven. You cannot audit a probability distribution

The whole enterprise stick is revolves around the fact that someone, from the (external) audit department can prove afterward why it did what it did. GDPR, SOX, HIPAA, ISO 27001, the EU AI Act… pffff…. every single one of these Sisyphus frameworks simply assumes that a deterministic audit trail exists.

Um, yeah, an LLM’s output is, by construction, a sample that is drawn from a probability distribution, and there is no clean way to turn a statistically most probable next token (given the training and the context at that moment) into an explanation that a regulator or a judge actually needs.

Researchers describe this as the verification gap rooted in genuine computational intractability.

Explanation, it is expensive to audit an opaque model at the level of depth that these frameworks demand. Europe, with their AI Act, having legislation before the industry even uttered the word “agentic” has no clue of what they’re asking companies to do. . . Europe and their 60k bureaucrats are Zeus, and they invented these frameworks (the big Rock), that we, the practitioners (Sisyphus) have to carry with us up and down the Olympus, every time we build a new agent.

What’s making it worse is that the we’re actually runing into actual mathematical limits on what can be verified at all. Daniel Sunny and Ido with-a-long-name, in their ‘26 paper “A neuro-symbolic framework for legal accountability in public sector AI”, built a verification layer like my OCG to check if an AI-generated explanation actually held up under the law. They didn’t use a mock up dataset either, fifty real contested cases from California’s CalFresh welfare program.

Article content

The result was that the system was right 97.7% of the time on average. That sounds great of course, but the moment it flagged something as wrong, it could only correctly point at which law got broken 51 to 83% of the time, depending on the topic, and the weird thing was that sometimes the agency’s explanation read perfectly reasonable to a human, but it didn’t match the actual law at all.

So we have a verification problem. And “it sounded convincing” doesn’t hold up in court.

I’ve sat in enough audit prep meetings to know what actually happens here. Someone asks for the reasoning trail behind an agentic decision from four months ago, and what gets produced is a log of prompts and outputs, which is not an explanation. It only tells you what the model said, not why that continuation was more probable than the ones it didn’t choose. Every one of these meetings ends with a document that looks like an explanation and functions as a guess in a suit and tie.

That is why the Enterprise AI industry is hell-bent on explainability but, as I keep on saying, this is the wrong approach. It is trying to catch a thief, while you have to prevent theft in the first place. That is why the future of Enterprise AI lies in building an actionable knowledge layer, based on an ontology and expressed through a knowledge graph (Hi, @MARIO).

Article content

Reason twelve. The whole thing is an open door

Most AI security research to date focuses on single-turn attacks and jailbreaks or crimes like prompt injections. You know, the stuff that happens inside one exchange. But with agentic AI, we have long-horizon agentic workflows, and now we face an entirely different threat model.

Tanqiu Jiang and his homies (AgentLAB: Benchmarking LLM Agents against long-horizon attacks), tested this directly across 28 environments and 644 security test cases that covered attacks like intent hijacking, tool chaining abuse, task injection hidden inside ordinary-looking tool output, and deliberate objective drift induced by an attacker, and memory poisoning that corrupts what the agent believes to be true across sessions.

But none of the defenses that were actually built for single-turn interactions did anything meaningful against this, because the attack doesn’t live in one exchange, it lives in the accumulated trust built up over dozens or hundreds of steps. And yes, this is the exact same architectural surface that makes drift and hallucination and coordination failure possible in the first place. Every reason above that makes these systems unreliable is also, from an attacker’s point of view, a door somebody left wide open.

Article content

So what does actually work

Well, I could’ve gone on and on with bashing the LLM, but that wouldn’t be fair to it. And no, I am not preaching “throw the LLM in the sea,” to be clear, even though I’ve joked about that before.

This model is genuinely extraordinary at bounded, stateless tasks, in Zone I and Zone II work, you know, the stuff that covers a real chunk of what enterprises need, and I’ve built production systems on it since 2023 that do exactly what they’re supposed to inside that scope.

What all twelve reasons above point at is that the failure is architectural. It doesn’t scale away with the next parameter count, no matter what Sam “the Scam” or Dario AI-mon-ey tell you at the next launch event.

What actually closes the gap is a completely different animal.

We need a deterministic governance layer that treats compliance as load-bearing instead of a filter bolted on at the end. A symbolic reasoning engine that represents knowledge, rules, and goals in a form that can be checked, and certainly not a sampled thing. We also need a persistent state store that holds up across sessions instead of resetting to nothing every time. We also need formal verification that proves correctness before an irreversible action happens, instead of discovering the failure in a postmortem three weeks later.

Article content

Now, that is all part of a new type of AI, specifically built for Agentic workflows which I’m calling the Coordination Foundation Model. It is trained on coordination patterns instead of language, and it predicts the next valid state instead of the next plausible word. And when you pair it with something like the Ontological Compliance Gateway that I built that acts as a validation layer between probabilistic generation and deterministic execution, and then you actually catch the tool-use hallucination and the constraint drift before it becomes a purchase order in the wrong currency, instead of three weeks after.

I am not positioning this as a wondermodel that fixes everything by Q3, and I’m not selling you one. Encoding a real enterprise process into something a symbolic engine can actually check is slower and less glamorous than fine-tuning a prompt template, and it will not produce a demo that fits into a ninety-second sales call.

It means someone has to sit down and formally define what a valid state transition looks like for a given workflow, and it is truly an unglamorous piece of modeling work that most organizations have spent the last two years avoiding by throwing a bigger model at the problem instead.

But the uncomfortable truth is though, that modeling work was always going to be necessary, LLM-centric or not, because you cannot govern a process you’ve never actually specified. A Coordination Foundation Model just makes that specification the load-bearing wall instead of an afterthought stapled onto an agent’s system prompt.

So I built these things. They work. They’re the evolution of enterprise AI.

And then I saw the next revolution of AI. And it gave me the same feeling as I got in November 30th 2022 when ChatGPT 3.5 hit the market. And it is called Evolver .AI.

Do not watch their website, it sucks.

Their platform is the answer to all of this, and you won’t have to rethink your processes and put them in an ontology. This thing is next level.

And no, this will not be a write-up on it, because I will save that for the next blog, because it is worthy of having its own place in history, not as a footnote of an LLM diss.

Article content

So, I’m not saying you should abandon agentic AI with the current LLM’s. Our factory has proven that it is possible, because we know what NOT to do, and we keep our cost low (for more info, read previous blogs), and now there’s a way to target Zone III. If you don’t want to part with your current investments in your LLM setup/harness etc, you can use the CoordinationFM (foundational model) and the OCG. But when you want to “boldly go where no Enterprise has gone before”, you invite Mario, their Chief Scientist to talk to you about what he and his friends built. No, they’re certainly no small first-time Silicon Valley, Harry, Dick and Tom at the back of their shed type of startup, but instead they’re Nicola, Luis and Mario in their pristine white Research Lab, run by some of the most brilliant people in tech and business.

But don’t look at their website.

Just don’t.

It’s not bad, it’s beautiful even, and I get why they’ve chosen this route, but this is not telling the real story. That is what I leave for next week’s blog.

Signing off,

Marco


For my daytime job, I’m a researcher and factory builder at Eigenvector, a commercial research lab operating at the frontier of unsexy AI. In the evenings, I write about why building this stuff is a lot harder than the vendors would have you believe.

Article content

Become an AI Expert !

Sign up to receive insider articles in your inbox, every week.

✔️ We scour 75+ sources daily

✔️ Read by CEO, Scientists, Business Owners, and more

✔️ Join thousands of subscribers

✔️ No clickbait - 100% free

We don’t spam! Read our privacy policy for more info.

Leave a Reply

Up ↑

Discover more from TechTonic Shifts

Subscribe now to keep reading and get access to the full archive.

Continue reading