Draft essay

Why Token Maxing Doesn’t Move the Needle

More tokens can create more output. That does not mean they create more value. The missing step is conversion.

May 2026 · Working draft · Not yet published

I keep seeing a pattern that feels suspiciously familiar. A team starts spending more on AI, running more prompts, testing more models, buying more access, and watching more output pour out of the machine. The whole thing feels like momentum. It has the emotional texture of progress. Dashboards move. Drafts appear. Experiments multiply. You can practically hear the future whirring in the room. But that feeling is often doing a lot of fraudulent work. Because if you look one layer deeper, what you usually find is not leverage. You find activity.

That’s the trap I’d call token maxing. It’s the belief that because intelligence has become easier to produce, consuming more of it will naturally produce more economic value. And I don’t think that’s true. Or at least, not by default. More model calls can absolutely give you more summaries, more code suggestions, more synthetic analysis, more rough drafts, more options, more motion. But output and value are not the same thing. If the output never gets translated into a better decision, a shipped workflow improvement, a reusable internal asset, a measurable customer result, or an actual financial gain, then the extra output is mostly just expensive fog.

This is part of why the argument connects so directly to what I wrote in The Real Bottleneck in Agentic Engineering. In that piece, the problem wasn’t that agents couldn’t generate enough. The problem was that they could generate faster than a human could safely review. The bottleneck moved downstream. And that same shape keeps showing up here. The limit is not raw intelligence generation anymore. The limit is review, absorption, prioritization, trust, workflow change, and whether the organization can actually convert machine output into something the world will reward.

Blunt version: intelligence is getting cheap. Conversion is still hard. That’s where the value hides.

I’ve seen versions of this in the physical world too. When I wrote OpenClaw on a Robotic Shell, the hard part was not having a clever enough model sitting in a prompt box. The hard part was making the whole system survive contact with batteries, latency, control paths, and actual hardware constraints. Likewise, in Shellsensor Virtual World Version 3, the most interesting lesson was not that we could generate richer ideas with more tooling. It was that a cheaper, sharper harness changed what was actually worth building. In other words, the important thing was not abundance. It was which parts of the chain were still scarce and therefore still decisive.

That’s the part a lot of teams are about to learn the hard way. AI makes it incredibly easy to produce more. More content. More analysis. More candidate answers. More pseudo-work with a professional haircut. But the scarce things are still scarce: distribution, customer access, trust, judgment, operational discipline, and the willingness to rewire a workflow instead of admiring a demo. If those scarce elements do not change, then increasing token consumption mostly increases the volume of material passing through the system without materially changing the outcome of the system itself.

So when I look at AI spend, I think the useful question is no longer, “How much intelligence can we generate?” That question is getting cheaper by the month. The better question is, “Where does this convert?” Where does the draft become a shipped artifact? Where does the suggestion become a decision? Where does the analysis become a savings? Where does the experiment become a repeatable capability? If a team cannot point to that conversion layer with some precision, then the token bill is probably not evidence of transformation. It’s evidence that the machine is busy.

The companies that win here are probably not going to be the ones that brag the most about usage. They’ll be the ones that can take cheap intelligence and force it through the rest of the value chain: review, trust, workflow integration, customer impact, and economic capture. That’s where the real moat still is. Smart alone is not scarce anymore. Scarce is what happens after smart.

Related work referenced in this draft