Every technology ends up judged by the same ratio: how much is lost between the source of power and the useful work it does. A steam engine loses most of its heat to the boiler room. A car loses most of its fuel to friction and exhaust. A company loses most of its day to meetings, approvals, and waiting on someone else to finish their part first. Industrial progress is less a story about new sources of energy than about closing the gap between the energy spent and the work returned.

AI is running the same story, and the gap is being closed on three fronts at once: electricity into compute, compute into a model, a model into an answer. The fourth front, an answer into something a company actually does, has barely started. It is also the one that decides whether any of this shows up in results.

The hardware layer already ran this play

At GTC in March 2026, Jensen Huang compressed the economics of AI into one line: “if you have the wrong architecture, even if it’s free, it’s not cheap enough.” That is the friction argument in its purest form: the cost that matters is waste, not scarcity. Nvidia’s Vera Rubin architecture is organized around a single metric, tokens per watt, and every generation is aimed at cutting the loss between the wall socket and the finished token rather than adding more power.

One precision worth keeping: a token is a unit of output, not automatically a unit of intelligence. A model that needs ten tokens to do what another does in three is not more capable for having produced more of them. The defensible claim is narrower: the cost of producing useful intelligence is falling fast, and the infrastructure race is organized around making that decline steeper.

Cheaper intelligence multiplies demand

One layer up, the same hunt has already paid off. Sam Altman puts the decline at more than 10x per year for five years running, heading toward intelligence “too cheap to meter.” Argue with the exact pace if you like, the direction is not in dispute.

The instinct is to read that as good news that settles itself. Satya Nadella named why it doesn’t. When DeepSeek shipped comparable work at a fraction of the training cost, he posted one line: “Jevons paradox strikes again! As AI gets more efficient and accessible, we will see its use skyrocket, turning it into a commodity we just can’t get enough of.” Jevons’s observation, made first about coal and steam, is that efficiency increases consumption, because people spend the savings on doing more. Cheaper intelligence doesn’t shrink the amount of AI-driven work arriving at a company’s door, it multiplies it, faster than most companies are staffed to receive it.

As intelligence becomes abundant, the scarce resource shifts to the organization’s capacity to absorb it.

Meet the chokepoints

Here is what that looks like on the ground. MIT’s Project NANDA estimated in 2025 that roughly 95% of enterprise generative-AI pilots produced no measurable return, against tens of billions invested. Quibble with the number, the finding survives the quibbling. The cause was not model quality. It was the gap between what the tools can do and what organizations are built to let them do.

Name the gap precisely. Every piece of work an AI produces still has to squeeze through the same narrow passages the human version did: the review queue, the deployment pipeline, the data-access request, the identity check, the sign-off. These are chokepoints, and every one of them was sized for a worker who logs in once a morning and waits for a ticket. A model drafts the contract in seconds, and the draft then waits three days in the same approval chain as before. Individuals are seeing the promised 10x. Companies are seeing something closer to 20%. The missing 9.8x is sitting at the chokepoints.

A company moves at the speed of its chokepoints, not its people. That is not a new discovery. Eliyahu Goldratt built a whole management theory on it forty years ago. The Theory of Constraints, laid out in The Goal, begins from one observation: a system’s throughput is set by its constraint and by nothing else. His rule for factories transfers verbatim: an hour lost at a bottleneck is an hour lost for the entire system, and an hour saved at a non-bottleneck is a mirage.

Read the last decade through that rule and the MIT number stops being surprising. Faster chips, cheaper tokens, smarter models: spectacular savings, every one of them at a non-bottleneck. The gains are real at the stage where they happen and a mirage at the level of the company, because the constraint sits elsewhere. Cheap intelligence just made the constraint visible, since the chokepoints are now the only slow thing left. And the status-quo fix, hiring more reviewers and approvers to widen the chokepoints with humans, is the one move guaranteed not to scale against a workload that doubles again next year.

Goldratt made one more prediction, and it is the shape of this whole story: break a constraint and the constraint moves. It has moved three times already, electricity into compute, compute into intelligence, intelligence into action. The first two were technological constraints, attacked in public at conferences and in benchmark papers. The third is organizational, which is why it is the one still losing money for almost everyone running a pilot instead of a program.

A chokepoint is not judgment

The reflex under pressure is to treat this as a tradeoff: move fast by cutting the checks, or stay safe by keeping the queues. Wrong framing. Most of what sits at a chokepoint is queueing, not judgment: the manual approval nobody reads closely because forty more are waiting behind it, the standing permission granted once and never revisited, the review step that exists because it always has. A reviewer checking one request an hour was never going to keep pace with a system generating decisions continuously, and pretending otherwise is what turns a security step into a rubber stamp.

The judgment has to stay, and it has to get sharper as the volume grows. A high-risk action, real financial exposure, sensitive data, an irreversible change, deserves exactly the scrutiny it gets today. A routine request repeated a thousand times a day never needed a human, and every hour spent reviewing it is an hour not spent watching for the request that actually matters. Controls built for a slower, human-paced world don’t hold against one that runs continuously. The controls themselves have to run at the speed of the systems they govern, continuous and automatic, built to enable governed action rather than only to say no. That is a harder job than gatekeeping, not an easier one.

Cut the chokepoint, keep the judgment. Same people, no chokepoints. That distinction is the argument this entire series is built around.

What comes next

The hardware fight is public. The model fight is running. The enterprise fight has barely begun. Most companies have started experimenting, and far fewer have started redesigning around what AI actually makes possible. The winners of this transition are unlikely to be the ones that simply moved fastest. They will be the ones that figured out which friction was protecting them and which was only in the way, then removed only the second kind.

We spent the last decade making AI intelligent. The next few years will be about making companies fast enough to use it.